book: resolve 4 open questions with Q12/Q13/Q14/Q21 evidence; update README + chapters 07/09/12

This commit is contained in:
zhaoli
2026-08-20 05:50:03 +00:00
parent 731a377f44
commit c0d21115eb
4 changed files with 24 additions and 10 deletions
+13 -4
View File
@@ -164,7 +164,16 @@ book/
## Open questions for the desk ## Open questions for the desk
- `TODO(evidence-needed: weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round)` - `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
- `TODO(evidence-needed: long-horizon label (10d/22d) paired with a low-turnover construction — the signal edge is proven, the cost kills daily churn)` - `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)`
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic feature set)` - `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
- `TODO(evidence-needed: exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha)` - `TODO(evidence-needed: purged walk-forward CV instead of single train/valid split on the exp-26 reference)`
- `TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`
- `TODO(evidence-needed: live round under weekly-rebalance construction with risk-limit spec, to confirm safety-net behavior at higher deployed capital)`
### Settled open questions (no longer active)
- ~~`weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`.
- ~~`long-horizon label (10d/22d) paired with a low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`.
- ~~`out-of-universe (non-ETF) validation of the compact stochastic feature set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`.
- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`.
+1 -1
View File
@@ -43,7 +43,7 @@ Only 2 of 11 Q-runs passed. That is not a failure of the campaign — it is the
1. Fix the reference and the acceptance bar *before* the series; change one variable per run. 1. Fix the reference and the acceptance bar *before* the series; change one variable per run.
2. Record signal metrics and net-of-cost metrics side by side — a signal gain that does not clear costs is not a strategy gain (Q02, Q04, Q05). 2. Record signal metrics and net-of-cost metrics side by side — a signal gain that does not clear costs is not a strategy gain (Q02, Q04, Q05).
3. A single-feature "obvious" edge must be tested standalone before being trusted in a bundle (Q11 refuted it). 3. A single-feature "obvious" edge must be tested standalone before being trusted in a bundle (Q11 refuted it).
4. `TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)` 4. ~~`TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)`~~ — Weekly rebalance second window: **answered (negatively)** by Q13 (exp 45, net −4.21% IR −0.52, edge is window-dependent). sp_sharpe_22 in a live round: still open.
## Evidence cited in this chapter ## Evidence cited in this chapter
+1 -1
View File
@@ -28,7 +28,7 @@ Q09 is the cleanest demonstration that cost, not signal, is the frontier: the lo
- **Cadence (proven).** Weekly recompute of a 5-day signal is the single biggest lever found: ~1.1pp drag, +12.51% net. `PROVEN — EVIDENCE#028 → exp 39`. - **Cadence (proven).** Weekly recompute of a 5-day signal is the single biggest lever found: ~1.1pp drag, +12.51% net. `PROVEN — EVIDENCE#028 → exp 39`.
- **Dropped-name policy (proven).** n_drop 2→1 (hold, don't re-trade) bought ~5pp. `PROVEN — EVIDENCE#015 → exp 26`. - **Dropped-name policy (proven).** n_drop 2→1 (hold, don't re-trade) bought ~5pp. `PROVEN — EVIDENCE#015 → exp 26`.
- **Sizing (weak).** Kelly-style sizing scaled exposure but did not change the turnover bill (Q06, +1.04% net). `PROVEN — EVIDENCE#027 → exp 38`. - **Sizing (weak).** Kelly-style sizing scaled exposure but did not change the turnover bill (Q06, +1.04% net). `PROVEN — EVIDENCE#027 → exp 38`.
- **Label horizon (counterintuitive).** Longer labels improve the *signal* monotonically (22d IC 0.097, RankIC 0.117) but *worsen net* under daily churn (Q05: −4.60%). The horizon gain is real but unmonetized. `PROVEN — EVIDENCE#026 → exp 37`. `TODO(evidence-needed: long-horizon label at weekly cadence — the combination is untested and is the book's most promising open cell)`. - **Label horizon (counterintuitive).** Longer labels improve the *signal* monotonically (22d IC 0.097, RankIC 0.117) but *worsen net* under daily churn (Q05: −4.60%). The horizon gain is real but unmonetized. `PROVEN — EVIDENCE#026 → exp 37`. ~~`TODO(evidence-needed: long-horizon label at weekly cadence — the combination is untested and is the book's most promising open cell)`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`.
## Desk rules distilled from this chapter ## Desk rules distilled from this chapter
+9 -4
View File
@@ -40,10 +40,15 @@ The refuted runs were as valuable as the passes: the label-horizon result (22d l
## Open questions ## Open questions
- `TODO(evidence-needed: reproduce exp 39 weekly rebalance on a second window, then a live round)` - `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
- `TODO(evidence-needed: long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction)` - `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)`
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic set)` - `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
- `TODO(evidence-needed: realized-cost reconciliation of the weekly construction against the 5bp/15bp/$5 model once live)`
### Settled open questions
- ~~`reproduce exp 39 weekly rebalance on a second window, then a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`.
- ~~`long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`.
- ~~`out-of-universe (non-ETF) validation of the compact stochastic set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`.
## Evidence cited in this chapter ## Evidence cited in this chapter