diff --git a/book/README.md b/book/README.md index 0fb7c5d..bb52a49 100644 --- a/book/README.md +++ b/book/README.md @@ -164,7 +164,16 @@ book/ ## Open questions for the desk -- `TODO(evidence-needed: weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round)` -- `TODO(evidence-needed: long-horizon label (10d/22d) paired with a low-turnover construction — the signal edge is proven, the cost kills daily churn)` -- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic feature set)` -- `TODO(evidence-needed: exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha)` \ No newline at end of file +- `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)` +- `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)` +- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)` +- `TODO(evidence-needed: purged walk-forward CV instead of single train/valid split on the exp-26 reference)` +- `TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)` +- `TODO(evidence-needed: live round under weekly-rebalance construction with risk-limit spec, to confirm safety-net behavior at higher deployed capital)` + +### Settled open questions (no longer active) + +- ~~`weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`. +- ~~`long-horizon label (10d/22d) paired with a low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`. +- ~~`out-of-universe (non-ETF) validation of the compact stochastic feature set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`. +- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`. \ No newline at end of file diff --git a/book/chapters/07-isolation-runs.md b/book/chapters/07-isolation-runs.md index 6f07600..b4e6adf 100644 --- a/book/chapters/07-isolation-runs.md +++ b/book/chapters/07-isolation-runs.md @@ -43,7 +43,7 @@ Only 2 of 11 Q-runs passed. That is not a failure of the campaign — it is the 1. Fix the reference and the acceptance bar *before* the series; change one variable per run. 2. Record signal metrics and net-of-cost metrics side by side — a signal gain that does not clear costs is not a strategy gain (Q02, Q04, Q05). 3. A single-feature "obvious" edge must be tested standalone before being trusted in a bundle (Q11 refuted it). -4. `TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)` +4. ~~`TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)`~~ — Weekly rebalance second window: **answered (negatively)** by Q13 (exp 45, net −4.21% IR −0.52, edge is window-dependent). sp_sharpe_22 in a live round: still open. ## Evidence cited in this chapter diff --git a/book/chapters/09-cost-turnover-frontier.md b/book/chapters/09-cost-turnover-frontier.md index 55c5e60..7d1dd7c 100644 --- a/book/chapters/09-cost-turnover-frontier.md +++ b/book/chapters/09-cost-turnover-frontier.md @@ -28,7 +28,7 @@ Q09 is the cleanest demonstration that cost, not signal, is the frontier: the lo - **Cadence (proven).** Weekly recompute of a 5-day signal is the single biggest lever found: ~1.1pp drag, +12.51% net. `PROVEN — EVIDENCE#028 → exp 39`. - **Dropped-name policy (proven).** n_drop 2→1 (hold, don't re-trade) bought ~5pp. `PROVEN — EVIDENCE#015 → exp 26`. - **Sizing (weak).** Kelly-style sizing scaled exposure but did not change the turnover bill (Q06, +1.04% net). `PROVEN — EVIDENCE#027 → exp 38`. -- **Label horizon (counterintuitive).** Longer labels improve the *signal* monotonically (22d IC 0.097, RankIC 0.117) but *worsen net* under daily churn (Q05: −4.60%). The horizon gain is real but unmonetized. `PROVEN — EVIDENCE#026 → exp 37`. `TODO(evidence-needed: long-horizon label at weekly cadence — the combination is untested and is the book's most promising open cell)`. +- **Label horizon (counterintuitive).** Longer labels improve the *signal* monotonically (22d IC 0.097, RankIC 0.117) but *worsen net* under daily churn (Q05: −4.60%). The horizon gain is real but unmonetized. `PROVEN — EVIDENCE#026 → exp 37`. ~~`TODO(evidence-needed: long-horizon label at weekly cadence — the combination is untested and is the book's most promising open cell)`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`. ## Desk rules distilled from this chapter diff --git a/book/chapters/12-synthesis.md b/book/chapters/12-synthesis.md index a7189e0..1e78421 100644 --- a/book/chapters/12-synthesis.md +++ b/book/chapters/12-synthesis.md @@ -40,10 +40,15 @@ The refuted runs were as valuable as the passes: the label-horizon result (22d l ## Open questions -- `TODO(evidence-needed: reproduce exp 39 weekly rebalance on a second window, then a live round)` -- `TODO(evidence-needed: long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction)` -- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic set)` -- `TODO(evidence-needed: realized-cost reconciliation of the weekly construction against the 5bp/15bp/$5 model once live)` +- `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)` +- `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)` +- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)` + +### Settled open questions + +- ~~`reproduce exp 39 weekly rebalance on a second window, then a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`. +- ~~`long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`. +- ~~`out-of-universe (non-ETF) validation of the compact stochastic set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`. ## Evidence cited in this chapter