book: add EVIDENCE#036-041, update claims for Q12-Q18 findings

- EVIDENCE#036: Q12 22d label + weekly rebalance still fails (−4.88%)
- EVIDENCE#037: Q13 weekly rebalance edge is window-dependent (−4.21%)
- EVIDENCE#038: Q15 single-seed worse (clean-lake conf of #016)
- EVIDENCE#039: Q16 HMM features degrade on clean data (−6.06%)
- EVIDENCE#040: Q17 realized-moments improve portfolio (+9.90% IR 0.99)
- EVIDENCE#041: Q18 OptimalStopControl loses to TopkDropout (−6.21%)

Claims updated: H2 OptStop PROVEN, moments claim REFUTED on clean data,
weekly edge window-dependent PROVEN, HMM refuted, seed count updated.
Open questions marked done.
This commit is contained in:
zhaoli
2026-08-20 05:31:13 +00:00
parent 6145cfeb62
commit 63763db38c
2 changed files with 17 additions and 6 deletions
+11 -6
View File
@@ -11,6 +11,7 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
| General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 on clean data | PROVEN | EVIDENCE#012 → exp 23 |
| Compact stochastic set is the clean-lake reference (RankIC 0.0663, RankICIR 0.2545) | PROVEN | EVIDENCE#013 → exp 24 |
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | PROVEN (refuted direction) | EVIDENCE#014 → exp 25 |
| HMM regime features (sp_hmm_p_regime1, sp_hmm_state) degrade both signal and portfolio on clean data | PROVEN (refuted direction) | EVIDENCE#039 → exp 47 (Q16) |
| Multi-horizon momentum (M1) degrades the reference | PROVEN (refuted direction) | EVIDENCE#017 → exp 29 |
| GARCH(1,1) vol-regime features add no signal | PROVEN (refuted direction) | EVIDENCE#019 → exp 31 |
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | PROVEN (reproduced on the compact set) | EVIDENCE#018 → exp 30; EVIDENCE#022 → exp 33 (Q01 repro) |
@@ -19,7 +20,7 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
| Long-horizon signal gains never monetize under daily-rebalance turnover (net worsens with label length) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) |
| Standalone 5d reversal (single feature sp_trend_slope_5) does not reproduce — model learns positive IC, no reversal | PROVEN (refuted direction) | EVIDENCE#032 → exp 43 (Q11) |
| Dropping model-specific feature families (ou, hmm) improves the rank signal | HYPOTHESIS (idea: pre-clean-lake, exp 9) | EVIDENCE#003 → exp 9 |
| Adding moment/volatility families regresses the signal | HYPOTHESIS (idea: pre-clean-lake, exp 11) | EVIDENCE#004 → exp 11 |
| Adding moment/volatility families regresses the signal | REFUTED on clean data (pre-reset EVIDENCE#004 was wrong — dirty-lake artifact). Clean-lake Q17: realized-moments (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) improve portfolio: net +9.90% IR 0.99 vs baseline +2.13% IR 0.21. Needs reproduction. | EVIDENCE#004 (pre-reset, superseded); EVIDENCE#040 → exp 48 (Q17, HYPOTHESIS) |
| Baseline 1-day LGB signal is weak / costs erase most of the edge | HYPOTHESIS (idea: pre-clean-lake, exp 8) | EVIDENCE#001/002 → exp 8 |
| More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3+ supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 |
| Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge | REFUTED (single-feature trend-slope reversal tested, no reversal learned) | EVIDENCE#032 → exp 43 (Q11) |
@@ -30,7 +31,7 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
| Claim | Status | Evidence |
|-------|--------|----------|
| Seed count is load-bearing: 2 seeds < 5 seeds on clean data | PROVEN | EVIDENCE#016 → exp 28 |
| Seed count is load-bearing: 1-seed < 2-seed < 5-seed on clean data | PROVEN | EVIDENCE#016 → exp 28 (2-seed); EVIDENCE#038 → exp 46 (Q15, 1-seed confirmation) |
| n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics | PROVEN | EVIDENCE#015 → exp 26 |
| Cost drag is the binding constraint, not signal quality | PROVEN (clean data) | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) |
| 10-seed ensemble raises rank metrics (RankIC 0.0671, L/S Sharpe 4.58) but book stays negative net (−0.93%) | PROVEN | EVIDENCE#023 → exp 34 (Q02) |
@@ -42,11 +43,12 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
| Claim | Status | Evidence |
|-------|--------|----------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | HYPOTHESIS (idea: pre-clean-lake exp 13/14; not re-tested post-reset) | EVIDENCE#006/007 |
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | PROVEN | EVIDENCE#006/#007 (pre-reset); EVIDENCE#041 → exp 49 (Q18, clean-lake confirmation: net −6.21% IR −0.64 vs +2.13% IR 0.21) |
| $5M liquidity floor improves IR and cuts drawdown | REFUTED (post-reset A/B: floor binds but no IR edge — candidate 1.512 < baseline 1.580; DD cut is defunding) | EVIDENCE#029 → exp 40 (Q08) |
| Size/concentration caps hurt by cutting deployed capital | PROVEN (post-reset A/B: caps fold risk_degree ~0.0095, deploy ~$9.5k of $1M) | EVIDENCE#029 → exp 40 (Q08) |
| Risk-limit gates are a safety net, not an alpha lever | PROVEN | EVIDENCE#029 → exp 40 (Q08) |
| Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) | PROVEN | EVIDENCE#028 → exp 39 (Q07) |
| Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) | PROVEN (single window: 2026-01-04..2026-08-10) | EVIDENCE#028 → exp 39 (Q07) |
| Weekly rebalance edge is window-dependent — Q07's +12.51% does not generalize to the 2025 OOS window (Q13: net −4.21% IR −0.52, IC 0.031 vs 0.050) | PROVEN | EVIDENCE#037 → exp 45 (Q13) |
| Turnover reduction (weekly) ≫ sizing (Kelly) ≫ gates (regime/risk-limit) as a performance lever | PROVEN | EVIDENCE#028 → exp 39 (Q07); EVIDENCE#027 → exp 38 (Q06); EVIDENCE#029/031 → exp 40/42 (Q08/Q10) |
| Widening the book (topk 10→20) adds no net edge (−1.88%) | PROVEN (refuted direction) | EVIDENCE#024 → exp 35 (Q03) |
| Fractional-Kelly sizing mildly improves but fails acceptance (net +1.04%, IR 0.11) | PROVEN (refuted direction) | EVIDENCE#027 → exp 38 (Q06) |
@@ -80,8 +82,11 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
- exp 30 M2 Sharpe-drift: DONE — reproduced on the compact set by Q01 (exp 33), promoted to PROVEN.
- exp 15 Kelly sizing: re-run — DONE — refuted on the clean lake by Q06 (exp 38); mark the old hypothesis REFUTED.
- exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal — DONE — refuted as an IR lever by Q08 (exp 40); keep as safety net only.
- Weekly rebalance: reproduce on a second window / take to a live round.
- Weekly rebalance: reproduce on a second window / take to a live round. — DONE — refuted by Q13 (exp 45); edge is window-dependent (net −4.21% on 2025 OOS). Q07's +12.51% was window-specific.
- Out-of-universe validation: non-ETF universe for the compact stochastic feature set. — DONE — refuted by Q14 (exp 50); RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names.
- Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination.
- Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination. — NEXT (Q21)
- HMM features on clean data: — DONE — refuted by Q16 (exp 47); IC 0.030, net −6.06%. HMM adds noise, not signal.
- Realized-moments on clean data: — DONE — confirmed by Q17 (exp 48); net +9.90% IR 0.99. Needs reproduction.
- OptimalStopControl on clean data: — DONE — refuted by Q18 (exp 49); net −6.21% vs TopkDropout +2.13%.
- Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel.
- Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.
+6
View File
@@ -52,6 +52,12 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
| EVIDENCE#033 | Q14 out-of-universe validation: compact stochastic set on 30 liquid single-stock names (AAPL,MSFT,NVDA,…). RankIC −0.0198 (needed >0.03), ICIR −0.073 (needed >0.15) — signal is noise on this universe. Net P&L positive (+10.02% ann, IR 0.668, maxDD −6.67%) but that is top-10 concentration luck, not predictive signal. Train RankIC 0.316 shows the model overfits to the 50-ETF panel. | exp 50, run `809ff460…` (mlflow exp 50 `tac-rd-q14-out-of-universe`), branch `exp/50-q14-compact-stochastic-set-generalizes-t` | yes — Q14 FAIL (signal does not generalize cross-universe) |
| EVIDENCE#034 | Q19 variance-ratio study (Lo-MacKinlay robust VR): 71-ETF panel, 2015–2026. Median VR < 1 at all horizons — 5d: 0.925, 10d: 0.900, 20d: 0.884. 37–47% of ETFs have VR < 1 with |z| > 2 (significant mean-reversion). Only 1–3% show significant momentum. Assets are mean-reverting at short horizons on the clean lake. Note: pooled trend_slope_5 beta is strongly positive (+3.80, t=237) — the cross-sectional signal does NOT capture time-series mean-reversion. | scripted study, `book/data/evidence/q19-vr/vr_study.py`, VR_stats.csv, VR_summary.json | yes — Q19 PROVEN (market-structure claim) |
| EVIDENCE#035 | Q20 effective independent names: eigenvalue analysis on 71-ETF correlation matrix (test window 2026-01-04 to 2026-08-10). Participation ratio = 4.46. Top-4 eigenvalues explain 66.8% of variance. 4 eigenvalues above Marchenko-Pastur bound (2.86). The 50-ETF book has ≈4.5 effective independent names — confirming the chat-derived claim. This explains why topk 10→20 adds no breadth (EVIDENCE#024). | scripted study, `book/data/evidence/q20-effective-names/eigenanalysis.py`, eigenanalysis_50etf.csv, eigen_summary_50etf.json | yes — Q20 PROVEN (diversification claim) |
| EVIDENCE#036 | Q12 22d label + weekly rebalance: same IC/RankIC as Q05 (IC 0.097, RankIC 0.117 — identical training), but weekly recompute cannot rescue the stale signal. Net −4.88% (IR −0.566), gross +1.46%, maxDD −10.49%. The 22d label's problem is not daily turnover alone — the signal itself is stale. | exp 44, run `aed45c54…` (mlflow exp 44), branch `exp/44-q12-label22d-weekly` | yes — Q12 FAIL (redundant with Q05, confirms signal-stale hypothesis) |
| EVIDENCE#037 | Q13 weekly rebalance on 2025 OOS window (train→2024-08-30, test 2025-01-02..2025-12-31): edge is window-dependent. IC 0.031 (vs Q07's 0.050), RankIC 0.073 (vs 0.066), L/S Sharpe 1.19 (vs 4.54). Net −4.21% (IR −0.523), maxDD −10.66%. Q07's +12.51% (IR 1.24) was specific to the 2026-01-04..2026-08-10 window. Weekly rebalance is not a robust edge. | exp 45, run `e5ac7a5d…` (mlflow exp 45), branch `exp/45-q13-weekly-oos` | yes — Q13 FAIL (limits Q07's generalizability) |
| EVIDENCE#038 | Q15 single-seed vs 5-seed: 1 seed loses to 5 seeds on every metric. RankIC 0.044 vs 0.066, RankICIR 0.160 vs 0.255, net −2.89% (IR −0.278) vs +12.51% (IR 1.24). Clean-lake confirmation of EVIDENCE#016 (2-seed < 5-seed). Seed count is load-bearing. | exp 46, run `8d49e0be…` (mlflow exp 46), branch `exp/46-q15-single-seed` | yes — Q15 FAIL (confirms EVIDENCE#016) |
| EVIDENCE#039 | Q16 HMM features (sp_hmm_p_regime1, sp_hmm_state) on clean data: degrades both signal and portfolio. IC 0.030 (vs 0.050 baseline), RankIC 0.048 (vs 0.066), net −6.06% (IR −0.623), L/S Sharpe 1.23 (vs 4.54). HMM regime detection adds noise, not signal. | exp 47, run `ff092e1c…` (mlflow exp 47), branch `exp/47-q16-hmm` | yes — Q16 FAIL (HMM refuted on clean data) |
| EVIDENCE#040 | Q17 realized-moments features (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) on clean data: improves portfolio over baseline. Net +9.90% (IR 0.990), gross +14.61%, maxDD −6.49% vs baseline net +2.13% (IR 0.21). IC 0.039 (vs 0.050), RankIC 0.060 (vs 0.066) — signal metrics slightly lower but portfolio construction benefits from moment conditioning. Contradicts pre-reset EVIDENCE#004 (which was inflated by dirty data). Single run, unreproduced. | exp 48, run `e62ce326…` (mlflow exp 48), branch `exp/48-q17-moments` | yes — Q17 HYPOTHESIS (needs reproduction) |
| EVIDENCE#041 | Q18 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) vs TopkDropout on clean data: same signal (IC 0.050, RankIC 0.066 — identical model), worse portfolio. Net −6.21% (IR −0.640) vs baseline +2.13% (IR 0.21). Cost drag ~8.3pp. Clean-lake confirmation of pre-reset EVIDENCE#006/#007. | exp 49, run `f140dcb8…` (mlflow exp 49), branch `exp/49-q18-optstop` | yes — Q18 FAIL (confirms EVIDENCE#006/#007 on clean data) |
## Live execution trail