book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters

- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
This commit is contained in:
zhaoli
2026-08-20 01:23:21 +00:00
parent 436692a620
commit 06fb1e8ee9
15 changed files with 847 additions and 36 deletions
+26 -11
View File
@@ -13,12 +13,16 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | PROVEN (refuted direction) | EVIDENCE#014 → exp 25 |
| Multi-horizon momentum (M1) degrades the reference | PROVEN (refuted direction) | EVIDENCE#017 → exp 29 |
| GARCH(1,1) vol-regime features add no signal | PROVEN (refuted direction) | EVIDENCE#019 → exp 31 |
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | HYPOTHESIS (one clean-lake run, unreproduced) | EVIDENCE#018 → exp 30 |
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | PROVEN (reproduced on the compact set) | EVIDENCE#018 → exp 30; EVIDENCE#022 → exp 33 (Q01 repro) |
| Adding sp_sharpe_22 to the compact set reproduces the M2 edge (net +6.53%, IR 0.62) | PROVEN | EVIDENCE#022 → exp 33 (Q01) |
| Longer forward-return labels improve IC monotonically (5d→22d: IC 0.050→0.097, RankIC 0.066→0.117) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) |
| Long-horizon signal gains never monetize under daily-rebalance turnover (net worsens with label length) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) |
| Standalone 5d reversal (single feature sp_trend_slope_5) does not reproduce — model learns positive IC, no reversal | PROVEN (refuted direction) | EVIDENCE#032 → exp 43 (Q11) |
| Dropping model-specific feature families (ou, hmm) improves the rank signal | HYPOTHESIS (idea: pre-clean-lake, exp 9) | EVIDENCE#003 → exp 9 |
| Adding moment/volatility families regresses the signal | HYPOTHESIS (idea: pre-clean-lake, exp 11) | EVIDENCE#004 → exp 11 |
| Baseline 1-day LGB signal is weak / costs erase most of the edge | HYPOTHESIS (idea: pre-clean-lake, exp 8) | EVIDENCE#001/002 → exp 8 |
| More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3+ supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 |
| Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge | HYPOTHESIS (chat-derived clean-data study; see book/references/chat-ideas.md) | — |
| Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge | REFUTED (single-feature trend-slope reversal tested, no reversal learned) | EVIDENCE#032 → exp 43 (Q11) |
| Assets are submartingales long-horizon / mean-reverting short-horizon (VR<1 at 5–20d) | HYPOTHESIS (chat-derived martingale study, exp 19 never closed) | book/data/chat_mining/martingale-study.txt |
## Model
@@ -28,18 +32,27 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
| Seed count is load-bearing: 2 seeds < 5 seeds on clean data | PROVEN | EVIDENCE#016 → exp 28 |
| n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics | PROVEN | EVIDENCE#015 → exp 26 |
| Cost drag is the binding constraint, not signal quality | PROVEN (clean data) | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) |
| 10-seed ensemble raises rank metrics (RankIC 0.0671, L/S Sharpe 4.58) but book stays negative net (−0.93%) | PROVEN | EVIDENCE#023 → exp 34 (Q02) |
| More seeds raise signal breadth but do not cure the cost problem | PROVEN | EVIDENCE#023 → exp 34 (Q02) |
| 5-seed RankIC ensemble raises performance vs single model on ablated set | HYPOTHESIS (pre-clean-lake exp 12 idea; re-validated directionally by exp 22–24 but not as a clean A/B) | EVIDENCE#005 |
| Fractional-Kelly sizing beats equal-weight top-k net of costs | HYPOTHESIS (exp 15 never finished) | run never completed |
| Fractional-Kelly sizing beats equal-weight top-k net of costs | REFUTED (net +1.04% IR 0.11 < acceptance; mild improvement only) | EVIDENCE#027 → exp 38 (Q06) |
## Portfolio construction & risk
| Claim | Status | Evidence |
|-------|--------|----------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | HYPOTHESIS (idea: pre-clean-lake exp 13/14; not re-tested post-reset) | EVIDENCE#006/007 |
| $5M liquidity floor improves IR and cuts drawdown | HYPOTHESIS (idea: pre-clean-lake exp 18; not comparable post-reset) | EVIDENCE#008 → exp 18 |
| Size/concentration caps hurt by cutting deployed capital | HYPOTHESIS (idea: pre-clean-lake exp 18) | EVIDENCE#008 → exp 18 |
| Entry/risk gates (momentum, HMM) are byte-identical no-ops | HYPOTHESIS (idea: pre-clean-lake exp 20) | EVIDENCE#009 → exp 20 |
| Signal quality is the bottleneck, not the execution/risk layer | HYPOTHESIS (idea: pre-clean-lake exp 20; round-3 live is consistent but short) | EVIDENCE#009 → exp 20 |
| $5M liquidity floor improves IR and cuts drawdown | REFUTED (post-reset A/B: floor binds but no IR edge — candidate 1.512 < baseline 1.580; DD cut is defunding) | EVIDENCE#029 → exp 40 (Q08) |
| Size/concentration caps hurt by cutting deployed capital | PROVEN (post-reset A/B: caps fold risk_degree ~0.0095, deploy ~$9.5k of $1M) | EVIDENCE#029 → exp 40 (Q08) |
| Risk-limit gates are a safety net, not an alpha lever | PROVEN | EVIDENCE#029 → exp 40 (Q08) |
| Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) | PROVEN | EVIDENCE#028 → exp 39 (Q07) |
| Turnover reduction (weekly) ≫ sizing (Kelly) ≫ gates (regime/risk-limit) as a performance lever | PROVEN | EVIDENCE#028 → exp 39 (Q07); EVIDENCE#027 → exp 38 (Q06); EVIDENCE#029/031 → exp 40/42 (Q08/Q10) |
| Widening the book (topk 10→20) adds no net edge (−1.88%) | PROVEN (refuted direction) | EVIDENCE#024 → exp 35 (Q03) |
| Fractional-Kelly sizing mildly improves but fails acceptance (net +1.04%, IR 0.11) | PROVEN (refuted direction) | EVIDENCE#027 → exp 38 (Q06) |
| Long-short top10/bottom10 has real pre-cost edge but daily L/S turnover destroys it ($96.7k cost ≈ 9.7% NAV, fill rate 0.40) | PROVEN (refuted direction) | EVIDENCE#030 → exp 41 (Q09) |
| HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5) meets only the drawdown leg; churns and erases gross | PROVEN (refuted direction) | EVIDENCE#031 → exp 42 (Q10) |
| Entry/risk gates (momentum, HMM) are byte-identical no-ops | PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) | EVIDENCE#031 → exp 42 (Q10) |
| Signal quality is the bottleneck, not the execution/risk layer | PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) | EVIDENCE#022–032 → exp 33–43 |
## Data & reproducibility
@@ -63,7 +76,9 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
## Open questions (settled by further experiments)
- exp 30 M2 Sharpe-drift: reproduce on a second window before promoting past HYPOTHESIS.
- exp 15 Kelly sizing: re-run on the clean lake.
- exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal (exp 26 lineage).
- Out-of-universe validation: non-ETF universe for the compact stochastic feature set.
- exp 30 M2 Sharpe-drift: DONE — reproduced on the compact set by Q01 (exp 33), promoted to PROVEN.
- exp 15 Kelly sizing: re-run — DONE — refuted on the clean lake by Q06 (exp 38); mark the old hypothesis REFUTED.
- exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal — DONE — refuted as an IR lever by Q08 (exp 40); keep as safety net only.
- Weekly rebalance: reproduce on a second window / take to a live round.
- Out-of-universe validation: non-ETF universe for the compact stochastic feature set.
- Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination.