# CLAIMS.md — Proven vs Hypothesis Matrix The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: `PROVEN` (reproduced from a recorded run **on the clean lake** / reconciled post-reset round), `HYPOTHESIS` (plausible, tested once or never, or pre-clean-lake / chat-derived idea), `REFUTED` (tested on clean data and contradicted), `REFERENCED` (external citation). **Boundary rule:** only claims traceable to exp 21+ or post-reset live rounds may be `PROVEN`. Pre-clean-lake experiments (exp 8–18) and opencode chat transcripts are idea sources — their claims are `HYPOTHESIS` at best and are marked `(idea: pre-clean-lake)`. ## Signal & features | Claim | Status | Evidence | |-------|--------|----------| | General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 on clean data | PROVEN | EVIDENCE#012 → exp 23 | | Compact stochastic set is the clean-lake reference (RankIC 0.0663, RankICIR 0.2545) | PROVEN | EVIDENCE#013 → exp 24 | | Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | PROVEN (refuted direction) | EVIDENCE#014 → exp 25 | | HMM regime features (sp_hmm_p_regime1, sp_hmm_state) degrade both signal and portfolio on clean data | PROVEN (refuted direction) | EVIDENCE#039 → exp 47 (Q16) | | Multi-horizon momentum (M1) degrades the reference | PROVEN (refuted direction) | EVIDENCE#017 → exp 29 | | GARCH(1,1) vol-regime features add no signal | PROVEN (refuted direction) | EVIDENCE#019 → exp 31 | | Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | PROVEN (reproduced on the compact set) | EVIDENCE#018 → exp 30; EVIDENCE#022 → exp 33 (Q01 repro) | | Adding sp_sharpe_22 to the compact set reproduces the M2 edge (net +6.53%, IR 0.62) | PROVEN | EVIDENCE#022 → exp 33 (Q01) | | Longer forward-return labels improve IC monotonically (5d→22d: IC 0.050→0.097, RankIC 0.066→0.117) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) | | Long-horizon signal gains never monetize under daily-rebalance turnover (net worsens with label length) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) | | Standalone 5d reversal (single feature sp_trend_slope_5) does not reproduce — model learns positive IC, no reversal | PROVEN (refuted direction) | EVIDENCE#032 → exp 43 (Q11) | | Dropping model-specific feature families (ou, hmm) improves the rank signal | HYPOTHESIS (idea: pre-clean-lake, exp 9) | EVIDENCE#003 → exp 9 | | Adding moment/volatility families regresses the signal | REFUTED on clean data (pre-reset EVIDENCE#004 was wrong — dirty-lake artifact). Clean-lake Q17: realized-moments (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) improve portfolio: net +9.90% IR 0.99 vs baseline +2.13% IR 0.21. Needs reproduction. | EVIDENCE#004 (pre-reset, superseded); EVIDENCE#040 → exp 48 (Q17, HYPOTHESIS) | | Baseline 1-day LGB signal is weak / costs erase most of the edge | HYPOTHESIS (idea: pre-clean-lake, exp 8) | EVIDENCE#001/002 → exp 8 | | More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3+ supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 | | Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge | REFUTED (single-feature trend-slope reversal tested, no reversal learned) | EVIDENCE#032 → exp 43 (Q11) | | Compact stochastic set generalizes to liquid single-stock names | REFUTED (out-of-universe RankIC −0.02, ICIR −0.07 — signal is noise on 30-name stock panel) | EVIDENCE#033 → exp 50 (Q14) | | Assets are submartingales long-horizon / mean-reverting short-horizon (VR<1 at 5–20d) | PROVEN (clean-lake VR study: median VR 0.88–0.92 across 5–20d, 37–47% of ETFs significantly mean-reverting) | EVIDENCE#034 → Q19 VR study | ## Model | Claim | Status | Evidence | |-------|--------|----------| | Seed count is load-bearing: 1-seed < 2-seed < 5-seed on clean data | PROVEN | EVIDENCE#016 → exp 28 (2-seed); EVIDENCE#038 → exp 46 (Q15, 1-seed confirmation) | | n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics | PROVEN | EVIDENCE#015 → exp 26 | | Cost drag is the binding constraint, not signal quality | PROVEN (clean data) | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) | | 10-seed ensemble raises rank metrics (RankIC 0.0671, L/S Sharpe 4.58) but book stays negative net (−0.93%) | PROVEN | EVIDENCE#023 → exp 34 (Q02) | | More seeds raise signal breadth but do not cure the cost problem | PROVEN | EVIDENCE#023 → exp 34 (Q02) | | 5-seed RankIC ensemble raises performance vs single model on ablated set | HYPOTHESIS (pre-clean-lake exp 12 idea; re-validated directionally by exp 22–24 but not as a clean A/B) | EVIDENCE#005 | | Fractional-Kelly sizing beats equal-weight top-k net of costs | REFUTED (net +1.04% IR 0.11 < acceptance; mild improvement only) | EVIDENCE#027 → exp 38 (Q06) | ## Portfolio construction & risk | Claim | Status | Evidence | |-------|--------|----------| | TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | PROVEN | EVIDENCE#006/#007 (pre-reset); EVIDENCE#041 → exp 49 (Q18, clean-lake confirmation: net −6.21% IR −0.64 vs +2.13% IR 0.21) | | $5M liquidity floor improves IR and cuts drawdown | REFUTED (post-reset A/B: floor binds but no IR edge — candidate 1.512 < baseline 1.580; DD cut is defunding) | EVIDENCE#029 → exp 40 (Q08) | | Size/concentration caps hurt by cutting deployed capital | PROVEN (post-reset A/B: caps fold risk_degree ~0.0095, deploy ~$9.5k of $1M) | EVIDENCE#029 → exp 40 (Q08) | | Risk-limit gates are a safety net, not an alpha lever | PROVEN | EVIDENCE#029 → exp 40 (Q08) | | Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) | PROVEN (single window: 2026-01-04..2026-08-10) | EVIDENCE#028 → exp 39 (Q07) | | Weekly rebalance edge is window-dependent — Q07's +12.51% does not generalize to the 2025 OOS window (Q13: net −4.21% IR −0.52, IC 0.031 vs 0.050) | PROVEN | EVIDENCE#037 → exp 45 (Q13) | | Weekly rebalance is a universal cost lever — delivers ~10pp improvement across label horizons (5d: +10.38pp via Q07, 10d: +11.11pp via Q21) but the 5d label remains the sweet spot (IR 1.24 vs 0.148) | PROVEN | EVIDENCE#028 → exp 39 (Q07); EVIDENCE#042 → exp 51 (Q21) | | Turnover reduction (weekly) ≫ sizing (Kelly) ≫ gates (regime/risk-limit) as a performance lever | PROVEN | EVIDENCE#028 → exp 39 (Q07); EVIDENCE#027 → exp 38 (Q06); EVIDENCE#029/031 → exp 40/42 (Q08/Q10) | | Widening the book (topk 10→20) adds no net edge (−1.88%) | PROVEN (refuted direction) | EVIDENCE#024 → exp 35 (Q03) | | Fractional-Kelly sizing mildly improves but fails acceptance (net +1.04%, IR 0.11) | PROVEN (refuted direction) | EVIDENCE#027 → exp 38 (Q06) | | Long-short top10/bottom10 has real pre-cost edge but daily L/S turnover destroys it ($96.7k cost ≈ 9.7% NAV, fill rate 0.40) | PROVEN (refuted direction) | EVIDENCE#030 → exp 41 (Q09) | | HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5) meets only the drawdown leg; churns and erases gross | PROVEN (refuted direction) | EVIDENCE#031 → exp 42 (Q10) | | Entry/risk gates (momentum, HMM) are byte-identical no-ops | PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) | EVIDENCE#031 → exp 42 (Q10) | | Signal quality is the bottleneck, not the execution/risk layer | PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) | EVIDENCE#022–032 → exp 33–43 | ## Data & reproducibility | Claim | Status | Evidence | |-------|--------|----------| | The pre-reset reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | PROVEN | EVIDENCE#010 → exp 21 | | Old-lake data quality inflated the signal and backtest | PROVEN | EVIDENCE#010 → exp 21 | | Signal work must be re-validated after any data rebuild | PROVEN (exp 21) / HYPOTHESIS (generality) | EVIDENCE#010 | | Silent NaN-drop (feature-provider path mismatch, stale coverage, mid-experiment regeneration) is a first-order pipeline failure class | HYPOTHESIS (chat-documented failure modes; partially re-validated by exp 22 fix) | book/data/chat_mining/*.txt + EVIDENCE#010/011 | | Pre-reset experiment baselines are not comparable to post-reset runs | PROVEN | EVIDENCE#009/010 (exp 20 R0 note, exp 21) | ## Live execution | Claim | Status | Evidence | |-------|--------|----------| | Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled | PROVEN | EVIDENCE#020 → round 3 | | Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 | PROVEN | EVIDENCE#020 → round 3 metrics | | Execution claims trace to round_id + reconcile, not backtest | PROVEN (methodology, round 3 settled) | EVIDENCE#020 | | 50-ETF panel results generalize to other universes | REFUTED (Q14: single-stock universe RankIC −0.02, ICIR −0.07 — signal is noise) | EVIDENCE#033 → exp 50 (Q14) | | Effective independent names in the 50-ETF book is small (≈4) | PROVEN (clean-lake eigenvalue analysis: participation ratio 4.46, top-4 explain 66.8% var, 4 signal eigenvalues above Marchenko-Pastur bound) | EVIDENCE#035 → Q20 eigenanalysis | ## Open questions (settled by further experiments) - exp 30 M2 Sharpe-drift: DONE — reproduced on the compact set by Q01 (exp 33), promoted to PROVEN. - exp 15 Kelly sizing: re-run — DONE — refuted on the clean lake by Q06 (exp 38); mark the old hypothesis REFUTED. - exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal — DONE — refuted as an IR lever by Q08 (exp 40); keep as safety net only. - Weekly rebalance: reproduce on a second window / take to a live round. — DONE — refuted by Q13 (exp 45); edge is window-dependent (net −4.21% on 2025 OOS). Q07's +12.51% was window-specific. - Out-of-universe validation: non-ETF universe for the compact stochastic feature set. — DONE — refuted by Q14 (exp 50); RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. - Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination. — DONE — partially refuted by Q21 (exp 51): 10d+weekly net +1.19% IR 0.148 (below 0.5 acceptance). Weekly rebalance is a universal cost lever (~10pp improvement for both 5d and 10d labels) but the 5d label remains the sweet spot. The 10d label's signal quality (IC 0.093) is strong but not enough to overcome the higher turnover. - HMM features on clean data: — DONE — refuted by Q16 (exp 47); IC 0.030, net −6.06%. HMM adds noise, not signal. - Realized-moments on clean data: — DONE — confirmed by Q17 (exp 48); net +9.90% IR 0.99. Needs reproduction. - OptimalStopControl on clean data: — DONE — refuted by Q18 (exp 49); net −6.21% vs TopkDropout +2.13%. - Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel. - Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.