Files
tac-exp-dev/book/CLAIMS.md
T
zhaoli 63763db38c book: add EVIDENCE#036-041, update claims for Q12-Q18 findings
- EVIDENCE#036: Q12 22d label + weekly rebalance still fails (−4.88%)
- EVIDENCE#037: Q13 weekly rebalance edge is window-dependent (−4.21%)
- EVIDENCE#038: Q15 single-seed worse (clean-lake conf of #016)
- EVIDENCE#039: Q16 HMM features degrade on clean data (−6.06%)
- EVIDENCE#040: Q17 realized-moments improve portfolio (+9.90% IR 0.99)
- EVIDENCE#041: Q18 OptimalStopControl loses to TopkDropout (−6.21%)

Claims updated: H2 OptStop PROVEN, moments claim REFUTED on clean data,
weekly edge window-dependent PROVEN, HMM refuted, seed count updated.
Open questions marked done.
2026-08-20 05:31:13 +00:00

10 KiB
Raw Blame History

CLAIMS.md — Proven vs Hypothesis Matrix

The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: PROVEN (reproduced from a recorded run on the clean lake / reconciled post-reset round), HYPOTHESIS (plausible, tested once or never, or pre-clean-lake / chat-derived idea), REFUTED (tested on clean data and contradicted), REFERENCED (external citation).

Boundary rule: only claims traceable to exp 21+ or post-reset live rounds may be PROVEN. Pre-clean-lake experiments (exp 8–18) and opencode chat transcripts are idea sources — their claims are HYPOTHESIS at best and are marked (idea: pre-clean-lake).

Signal & features

Claim Status Evidence
General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 on clean data PROVEN EVIDENCE#012 → exp 23
Compact stochastic set is the clean-lake reference (RankIC 0.0663, RankICIR 0.2545) PROVEN EVIDENCE#013 → exp 24
Adding OU mean-reversion (sp_ou_zscore) hurts on clean data PROVEN (refuted direction) EVIDENCE#014 → exp 25
HMM regime features (sp_hmm_p_regime1, sp_hmm_state) degrade both signal and portfolio on clean data PROVEN (refuted direction) EVIDENCE#039 → exp 47 (Q16)
Multi-horizon momentum (M1) degrades the reference PROVEN (refuted direction) EVIDENCE#017 → exp 29
GARCH(1,1) vol-regime features add no signal PROVEN (refuted direction) EVIDENCE#019 → exp 31
Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics PROVEN (reproduced on the compact set) EVIDENCE#018 → exp 30; EVIDENCE#022 → exp 33 (Q01 repro)
Adding sp_sharpe_22 to the compact set reproduces the M2 edge (net +6.53%, IR 0.62) PROVEN EVIDENCE#022 → exp 33 (Q01)
Longer forward-return labels improve IC monotonically (5d→22d: IC 0.050→0.097, RankIC 0.066→0.117) PROVEN EVIDENCE#025/026 → exp 36/37 (Q04/Q05)
Long-horizon signal gains never monetize under daily-rebalance turnover (net worsens with label length) PROVEN EVIDENCE#025/026 → exp 36/37 (Q04/Q05)
Standalone 5d reversal (single feature sp_trend_slope_5) does not reproduce — model learns positive IC, no reversal PROVEN (refuted direction) EVIDENCE#032 → exp 43 (Q11)
Dropping model-specific feature families (ou, hmm) improves the rank signal HYPOTHESIS (idea: pre-clean-lake, exp 9) EVIDENCE#003 → exp 9
Adding moment/volatility families regresses the signal REFUTED on clean data (pre-reset EVIDENCE#004 was wrong — dirty-lake artifact). Clean-lake Q17: realized-moments (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) improve portfolio: net +9.90% IR 0.99 vs baseline +2.13% IR 0.21. Needs reproduction. EVIDENCE#004 (pre-reset, superseded); EVIDENCE#040 → exp 48 (Q17, HYPOTHESIS)
Baseline 1-day LGB signal is weak / costs erase most of the edge HYPOTHESIS (idea: pre-clean-lake, exp 8) EVIDENCE#001/002 → exp 8
More features ≠ better signal on a small (50-name) cross-section HYPOTHESIS (3+ supporting runs, panel-specific) EVIDENCE#003/004/014/017/019
Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge REFUTED (single-feature trend-slope reversal tested, no reversal learned) EVIDENCE#032 → exp 43 (Q11)
Compact stochastic set generalizes to liquid single-stock names REFUTED (out-of-universe RankIC −0.02, ICIR −0.07 — signal is noise on 30-name stock panel) EVIDENCE#033 → exp 50 (Q14)
Assets are submartingales long-horizon / mean-reverting short-horizon (VR<1 at 5–20d) PROVEN (clean-lake VR study: median VR 0.88–0.92 across 5–20d, 37–47% of ETFs significantly mean-reverting) EVIDENCE#034 → Q19 VR study

Model

Claim Status Evidence
Seed count is load-bearing: 1-seed < 2-seed < 5-seed on clean data PROVEN EVIDENCE#016 → exp 28 (2-seed); EVIDENCE#038 → exp 46 (Q15, 1-seed confirmation)
n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics PROVEN EVIDENCE#015 → exp 26
Cost drag is the binding constraint, not signal quality PROVEN (clean data) EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop)
10-seed ensemble raises rank metrics (RankIC 0.0671, L/S Sharpe 4.58) but book stays negative net (−0.93%) PROVEN EVIDENCE#023 → exp 34 (Q02)
More seeds raise signal breadth but do not cure the cost problem PROVEN EVIDENCE#023 → exp 34 (Q02)
5-seed RankIC ensemble raises performance vs single model on ablated set HYPOTHESIS (pre-clean-lake exp 12 idea; re-validated directionally by exp 22–24 but not as a clean A/B) EVIDENCE#005
Fractional-Kelly sizing beats equal-weight top-k net of costs REFUTED (net +1.04% IR 0.11 < acceptance; mild improvement only) EVIDENCE#027 → exp 38 (Q06)

Portfolio construction & risk

Claim Status Evidence
TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal PROVEN EVIDENCE#006/#007 (pre-reset); EVIDENCE#041 → exp 49 (Q18, clean-lake confirmation: net −6.21% IR −0.64 vs +2.13% IR 0.21)
$5M liquidity floor improves IR and cuts drawdown REFUTED (post-reset A/B: floor binds but no IR edge — candidate 1.512 < baseline 1.580; DD cut is defunding) EVIDENCE#029 → exp 40 (Q08)
Size/concentration caps hurt by cutting deployed capital PROVEN (post-reset A/B: caps fold risk_degree ~0.0095, deploy ~$9.5k of $1M) EVIDENCE#029 → exp 40 (Q08)
Risk-limit gates are a safety net, not an alpha lever PROVEN EVIDENCE#029 → exp 40 (Q08)
Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) PROVEN (single window: 2026-01-04..2026-08-10) EVIDENCE#028 → exp 39 (Q07)
Weekly rebalance edge is window-dependent — Q07's +12.51% does not generalize to the 2025 OOS window (Q13: net −4.21% IR −0.52, IC 0.031 vs 0.050) PROVEN EVIDENCE#037 → exp 45 (Q13)
Turnover reduction (weekly) ≫ sizing (Kelly) ≫ gates (regime/risk-limit) as a performance lever PROVEN EVIDENCE#028 → exp 39 (Q07); EVIDENCE#027 → exp 38 (Q06); EVIDENCE#029/031 → exp 40/42 (Q08/Q10)
Widening the book (topk 10→20) adds no net edge (−1.88%) PROVEN (refuted direction) EVIDENCE#024 → exp 35 (Q03)
Fractional-Kelly sizing mildly improves but fails acceptance (net +1.04%, IR 0.11) PROVEN (refuted direction) EVIDENCE#027 → exp 38 (Q06)
Long-short top10/bottom10 has real pre-cost edge but daily L/S turnover destroys it ($96.7k cost ≈ 9.7% NAV, fill rate 0.40) PROVEN (refuted direction) EVIDENCE#030 → exp 41 (Q09)
HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5) meets only the drawdown leg; churns and erases gross PROVEN (refuted direction) EVIDENCE#031 → exp 42 (Q10)
Entry/risk gates (momentum, HMM) are byte-identical no-ops PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) EVIDENCE#031 → exp 42 (Q10)
Signal quality is the bottleneck, not the execution/risk layer PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) EVIDENCE#022–032 → exp 33–43

Data & reproducibility

Claim Status Evidence
The pre-reset reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) PROVEN EVIDENCE#010 → exp 21
Old-lake data quality inflated the signal and backtest PROVEN EVIDENCE#010 → exp 21
Signal work must be re-validated after any data rebuild PROVEN (exp 21) / HYPOTHESIS (generality) EVIDENCE#010
Silent NaN-drop (feature-provider path mismatch, stale coverage, mid-experiment regeneration) is a first-order pipeline failure class HYPOTHESIS (chat-documented failure modes; partially re-validated by exp 22 fix) book/data/chat_mining/*.txt + EVIDENCE#010/011
Pre-reset experiment baselines are not comparable to post-reset runs PROVEN EVIDENCE#009/010 (exp 20 R0 note, exp 21)

Live execution

Claim Status Evidence
Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled PROVEN EVIDENCE#020 → round 3
Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 PROVEN EVIDENCE#020 → round 3 metrics
Execution claims trace to round_id + reconcile, not backtest PROVEN (methodology, round 3 settled) EVIDENCE#020
50-ETF panel results generalize to other universes REFUTED (Q14: single-stock universe RankIC −0.02, ICIR −0.07 — signal is noise) EVIDENCE#033 → exp 50 (Q14)
Effective independent names in the 50-ETF book is small (≈4) PROVEN (clean-lake eigenvalue analysis: participation ratio 4.46, top-4 explain 66.8% var, 4 signal eigenvalues above Marchenko-Pastur bound) EVIDENCE#035 → Q20 eigenanalysis

Open questions (settled by further experiments)

  • exp 30 M2 Sharpe-drift: DONE — reproduced on the compact set by Q01 (exp 33), promoted to PROVEN.
  • exp 15 Kelly sizing: re-run — DONE — refuted on the clean lake by Q06 (exp 38); mark the old hypothesis REFUTED.
  • exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal — DONE — refuted as an IR lever by Q08 (exp 40); keep as safety net only.
  • Weekly rebalance: reproduce on a second window / take to a live round. — DONE — refuted by Q13 (exp 45); edge is window-dependent (net −4.21% on 2025 OOS). Q07's +12.51% was window-specific.
  • Out-of-universe validation: non-ETF universe for the compact stochastic feature set. — DONE — refuted by Q14 (exp 50); RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names.
  • Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination. — NEXT (Q21)
  • HMM features on clean data: — DONE — refuted by Q16 (exp 47); IC 0.030, net −6.06%. HMM adds noise, not signal.
  • Realized-moments on clean data: — DONE — confirmed by Q17 (exp 48); net +9.90% IR 0.99. Needs reproduction.
  • OptimalStopControl on clean data: — DONE — refuted by Q18 (exp 49); net −6.21% vs TopkDropout +2.13%.
  • Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel.
  • Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.