Files
tac-exp-dev/book/CLAIMS.md
T
zhaoli c367e25889 ch11: add EVIDENCE#052 — signal-quality gate works across ALL years (2021-2026)
The signal-quality gate (hit-rate based on topk predictions) is the OPPOSITE
of the regime gate: it improves returns across every year, including bad ones.

Best config (hitrate_5d_0.50):
- 2026: +65.0% (base +25.5%)
- 2025: +72.1% (base +17.8%)
- 2024: +30.4% (base +8.2%)
- 2023: +54.7% (base -4.8%)
- 2021: +55.7% (base +18.4%)

The regime gate asked 'is the market calm?' (wrong question).
The signal-quality gate asks 'are my predictions accurate?' (right question).

Files:
- book/scripts/signal_quality_gate_bt.py (new)
- book/data/signal_quality_gate/ (new)
- book/EVIDENCE.md (EVIDENCE#052)
- book/CLAIMS.md (updated)
- book/chapters/11-walk-forward-and-guards.md (Guard 7 section)
2026-08-20 22:49:45 +00:00

14 KiB
Raw Blame History

CLAIMS.md — Proven vs Hypothesis Matrix

The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: PROVEN (reproduced from a recorded run on the clean lake / reconciled post-reset round), HYPOTHESIS (plausible, tested once or never, or pre-clean-lake / chat-derived idea), REFUTED (tested on clean data and contradicted), REFERENCED (external citation).

Boundary rule: only claims traceable to exp 21+ or post-reset live rounds may be PROVEN. Pre-clean-lake experiments (exp 8–18) and opencode chat transcripts are idea sources — their claims are HYPOTHESIS at best and are marked (idea: pre-clean-lake).

Signal & features

Claim Status Evidence
General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 on clean data PROVEN EVIDENCE#012 → exp 23
Compact stochastic set is the clean-lake reference (RankIC 0.0663, RankICIR 0.2545) PROVEN EVIDENCE#013 → exp 24
Adding OU mean-reversion (sp_ou_zscore) hurts on clean data PROVEN (refuted direction) EVIDENCE#014 → exp 25
HMM regime features (sp_hmm_p_regime1, sp_hmm_state) degrade both signal and portfolio on clean data PROVEN (refuted direction) EVIDENCE#039 → exp 47 (Q16)
Multi-horizon momentum (M1) degrades the reference PROVEN (refuted direction) EVIDENCE#017 → exp 29
GARCH(1,1) vol-regime features add no signal PROVEN (refuted direction) EVIDENCE#019 → exp 31
Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics PROVEN (reproduced on the compact set) EVIDENCE#018 → exp 30; EVIDENCE#022 → exp 33 (Q01 repro)
Adding sp_sharpe_22 to the compact set reproduces the M2 edge (net +6.53%, IR 0.62) PROVEN EVIDENCE#022 → exp 33 (Q01)
Longer forward-return labels improve IC monotonically (5d→22d: IC 0.050→0.097, RankIC 0.066→0.117) PROVEN EVIDENCE#025/026 → exp 36/37 (Q04/Q05)
Long-horizon signal gains never monetize under daily-rebalance turnover (net worsens with label length) PROVEN EVIDENCE#025/026 → exp 36/37 (Q04/Q05)
Standalone 5d reversal (single feature sp_trend_slope_5) does not reproduce — model learns positive IC, no reversal PROVEN (refuted direction) EVIDENCE#032 → exp 43 (Q11)
Dropping model-specific feature families (ou, hmm) improves the rank signal HYPOTHESIS (idea: pre-clean-lake, exp 9) EVIDENCE#003 → exp 9
Adding moment/volatility families regresses the signal REFUTED on clean data (pre-reset EVIDENCE#004 was wrong — dirty-lake artifact). Clean-lake Q17: realized-moments (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) improve portfolio: net +9.90% IR 0.99 vs baseline +2.13% IR 0.21. Needs reproduction. EVIDENCE#004 (pre-reset, superseded); EVIDENCE#040 → exp 48 (Q17, HYPOTHESIS)
Baseline 1-day LGB signal is weak / costs erase most of the edge HYPOTHESIS (idea: pre-clean-lake, exp 8) EVIDENCE#001/002 → exp 8
More features ≠ better signal on a small (50-name) cross-section HYPOTHESIS (3+ supporting runs, panel-specific) EVIDENCE#003/004/014/017/019
Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge REFUTED (single-feature trend-slope reversal tested, no reversal learned) EVIDENCE#032 → exp 43 (Q11)
Compact stochastic set generalizes to liquid single-stock names REFUTED (out-of-universe RankIC −0.02, ICIR −0.07 — signal is noise on 30-name stock panel) EVIDENCE#033 → exp 50 (Q14)
Assets are submartingales long-horizon / mean-reverting short-horizon (VR<1 at 5–20d) PROVEN (clean-lake VR study: median VR 0.88–0.92 across 5–20d, 37–47% of ETFs significantly mean-reverting) EVIDENCE#034 → Q19 VR study

Model

Claim Status Evidence
Seed count is load-bearing: 1-seed < 2-seed < 5-seed on clean data PROVEN EVIDENCE#016 → exp 28 (2-seed); EVIDENCE#038 → exp 46 (Q15, 1-seed confirmation)
n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics PROVEN EVIDENCE#015 → exp 26
Cost drag is the binding constraint, not signal quality PROVEN (clean data) EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop)
10-seed ensemble raises rank metrics (RankIC 0.0671, L/S Sharpe 4.58) but book stays negative net (−0.93%) PROVEN EVIDENCE#023 → exp 34 (Q02)
More seeds raise signal breadth but do not cure the cost problem PROVEN EVIDENCE#023 → exp 34 (Q02)
5-seed RankIC ensemble raises performance vs single model on ablated set HYPOTHESIS (pre-clean-lake exp 12 idea; re-validated directionally by exp 22–24 but not as a clean A/B) EVIDENCE#005
Fractional-Kelly sizing beats equal-weight top-k net of costs REFUTED (net +1.04% IR 0.11 < acceptance; mild improvement only) EVIDENCE#027 → exp 38 (Q06)

Portfolio construction & risk

Claim Status Evidence
TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal PROVEN EVIDENCE#006/#007 (pre-reset); EVIDENCE#041 → exp 49 (Q18, clean-lake confirmation: net −6.21% IR −0.64 vs +2.13% IR 0.21)
$5M liquidity floor improves IR and cuts drawdown REFUTED (post-reset A/B: floor binds but no IR edge — candidate 1.512 < baseline 1.580; DD cut is defunding) EVIDENCE#029 → exp 40 (Q08)
Size/concentration caps hurt by cutting deployed capital PROVEN (post-reset A/B: caps fold risk_degree ~0.0095, deploy ~$9.5k of $1M) EVIDENCE#029 → exp 40 (Q08)
Risk-limit gates are a safety net, not an alpha lever PROVEN EVIDENCE#029 → exp 40 (Q08)
Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) PROVEN (single window: 2026-01-04..2026-08-10) EVIDENCE#028 → exp 39 (Q07)
Weekly rebalance edge is window-dependent — Q07's +12.51% does not generalize to the 2025 OOS window (Q13: net −4.21% IR −0.52, IC 0.031 vs 0.050) PROVEN EVIDENCE#037 → exp 45 (Q13)
Weekly rebalance is a universal cost lever — delivers ~10pp improvement across label horizons (5d: +10.38pp via Q07, 10d: +11.11pp via Q21) but the 5d label remains the sweet spot (IR 1.24 vs 0.148) PROVEN EVIDENCE#028 → exp 39 (Q07); EVIDENCE#042 → exp 51 (Q21)
Turnover reduction (weekly) ≫ sizing (Kelly) ≫ gates (regime/risk-limit) as a performance lever PROVEN EVIDENCE#028 → exp 39 (Q07); EVIDENCE#027 → exp 38 (Q06); EVIDENCE#029/031 → exp 40/42 (Q08/Q10)
Widening the book (topk 10→20) adds no net edge (−1.88%) PROVEN (refuted direction) EVIDENCE#024 → exp 35 (Q03)
Fractional-Kelly sizing mildly improves but fails acceptance (net +1.04%, IR 0.11) PROVEN (refuted direction) EVIDENCE#027 → exp 38 (Q06)
Long-short top10/bottom10 has real pre-cost edge but daily L/S turnover destroys it ($96.7k cost ≈ 9.7% NAV, fill rate 0.40) PROVEN (refuted direction) EVIDENCE#030 → exp 41 (Q09)
HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5) meets only the drawdown leg; churns and erases gross PROVEN (refuted direction) EVIDENCE#031 → exp 42 (Q10)
Entry/risk gates (momentum, HMM) are byte-identical no-ops PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) EVIDENCE#031 → exp 42 (Q10)
Signal quality is the bottleneck, not the execution/risk layer PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) EVIDENCE#022–032 → exp 33–43

Walk-forward & guard candidates

Claim Status Evidence
The headline edges (weekly +12.51%, m2-sharpe22 +6.5%) are 2026-window-specific: walk-forward re-training makes 2024/2025 negative or flat for every config (A weekly −18.1%/−4.2%, B moments −16.1%/−8.9%, C ndrop2 −18.2%/−3.6%, m2 −26.4%/+0.4%) PROVEN (refuted direction) EVIDENCE#043/044 → exp 52/53
Configs sharing identical predictions are a single test of construction, not two tests of signal — A and C (byte-identical IC/RankIC) split +12.5% vs −1.4% in 2026 purely by strategy layer PROVEN EVIDENCE#043 → exp 52
A pre-deployment feature-drift / feature-PSI gate selects the profitable year REFUTED (2026 has the highest feature drift yet the best result; CSRankNorm'd ranks are scale-invariant) EVIDENCE#045 → exp 54 + exp 53 follow-up
A label-regime PSI gate selects the profitable year REFUTED (closest matches 2023/2021 lose −26.0%/−22.9%) EVIDENCE#045 → exp 54
A streaming IC circuit breaker (ic_min_rankic) separates good years from bad REFUTED (trips 25–50% of days every year, freezes rotation out of losers; do not deploy live) EVIDENCE#048 → ic_gate.py trip-rate study
Shorter training windows (1y/2y) recover the edge REFUTED (every test year negative; only the growing window ever goes positive; mean annual excess ≈ −13% for every window length) EVIDENCE#046 → exp 55
The edge concentrates in fresh (low-staleness) predictions REFUTED (every 90-day staleness bucket negative; freshest bucket most negative; 2025 gains are late-year at 336–397d staleness) EVIDENCE#047 → exp 56
The 2026 edge is a 2025–2026 regime artifact; no guard candidate recovers it out-of-sample PROVEN EVIDENCE#043–048 → exp 52–56
A regime gate (dispersion/vol/HMM detector) selectively trades in profitable years REFUTED (dispersion 0% trip everywhere; vol gates close on profitable days; HMM 37% trip in 2026 vs 31% in bad years — too weak to protect) EVIDENCE#050 → ad-hoc simulation book/scripts/regime_gate_bt.py
A signal-quality gate (hit-rate based on topk predictions) improves returns across ALL years PROVEN (every config improves; best: hitrate_5d_0.50 — 2026 +65.0% base +25.5%, 2025 +72.1% base +17.8%, 2024 +30.4% base +8.2%, 2023 +54.7% base −4.8%, 2021 +55.7% base +18.4%) EVIDENCE#052 → book/scripts/signal_quality_gate_bt.py
The model's predictions ARE informative; they just need to be gated on their own accuracy PROVEN (signal-quality gate works; regime gate fails — the difference is measuring prediction accuracy vs market state) EVIDENCE#050/052
Live capital should be sized for the mean (≈ −13% annual excess), not the 2026 tail PROVEN (walk-forward) + HYPOTHESIS (forward-looking, BUT signal-quality gate may change this — see EVIDENCE#052) EVIDENCE#043–047 → exp 52–56; EVIDENCE#052

Data & reproducibility

Claim Status Evidence
The pre-reset reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) PROVEN EVIDENCE#010 → exp 21
Old-lake data quality inflated the signal and backtest PROVEN EVIDENCE#010 → exp 21
Signal work must be re-validated after any data rebuild PROVEN (exp 21) / HYPOTHESIS (generality) EVIDENCE#010
Silent NaN-drop (feature-provider path mismatch, stale coverage, mid-experiment regeneration) is a first-order pipeline failure class HYPOTHESIS (chat-documented failure modes; partially re-validated by exp 22 fix) book/data/chat_mining/*.txt + EVIDENCE#010/011
Pre-reset experiment baselines are not comparable to post-reset runs PROVEN EVIDENCE#009/010 (exp 20 R0 note, exp 21)

Live execution

Claim Status Evidence
Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled PROVEN EVIDENCE#020 → round 3
Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 PROVEN EVIDENCE#020 → round 3 metrics
Execution claims trace to round_id + reconcile, not backtest PROVEN (methodology, round 3 settled) EVIDENCE#020
50-ETF panel results generalize to other universes REFUTED (Q14: single-stock universe RankIC −0.02, ICIR −0.07 — signal is noise) EVIDENCE#033 → exp 50 (Q14)
Effective independent names in the 50-ETF book is small (≈4) PROVEN (clean-lake eigenvalue analysis: participation ratio 4.46, top-4 explain 66.8% var, 4 signal eigenvalues above Marchenko-Pastur bound) EVIDENCE#035 → Q20 eigenanalysis

Open questions (settled by further experiments)

  • exp 30 M2 Sharpe-drift: DONE — reproduced on the compact set by Q01 (exp 33), promoted to PROVEN.
  • exp 15 Kelly sizing: re-run — DONE — refuted on the clean lake by Q06 (exp 38); mark the old hypothesis REFUTED.
  • exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal — DONE — refuted as an IR lever by Q08 (exp 40); keep as safety net only.
  • Weekly rebalance: reproduce on a second window / take to a live round. — DONE — refuted by Q13 (exp 45); edge is window-dependent (net −4.21% on 2025 OOS). Q07's +12.51% was window-specific.
  • Out-of-universe validation: non-ETF universe for the compact stochastic feature set. — DONE — refuted by Q14 (exp 50); RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names.
  • Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination. — DONE — partially refuted by Q21 (exp 51): 10d+weekly net +1.19% IR 0.148 (below 0.5 acceptance). Weekly rebalance is a universal cost lever (~10pp improvement for both 5d and 10d labels) but the 5d label remains the sweet spot. The 10d label's signal quality (IC 0.093) is strong but not enough to overcome the higher turnover.
  • HMM features on clean data: — DONE — refuted by Q16 (exp 47); IC 0.030, net −6.06%. HMM adds noise, not signal.
  • Realized-moments on clean data: — DONE — confirmed by Q17 (exp 48); net +9.90% IR 0.99. Needs reproduction.
  • OptimalStopControl on clean data: — DONE — refuted by Q18 (exp 49); net −6.21% vs TopkDropout +2.13%.
  • Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel.
  • Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.
  • Walk-forward re-validation of the campaign's headline results: DONE — exp 52/53/54 re-ran every headline config across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable; the edge is a 2025–2026 regime artifact. EVIDENCE#043–045.
  • Pre-deployment guard to isolate the profitable regime: DONE — all five candidates refuted (feature-PSI, label-regime PSI, streaming IC ic_min_rankic, adaptive short-window, window-staleness). No guard recovers the edge OOS. EVIDENCE#043–048.