Files
tac-exp-dev/book/CLAIMS.md
T

4.5 KiB
Raw Blame History

CLAIMS.md — Proven vs Hypothesis Matrix

The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: PROVEN (reproduced from recorded run / reconciled round), HYPOTHESIS (plausible, tested once or never), REFUTED (tested and contradicted), REFERENCED (external citation).

Signal & features

Claim Status Evidence
Baseline 1-day LGB signal is weak on 2026 OOS (RankIC 0.040, ICIR 0.062) PROVEN EVIDENCE#001 → exp 8
Costs erase most of the baseline edge (+6.2% gross → +1.6% net) PROVEN EVIDENCE#002 → exp 8
Dropping model-specific feature families (ou, hmm) improves rank signal (RankIC 0.030→0.064) PROVEN EVIDENCE#003 → exp 9
Adding moment/volatility families regresses the signal PROVEN (refuted direction) EVIDENCE#004 → exp 11
Adding OU mean-reversion (sp_ou_zscore) hurts on clean data PROVEN (refuted direction) EVIDENCE#014 → exp 25
Multi-horizon momentum (M1) degrades the reference PROVEN (refuted direction) EVIDENCE#017 → exp 29
GARCH(1,1) vol-regime features add no signal PROVEN (refuted direction) EVIDENCE#019 → exp 31
Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics HYPOTHESIS (one run, unreproduced) EVIDENCE#018 → exp 30
More features ≠ better signal on a small (50-name) cross-section HYPOTHESIS (3 supporting runs, panel-specific) EVIDENCE#003/004/014/017/019
General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 PROVEN EVIDENCE#012 → exp 23

Model

Claim Status Evidence
5-seed RankIC ensemble raises performance vs single model on ablated set PROVEN (pre-reset); re-validated post-reset exp 22–24 EVIDENCE#005/011/013
Seed count is load-bearing: 2 seeds < 5 seeds on clean data PROVEN EVIDENCE#016 → exp 28
n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics PROVEN EVIDENCE#015 → exp 26
Cost drag is the binding constraint, not signal quality PROVEN EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop)
Fractional-Kelly sizing beats equal-weight top-k net of costs HYPOTHESIS (exp 15 never finished) run never completed

Portfolio construction & risk

Claim Status Evidence
TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal PROVEN EVIDENCE#006/007 → exp 13/14
Stop-control churns and bleeds costs (−11.3pp cost drag) PROVEN EVIDENCE#006 → exp 13
$5M liquidity floor improves net IR (0.81→0.98) and cuts drawdown (7.9%→5.4%) PROVEN (pre-clean-lake; not comparable post-reset) EVIDENCE#008 → exp 18
Size/concentration caps hurt by cutting deployed capital PROVEN (pre-clean-lake) EVIDENCE#008 → exp 18
Entry/risk gates (momentum, HMM) are byte-identical no-ops on the reference signal PROVEN EVIDENCE#009 → exp 20
Signal quality is the bottleneck, not the execution/risk layer PROVEN (on the exp-20 reference) EVIDENCE#009 → exp 20

Data & reproducibility

Claim Status Evidence
The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) PROVEN EVIDENCE#010 → exp 21
Old-lake data quality inflated the signal and backtest PROVEN EVIDENCE#010 → exp 21
Signal work must be re-validated after any data rebuild PROVEN (exp 21) / HYPOTHESIS (generality) EVIDENCE#010
Pre-reset experiment baselines are not comparable to post-reset runs PROVEN EVIDENCE#009/010 (exp 20 R0 note, exp 21)

Live execution

Claim Status Evidence
Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled PROVEN EVIDENCE#020 → round 3
Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 PROVEN EVIDENCE#020 → round 3 metrics
Execution claims trace to round_id + reconcile, not backtest PROVEN (methodology, round 3 settled) EVIDENCE#020
50-ETF panel results generalize to other universes HYPOTHESIS — TODO(evidence-needed) —

Open questions (settled by further experiments)

  • exp 30 M2 Sharpe-drift: reproduce on a second window before promoting past HYPOTHESIS.
  • exp 15 Kelly sizing: re-run on the clean lake.
  • exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal (exp 26 lineage).
  • Out-of-universe validation: non-ETF universe for the compact stochastic feature set.