Files
tac-exp-dev/book/CLAIMS.md
T

87 lines
9.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAIMS.md — Proven vs Hypothesis Matrix
The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: `PROVEN` (reproduced from a recorded run **on the clean lake** / reconciled post-reset round), `HYPOTHESIS` (plausible, tested once or never, or pre-clean-lake / chat-derived idea), `REFUTED` (tested on clean data and contradicted), `REFERENCED` (external citation).
**Boundary rule:** only claims traceable to exp 21+ or post-reset live rounds may be `PROVEN`. Pre-clean-lake experiments (exp 8–18) and opencode chat transcripts are idea sources — their claims are `HYPOTHESIS` at best and are marked `(idea: pre-clean-lake)`.
## Signal & features
| Claim | Status | Evidence |
|-------|--------|----------|
| General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 on clean data | PROVEN | EVIDENCE#012 → exp 23 |
| Compact stochastic set is the clean-lake reference (RankIC 0.0663, RankICIR 0.2545) | PROVEN | EVIDENCE#013 → exp 24 |
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | PROVEN (refuted direction) | EVIDENCE#014 → exp 25 |
| Multi-horizon momentum (M1) degrades the reference | PROVEN (refuted direction) | EVIDENCE#017 → exp 29 |
| GARCH(1,1) vol-regime features add no signal | PROVEN (refuted direction) | EVIDENCE#019 → exp 31 |
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | PROVEN (reproduced on the compact set) | EVIDENCE#018 → exp 30; EVIDENCE#022 → exp 33 (Q01 repro) |
| Adding sp_sharpe_22 to the compact set reproduces the M2 edge (net +6.53%, IR 0.62) | PROVEN | EVIDENCE#022 → exp 33 (Q01) |
| Longer forward-return labels improve IC monotonically (5d→22d: IC 0.050→0.097, RankIC 0.066→0.117) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) |
| Long-horizon signal gains never monetize under daily-rebalance turnover (net worsens with label length) | PROVEN | EVIDENCE#025/026 → exp 36/37 (Q04/Q05) |
| Standalone 5d reversal (single feature sp_trend_slope_5) does not reproduce — model learns positive IC, no reversal | PROVEN (refuted direction) | EVIDENCE#032 → exp 43 (Q11) |
| Dropping model-specific feature families (ou, hmm) improves the rank signal | HYPOTHESIS (idea: pre-clean-lake, exp 9) | EVIDENCE#003 → exp 9 |
| Adding moment/volatility families regresses the signal | HYPOTHESIS (idea: pre-clean-lake, exp 11) | EVIDENCE#004 → exp 11 |
| Baseline 1-day LGB signal is weak / costs erase most of the edge | HYPOTHESIS (idea: pre-clean-lake, exp 8) | EVIDENCE#001/002 → exp 8 |
| More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3+ supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 |
| Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge | REFUTED (single-feature trend-slope reversal tested, no reversal learned) | EVIDENCE#032 → exp 43 (Q11) |
| Compact stochastic set generalizes to liquid single-stock names | REFUTED (out-of-universe RankIC −0.02, ICIR −0.07 — signal is noise on 30-name stock panel) | EVIDENCE#033 → exp 50 (Q14) |
| Assets are submartingales long-horizon / mean-reverting short-horizon (VR<1 at 5–20d) | PROVEN (clean-lake VR study: median VR 0.88–0.92 across 5–20d, 37–47% of ETFs significantly mean-reverting) | EVIDENCE#034 → Q19 VR study |
## Model
| Claim | Status | Evidence |
|-------|--------|----------|
| Seed count is load-bearing: 2 seeds < 5 seeds on clean data | PROVEN | EVIDENCE#016 → exp 28 |
| n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics | PROVEN | EVIDENCE#015 → exp 26 |
| Cost drag is the binding constraint, not signal quality | PROVEN (clean data) | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) |
| 10-seed ensemble raises rank metrics (RankIC 0.0671, L/S Sharpe 4.58) but book stays negative net (−0.93%) | PROVEN | EVIDENCE#023 → exp 34 (Q02) |
| More seeds raise signal breadth but do not cure the cost problem | PROVEN | EVIDENCE#023 → exp 34 (Q02) |
| 5-seed RankIC ensemble raises performance vs single model on ablated set | HYPOTHESIS (pre-clean-lake exp 12 idea; re-validated directionally by exp 22–24 but not as a clean A/B) | EVIDENCE#005 |
| Fractional-Kelly sizing beats equal-weight top-k net of costs | REFUTED (net +1.04% IR 0.11 < acceptance; mild improvement only) | EVIDENCE#027 → exp 38 (Q06) |
## Portfolio construction & risk
| Claim | Status | Evidence |
|-------|--------|----------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | HYPOTHESIS (idea: pre-clean-lake exp 13/14; not re-tested post-reset) | EVIDENCE#006/007 |
| $5M liquidity floor improves IR and cuts drawdown | REFUTED (post-reset A/B: floor binds but no IR edge — candidate 1.512 < baseline 1.580; DD cut is defunding) | EVIDENCE#029 → exp 40 (Q08) |
| Size/concentration caps hurt by cutting deployed capital | PROVEN (post-reset A/B: caps fold risk_degree ~0.0095, deploy ~$9.5k of $1M) | EVIDENCE#029 → exp 40 (Q08) |
| Risk-limit gates are a safety net, not an alpha lever | PROVEN | EVIDENCE#029 → exp 40 (Q08) |
| Weekly rebalance of the same signal is the campaign's best construction (net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag) | PROVEN | EVIDENCE#028 → exp 39 (Q07) |
| Turnover reduction (weekly) ≫ sizing (Kelly) ≫ gates (regime/risk-limit) as a performance lever | PROVEN | EVIDENCE#028 → exp 39 (Q07); EVIDENCE#027 → exp 38 (Q06); EVIDENCE#029/031 → exp 40/42 (Q08/Q10) |
| Widening the book (topk 10→20) adds no net edge (−1.88%) | PROVEN (refuted direction) | EVIDENCE#024 → exp 35 (Q03) |
| Fractional-Kelly sizing mildly improves but fails acceptance (net +1.04%, IR 0.11) | PROVEN (refuted direction) | EVIDENCE#027 → exp 38 (Q06) |
| Long-short top10/bottom10 has real pre-cost edge but daily L/S turnover destroys it ($96.7k cost ≈ 9.7% NAV, fill rate 0.40) | PROVEN (refuted direction) | EVIDENCE#030 → exp 41 (Q09) |
| HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5) meets only the drawdown leg; churns and erases gross | PROVEN (refuted direction) | EVIDENCE#031 → exp 42 (Q10) |
| Entry/risk gates (momentum, HMM) are byte-identical no-ops | PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) | EVIDENCE#031 → exp 42 (Q10) |
| Signal quality is the bottleneck, not the execution/risk layer | PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) | EVIDENCE#022–032 → exp 33–43 |
## Data & reproducibility
| Claim | Status | Evidence |
|-------|--------|----------|
| The pre-reset reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | PROVEN | EVIDENCE#010 → exp 21 |
| Old-lake data quality inflated the signal and backtest | PROVEN | EVIDENCE#010 → exp 21 |
| Signal work must be re-validated after any data rebuild | PROVEN (exp 21) / HYPOTHESIS (generality) | EVIDENCE#010 |
| Silent NaN-drop (feature-provider path mismatch, stale coverage, mid-experiment regeneration) is a first-order pipeline failure class | HYPOTHESIS (chat-documented failure modes; partially re-validated by exp 22 fix) | book/data/chat_mining/*.txt + EVIDENCE#010/011 |
| Pre-reset experiment baselines are not comparable to post-reset runs | PROVEN | EVIDENCE#009/010 (exp 20 R0 note, exp 21) |
## Live execution
| Claim | Status | Evidence |
|-------|--------|----------|
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled | PROVEN | EVIDENCE#020 → round 3 |
| Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 | PROVEN | EVIDENCE#020 → round 3 metrics |
| Execution claims trace to round_id + reconcile, not backtest | PROVEN (methodology, round 3 settled) | EVIDENCE#020 |
| 50-ETF panel results generalize to other universes | REFUTED (Q14: single-stock universe RankIC −0.02, ICIR −0.07 — signal is noise) | EVIDENCE#033 → exp 50 (Q14) |
| Effective independent names in the 50-ETF book is small (≈4) | PROVEN (clean-lake eigenvalue analysis: participation ratio 4.46, top-4 explain 66.8% var, 4 signal eigenvalues above Marchenko-Pastur bound) | EVIDENCE#035 → Q20 eigenanalysis |
## Open questions (settled by further experiments)
- exp 30 M2 Sharpe-drift: DONE — reproduced on the compact set by Q01 (exp 33), promoted to PROVEN.
- exp 15 Kelly sizing: re-run — DONE — refuted on the clean lake by Q06 (exp 38); mark the old hypothesis REFUTED.
- exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal — DONE — refuted as an IR lever by Q08 (exp 40); keep as safety net only.
- Weekly rebalance: reproduce on a second window / take to a live round.
- Out-of-universe validation: non-ETF universe for the compact stochastic feature set. — DONE — refuted by Q14 (exp 50); RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names.
- Long-horizon label (10d/22d) with a matching low-turnover construction (e.g. weekly recompute) — signal says the edge is there, cost says daily churn kills it; untested combination.
- Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel.
- Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.