book: evidence boundary (clean-lake watermark) + ch02 cost reality + ch05 clean-lake reset — exp 21-31, chat mining
This commit is contained in:
+22
-17
@@ -1,50 +1,54 @@
|
||||
# CLAIMS.md — Proven vs Hypothesis Matrix
|
||||
|
||||
The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: `PROVEN` (reproduced from recorded run / reconciled round), `HYPOTHESIS` (plausible, tested once or never), `REFUTED` (tested and contradicted), `REFERENCED` (external citation).
|
||||
The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: `PROVEN` (reproduced from a recorded run **on the clean lake** / reconciled post-reset round), `HYPOTHESIS` (plausible, tested once or never, or pre-clean-lake / chat-derived idea), `REFUTED` (tested on clean data and contradicted), `REFERENCED` (external citation).
|
||||
|
||||
**Boundary rule:** only claims traceable to exp 21+ or post-reset live rounds may be `PROVEN`. Pre-clean-lake experiments (exp 8–18) and opencode chat transcripts are idea sources — their claims are `HYPOTHESIS` at best and are marked `(idea: pre-clean-lake)`.
|
||||
|
||||
## Signal & features
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC 0.040, ICIR 0.062) | PROVEN | EVIDENCE#001 → exp 8 |
|
||||
| Costs erase most of the baseline edge (+6.2% gross → +1.6% net) | PROVEN | EVIDENCE#002 → exp 8 |
|
||||
| Dropping model-specific feature families (ou, hmm) improves rank signal (RankIC 0.030→0.064) | PROVEN | EVIDENCE#003 → exp 9 |
|
||||
| Adding moment/volatility families regresses the signal | PROVEN (refuted direction) | EVIDENCE#004 → exp 11 |
|
||||
| General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 on clean data | PROVEN | EVIDENCE#012 → exp 23 |
|
||||
| Compact stochastic set is the clean-lake reference (RankIC 0.0663, RankICIR 0.2545) | PROVEN | EVIDENCE#013 → exp 24 |
|
||||
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | PROVEN (refuted direction) | EVIDENCE#014 → exp 25 |
|
||||
| Multi-horizon momentum (M1) degrades the reference | PROVEN (refuted direction) | EVIDENCE#017 → exp 29 |
|
||||
| GARCH(1,1) vol-regime features add no signal | PROVEN (refuted direction) | EVIDENCE#019 → exp 31 |
|
||||
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | HYPOTHESIS (one run, unreproduced) | EVIDENCE#018 → exp 30 |
|
||||
| More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3 supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 |
|
||||
| General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 | PROVEN | EVIDENCE#012 → exp 23 |
|
||||
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | HYPOTHESIS (one clean-lake run, unreproduced) | EVIDENCE#018 → exp 30 |
|
||||
| Dropping model-specific feature families (ou, hmm) improves the rank signal | HYPOTHESIS (idea: pre-clean-lake, exp 9) | EVIDENCE#003 → exp 9 |
|
||||
| Adding moment/volatility families regresses the signal | HYPOTHESIS (idea: pre-clean-lake, exp 11) | EVIDENCE#004 → exp 11 |
|
||||
| Baseline 1-day LGB signal is weak / costs erase most of the edge | HYPOTHESIS (idea: pre-clean-lake, exp 8) | EVIDENCE#001/002 → exp 8 |
|
||||
| More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3+ supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 |
|
||||
| Mean reversion (OU z-score, trend-slope reversal) is the stable single-feature edge | HYPOTHESIS (chat-derived clean-data study; see book/references/chat-ideas.md) | — |
|
||||
| Assets are submartingales long-horizon / mean-reverting short-horizon (VR<1 at 5–20d) | HYPOTHESIS (chat-derived martingale study, exp 19 never closed) | book/data/chat_mining/martingale-study.txt |
|
||||
|
||||
## Model
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| 5-seed RankIC ensemble raises performance vs single model on ablated set | PROVEN (pre-reset); re-validated post-reset exp 22–24 | EVIDENCE#005/011/013 |
|
||||
| Seed count is load-bearing: 2 seeds < 5 seeds on clean data | PROVEN | EVIDENCE#016 → exp 28 |
|
||||
| n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics | PROVEN | EVIDENCE#015 → exp 26 |
|
||||
| Cost drag is the binding constraint, not signal quality | PROVEN | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) |
|
||||
| Cost drag is the binding constraint, not signal quality | PROVEN (clean data) | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) |
|
||||
| 5-seed RankIC ensemble raises performance vs single model on ablated set | HYPOTHESIS (pre-clean-lake exp 12 idea; re-validated directionally by exp 22–24 but not as a clean A/B) | EVIDENCE#005 |
|
||||
| Fractional-Kelly sizing beats equal-weight top-k net of costs | HYPOTHESIS (exp 15 never finished) | run never completed |
|
||||
|
||||
## Portfolio construction & risk
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | PROVEN | EVIDENCE#006/007 → exp 13/14 |
|
||||
| Stop-control churns and bleeds costs (−11.3pp cost drag) | PROVEN | EVIDENCE#006 → exp 13 |
|
||||
| $5M liquidity floor improves net IR (0.81→0.98) and cuts drawdown (7.9%→5.4%) | PROVEN (pre-clean-lake; not comparable post-reset) | EVIDENCE#008 → exp 18 |
|
||||
| Size/concentration caps hurt by cutting deployed capital | PROVEN (pre-clean-lake) | EVIDENCE#008 → exp 18 |
|
||||
| Entry/risk gates (momentum, HMM) are byte-identical no-ops on the reference signal | PROVEN | EVIDENCE#009 → exp 20 |
|
||||
| Signal quality is the bottleneck, not the execution/risk layer | PROVEN (on the exp-20 reference) | EVIDENCE#009 → exp 20 |
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | HYPOTHESIS (idea: pre-clean-lake exp 13/14; not re-tested post-reset) | EVIDENCE#006/007 |
|
||||
| $5M liquidity floor improves IR and cuts drawdown | HYPOTHESIS (idea: pre-clean-lake exp 18; not comparable post-reset) | EVIDENCE#008 → exp 18 |
|
||||
| Size/concentration caps hurt by cutting deployed capital | HYPOTHESIS (idea: pre-clean-lake exp 18) | EVIDENCE#008 → exp 18 |
|
||||
| Entry/risk gates (momentum, HMM) are byte-identical no-ops | HYPOTHESIS (idea: pre-clean-lake exp 20) | EVIDENCE#009 → exp 20 |
|
||||
| Signal quality is the bottleneck, not the execution/risk layer | HYPOTHESIS (idea: pre-clean-lake exp 20; round-3 live is consistent but short) | EVIDENCE#009 → exp 20 |
|
||||
|
||||
## Data & reproducibility
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | PROVEN | EVIDENCE#010 → exp 21 |
|
||||
| The pre-reset reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | PROVEN | EVIDENCE#010 → exp 21 |
|
||||
| Old-lake data quality inflated the signal and backtest | PROVEN | EVIDENCE#010 → exp 21 |
|
||||
| Signal work must be re-validated after any data rebuild | PROVEN (exp 21) / HYPOTHESIS (generality) | EVIDENCE#010 |
|
||||
| Silent NaN-drop (feature-provider path mismatch, stale coverage, mid-experiment regeneration) is a first-order pipeline failure class | HYPOTHESIS (chat-documented failure modes; partially re-validated by exp 22 fix) | book/data/chat_mining/*.txt + EVIDENCE#010/011 |
|
||||
| Pre-reset experiment baselines are not comparable to post-reset runs | PROVEN | EVIDENCE#009/010 (exp 20 R0 note, exp 21) |
|
||||
|
||||
## Live execution
|
||||
@@ -55,6 +59,7 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
|
||||
| Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 | PROVEN | EVIDENCE#020 → round 3 metrics |
|
||||
| Execution claims trace to round_id + reconcile, not backtest | PROVEN (methodology, round 3 settled) | EVIDENCE#020 |
|
||||
| 50-ETF panel results generalize to other universes | HYPOTHESIS — TODO(evidence-needed) | — |
|
||||
| Effective independent names in the 50-ETF book is small (≈4) | HYPOTHESIS (chat-derived eigenvalue analysis, pre-reset) | book/data/chat_mining/exp-polluted-lake.txt |
|
||||
|
||||
## Open questions (settled by further experiments)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user