book: add ch 11 walk-forward + 5 refuted guards (exp 52-56); EVIDENCE#043-047; rename 12→13-synthesis
This commit is contained in:
+17
-1
@@ -58,6 +58,20 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
|
||||
| Entry/risk gates (momentum, HMM) are byte-identical no-ops | PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) | EVIDENCE#031 → exp 42 (Q10) |
|
||||
| Signal quality is the bottleneck, not the execution/risk layer | PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) | EVIDENCE#022–032 → exp 33–43 |
|
||||
|
||||
## Walk-forward & guard candidates
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| The headline edges (weekly +12.51%, m2-sharpe22 +6.5%) are 2026-window-specific: walk-forward re-training makes 2024/2025 negative or flat for every config (A weekly −18.1%/−4.2%, B moments −16.1%/−8.9%, C ndrop2 −18.2%/−3.6%, m2 −26.4%/+0.4%) | PROVEN (refuted direction) | EVIDENCE#043/044 → exp 52/53 |
|
||||
| Configs sharing identical predictions are a single test of construction, not two tests of signal — A and C (byte-identical IC/RankIC) split +12.5% vs −1.4% in 2026 purely by strategy layer | PROVEN | EVIDENCE#043 → exp 52 |
|
||||
| A pre-deployment feature-drift / feature-PSI gate selects the profitable year | REFUTED (2026 has the highest feature drift yet the best result; CSRankNorm'd ranks are scale-invariant) | EVIDENCE#045 → exp 54 + exp 53 follow-up |
|
||||
| A label-regime PSI gate selects the profitable year | REFUTED (closest matches 2023/2021 lose −26.0%/−22.9%) | EVIDENCE#045 → exp 54 |
|
||||
| A streaming IC circuit breaker (`ic_min_rankic`) separates good years from bad | REFUTED (trips 25–50% of days every year, freezes rotation out of losers; do not deploy live) | EVIDENCE#048 → ic_gate.py trip-rate study |
|
||||
| Shorter training windows (1y/2y) recover the edge | REFUTED (every test year negative; only the growing window ever goes positive; mean annual excess ≈ −13% for every window length) | EVIDENCE#046 → exp 55 |
|
||||
| The edge concentrates in fresh (low-staleness) predictions | REFUTED (every 90-day staleness bucket negative; freshest bucket most negative; 2025 gains are late-year at 336–397d staleness) | EVIDENCE#047 → exp 56 |
|
||||
| The 2026 edge is a 2025–2026 regime artifact; no guard candidate recovers it out-of-sample | PROVEN | EVIDENCE#043–048 → exp 52–56 |
|
||||
| Live capital should be sized for the mean (≈ −13% annual excess), not the 2026 tail | PROVEN (walk-forward) + HYPOTHESIS (forward-looking) | EVIDENCE#043–047 → exp 52–56 |
|
||||
|
||||
## Data & reproducibility
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
@@ -90,4 +104,6 @@ The running scoreboard of every quantitative claim in the book. Updated per chap
|
||||
- Realized-moments on clean data: — DONE — confirmed by Q17 (exp 48); net +9.90% IR 0.99. Needs reproduction.
|
||||
- OptimalStopControl on clean data: — DONE — refuted by Q18 (exp 49); net −6.21% vs TopkDropout +2.13%.
|
||||
- Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel.
|
||||
- Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.
|
||||
- Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim.
|
||||
- Walk-forward re-validation of the campaign's headline results: DONE — exp 52/53/54 re-ran every headline config across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable; the edge is a 2025–2026 regime artifact. EVIDENCE#043–045.
|
||||
- Pre-deployment guard to isolate the profitable regime: DONE — all five candidates refuted (feature-PSI, label-regime PSI, streaming IC `ic_min_rankic`, adaptive short-window, window-staleness). No guard recovers the edge OOS. EVIDENCE#043–048.
|
||||
Reference in New Issue
Block a user