From 18b61bbb8fcc219b7c34bed517c9f30d9ad748bc Mon Sep 17 00:00:00 2001 From: zhaoli Date: Thu, 20 Aug 2026 21:03:59 +0000 Subject: [PATCH] =?UTF-8?q?book:=20add=20ch=2011=20walk-forward=20+=205=20?= =?UTF-8?q?refuted=20guards=20(exp=2052-56);=20EVIDENCE#043-047;=20rename?= =?UTF-8?q?=2012=E2=86=9213-synthesis?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- book/CLAIMS.md | 18 +++- book/EVIDENCE.md | 9 +- book/README.md | 27 ++++-- book/chapters/00-intro.md | 5 +- .../01-metrics-and-market-structure.md | 4 +- book/chapters/02-research-loop.md | 4 +- book/chapters/11-walk-forward-and-guards.md | 87 +++++++++++++++++++ .../{12-synthesis.md => 13-synthesis.md} | 9 +- 8 files changed, 144 insertions(+), 19 deletions(-) create mode 100644 book/chapters/11-walk-forward-and-guards.md rename book/chapters/{12-synthesis.md => 13-synthesis.md} (80%) diff --git a/book/CLAIMS.md b/book/CLAIMS.md index d3ee1dd..977a278 100644 --- a/book/CLAIMS.md +++ b/book/CLAIMS.md @@ -58,6 +58,20 @@ The running scoreboard of every quantitative claim in the book. Updated per chap | Entry/risk gates (momentum, HMM) are byte-identical no-ops | PROVEN (clean-lake re-test: regime gate refuted; still only DD relief) | EVIDENCE#031 → exp 42 (Q10) | | Signal quality is the bottleneck, not the execution/risk layer | PROVEN (10 of 11 Q-runs refuted on signal/construction; weekly cost relief wins) | EVIDENCE#022–032 → exp 33–43 | +## Walk-forward & guard candidates + +| Claim | Status | Evidence | +|-------|--------|----------| +| The headline edges (weekly +12.51%, m2-sharpe22 +6.5%) are 2026-window-specific: walk-forward re-training makes 2024/2025 negative or flat for every config (A weekly −18.1%/−4.2%, B moments −16.1%/−8.9%, C ndrop2 −18.2%/−3.6%, m2 −26.4%/+0.4%) | PROVEN (refuted direction) | EVIDENCE#043/044 → exp 52/53 | +| Configs sharing identical predictions are a single test of construction, not two tests of signal — A and C (byte-identical IC/RankIC) split +12.5% vs −1.4% in 2026 purely by strategy layer | PROVEN | EVIDENCE#043 → exp 52 | +| A pre-deployment feature-drift / feature-PSI gate selects the profitable year | REFUTED (2026 has the highest feature drift yet the best result; CSRankNorm'd ranks are scale-invariant) | EVIDENCE#045 → exp 54 + exp 53 follow-up | +| A label-regime PSI gate selects the profitable year | REFUTED (closest matches 2023/2021 lose −26.0%/−22.9%) | EVIDENCE#045 → exp 54 | +| A streaming IC circuit breaker (`ic_min_rankic`) separates good years from bad | REFUTED (trips 25–50% of days every year, freezes rotation out of losers; do not deploy live) | EVIDENCE#048 → ic_gate.py trip-rate study | +| Shorter training windows (1y/2y) recover the edge | REFUTED (every test year negative; only the growing window ever goes positive; mean annual excess ≈ −13% for every window length) | EVIDENCE#046 → exp 55 | +| The edge concentrates in fresh (low-staleness) predictions | REFUTED (every 90-day staleness bucket negative; freshest bucket most negative; 2025 gains are late-year at 336–397d staleness) | EVIDENCE#047 → exp 56 | +| The 2026 edge is a 2025–2026 regime artifact; no guard candidate recovers it out-of-sample | PROVEN | EVIDENCE#043–048 → exp 52–56 | +| Live capital should be sized for the mean (≈ −13% annual excess), not the 2026 tail | PROVEN (walk-forward) + HYPOTHESIS (forward-looking) | EVIDENCE#043–047 → exp 52–56 | + ## Data & reproducibility | Claim | Status | Evidence | @@ -90,4 +104,6 @@ The running scoreboard of every quantitative claim in the book. Updated per chap - Realized-moments on clean data: — DONE — confirmed by Q17 (exp 48); net +9.90% IR 0.99. Needs reproduction. - OptimalStopControl on clean data: — DONE — refuted by Q18 (exp 49); net −6.21% vs TopkDropout +2.13%. - Martingale / variance-ratio study: DONE — PROVEN by Q19 scripted study; VR < 1 at 5–20d with significant z-stats for 37–47% of the panel. -- Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim. \ No newline at end of file +- Effective independent names: DONE — PROVEN by Q20 eigenvalue analysis; participation ratio ≈ 4.5, matching the chat-derived claim. +- Walk-forward re-validation of the campaign's headline results: DONE — exp 52/53/54 re-ran every headline config across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable; the edge is a 2025–2026 regime artifact. EVIDENCE#043–045. +- Pre-deployment guard to isolate the profitable regime: DONE — all five candidates refuted (feature-PSI, label-regime PSI, streaming IC `ic_min_rankic`, adaptive short-window, window-staleness). No guard recovers the edge OOS. EVIDENCE#043–048. \ No newline at end of file diff --git a/book/EVIDENCE.md b/book/EVIDENCE.md index 855b4cd..2514b1f 100644 --- a/book/EVIDENCE.md +++ b/book/EVIDENCE.md @@ -24,7 +24,7 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit | EVIDENCE#008 | Risk-limit A/B: $5M liquidity floor → net IR 0.81→0.98, cumDD 7.93%→5.44%; size cap 15% + conc 60% hurts (IR 0.816, ann 6.11%). | exp 18, run `28c7fa08…` (mlflow exp 21), branch `exp/18-risk-limit-control-on-the-reference-ense` | NOT usable as PROVEN — pre-clean-lake (idea: liquidity floor > concentration caps) | | EVIDENCE#009 | Improvement sweep (R1-R5): 4/5 refuted; R2 momentum gate and R3 HMM gate are byte-identical no-ops; R5 MA3/EWMA marginal (IR 0.049). Conclusion: signal quality is the bottleneck, not the execution/risk layer. | exp 20, run `958198a8…` (mlflow exp 21), branch `exp/20-improve-the-risk-limit-reference-signal` | NOT usable as PROVEN — pre-clean-lake (idea: gates are no-ops when signal is weak) | -## Post-reset period (exp 21–43) — canonical, current +## Post-reset period (exp 21–56) — canonical, current | ID | Claim | Source | Verified? | |----|-------|--------|-----------| @@ -59,6 +59,11 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit | EVIDENCE#040 | Q17 realized-moments features (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) on clean data: improves portfolio over baseline. Net +9.90% (IR 0.990), gross +14.61%, maxDD −6.49% vs baseline net +2.13% (IR 0.21). IC 0.039 (vs 0.050), RankIC 0.060 (vs 0.066) — signal metrics slightly lower but portfolio construction benefits from moment conditioning. Contradicts pre-reset EVIDENCE#004 (which was inflated by dirty data). Single run, unreproduced. | exp 48, run `e62ce326…` (mlflow exp 48), branch `exp/48-q17-moments` | yes — Q17 HYPOTHESIS (needs reproduction) | | EVIDENCE#041 | Q18 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) vs TopkDropout on clean data: same signal (IC 0.050, RankIC 0.066 — identical model), worse portfolio. Net −6.21% (IR −0.640) vs baseline +2.13% (IR 0.21). Cost drag ~8.3pp. Clean-lake confirmation of pre-reset EVIDENCE#006/#007. | exp 49, run `f140dcb8…` (mlflow exp 49), branch `exp/49-q18-optstop` | yes — Q18 FAIL (confirms EVIDENCE#006/#007 on clean data) | | EVIDENCE#042 | Q21 10d label + weekly rebalance: cost drag cut from 4.61pp (Q04 daily) to 1.05pp (weekly). Net flipped from −9.92% to +1.19% (IR 0.148, maxDD −4.78%). Signal identical to Q04 (IC 0.093, RankIC 0.096). Weekly rebalance delivers ~10pp improvement regardless of label horizon (5d: +10.38pp via Q07, 10d: +11.11pp via Q21). But IR 0.148 < 0.5 acceptance — 5d+weekly (Q07, IR 1.24) remains the best construction. | exp 51, run `046c93a6…` (mlflow exp 51), branch `exp/51-q21-test-10d-label--weekly-rebalance-q04` | yes — Q21 FAIL (below IR bar, but confirms weekly-rebalance universality) | +| EVIDENCE#043 | Walk-forward 3×3 (3 best configs × 2024/2025/2026): A weekly n_drop1 = −18.1% (IR −1.39) / −4.2% (IR −0.52) / **+12.5% (IR 1.25)**; B moments n_drop1 = −16.1% (IR −1.91) / −8.9% (IR −1.00) / **+9.2% (IR 0.94)**; C base n_drop2 = −18.2% (IR −1.97) / −3.6% (IR −0.51) / **−1.4% (IR −0.13)**. Only 2026 is profitable, and only for A/B. A and C share identical predictions (byte-identical IC/RankIC) — the strategy layer alone decides the outcome. Run A-2025 exactly replicated exp 45 (`e5ac7a5d`). The edge is a 2026-window-specific regime artifact. | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` | yes — walk-forward REFUTED (edge window-specific) | +| EVIDENCE#044 | m2-sharpe22 3-window: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the exp-33 reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — harness is reproducible; edge is recent-window-only. | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference `c7c12228` (exp 33) | yes — walk-forward REFUTED (edge recent-window-only) | +| EVIDENCE#045 | Label-regime transfer (2021/2023 — the closest label-regime PSI matches to 2026): 2023 −26.0% (IR −2.04, maxDD −30.8%), 2021 −22.9% (IR −2.26, maxDD −27.0%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two closest matches both lose ≈ a quarter. Feature-PSI gate also fails: 2026 has the highest feature drift yet the best result (CSRankNorm'd ranks are scale-invariant). No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) | yes — guard candidates 1+2 REFUTED | +| EVIDENCE#046 | Adaptive short-window retrain (1y/2y rolling windows): 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing 2016→prev-Aug window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length. | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` | yes — guard candidate 4 REFUTED | +| EVIDENCE#047 | Window-staleness isolation: pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in EVERY bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%) — the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug, Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). No staleness threshold isolates the edge. Account-based cumulative excess vs SPY: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38% (blotter `return` field excludes initial cost — use `account`). | exp 56, staleness analysis on exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` | yes — guard candidate 5 REFUTED | ## Live execution trail @@ -71,7 +76,7 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit | ID | Claim | Source | Verified? | |----|-------|--------|-----------| -| (none yet) | — | — | — | +| EVIDENCE#048 | Streaming IC circuit-breaker (`ic_min_rankic`, `ICGateTopkDropoutStrategy` in `tac_qlib/contrib/strategy/ic_gate.py`) trip-rate study: with thresholds 0.02–0.06, the gate trips on 25–50% of days in every year (2021–2026), freezing TopkDropout's rotation out of losers. A gate that trips every year cannot separate good years from bad. Do not deploy live. | ad-hoc scripted study on exp 52/53 pred/label artifacts, `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py`, `tac_qlib/tac_qlib/risk_limits.py` | yes — guard candidate 3 REFUTED | ## External references (book/references/) diff --git a/book/README.md b/book/README.md index bb52a49..49362fe 100644 --- a/book/README.md +++ b/book/README.md @@ -33,8 +33,9 @@ A quant-desk reader should be able to act on this book: replicate a signal pipel | 08 | Portfolio construction: dropout vs the rest | drafting | exp 13, 14, 15, 35, 38, 39, 41 | Turnover-sensitive construction bleeds the edge; weekly recompute wins | | 09 | The cost/turnover frontier | drafting | exp 26, 39, 41 | Cut turnover before adding signal; weekly rebalance is the proven lever | | 10 | Risk limits and gates that work | drafting | exp 18, 20, 40, 42 | Limits are a safety net, not alpha; gates churn without signal | -| 11 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 | -| 12 | Synthesis: how proved truth compounds | drafting | all, exp 33–43 | Cost relief > signal; the Q-campaign scoreboard | +| 11 | Walk-forward re-validation and guard candidates | drafting | exp 52–56 | The edge is a 2025–2026 regime artifact; all 5 guards refuted | +| 12 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 | +| 13 | Synthesis: how proved truth compounds | drafting | all, exp 33–43, 52–56 | Cost relief > signal; the Q-campaign scoreboard | Status legend: `drafting` → `in-review` → `done`. @@ -52,7 +53,7 @@ Each chapter opens with its claims. The inventory below is the working contract: ### 01 — Metrics: the vocabulary of a price series | Claim | Expected status | |-------|-----------------| -| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–12) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) | +| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–13) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) | | Generic scale-free statistics beat model-specific machinery on a small daily panel | `PROVEN` (exp 23/24/25/29/31) + `HYPOTHESIS` (generality) | | A statistic is only as good as the falsification it survives (null z-scores, reproduction) | `PROVEN` (exp 21 detection playbook) + `REFERENCED` | | The strongest single-feature signal (OU z-score) can be worthless inside a rank model — the "OU paradox" | `PROVEN` (exp 25) + open mechanism `TODO(evidence-needed)` | @@ -133,19 +134,28 @@ Each chapter opens with its claims. The inventory below is the working contract: | Post-reset A/B: the floor binds but adds no IR edge; DD relief is pure defunding | `PROVEN` — exp 40 (Q08) | | HMM regime gate meets only the drawdown leg and churns | `PROVEN` — exp 42 (Q10) | -### 11 — Live execution and reconciliation +### 11 — Walk-forward re-validation and guard candidates +| Claim | Expected status | +|-------|-----------------| +| The headline results (weekly +12.51%, m2-sharpe22 +6.5%) are 2026-window-specific; walk-forward re-training across 2024/2025 is negative or flat | `PROVEN` — exp 52/53 | +| A and C share identical predictions; the strategy layer alone decides the outcome | `PROVEN` — exp 52 | +| No pre-deployment measurable gate (feature-PSI, label-regime PSI, streaming IC, window length, staleness) selects a profitable year | `PROVEN` — exp 52–56, all 5 guards refuted | +| The edge is a 2025–2026 regime artifact; live capital must be cut until the regime returns | `PROVEN` (walk-forward) + `HYPOTHESIS` (forward-looking) | + +### 12 — Live execution and reconciliation | Claim | Expected status | |-------|-----------------| | Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 | | Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics | | Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology | -### 12 — Synthesis +### 13 — Synthesis | Claim | Expected status | |-------|-----------------| | The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26, 39 | | The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31, Q02–Q06, Q09–Q11) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) | | Turnover reduction is the dominant net-performance lever (weekly rebalance +12.51% vs daily −3.21%–+2.13%) | `PROVEN` — exp 26 vs 39 | +| The campaign's headline edges were a 2025–2026 regime artifact, not robust OOS | `PROVEN` — exp 52–56 | | Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) | ## Repository layout @@ -167,13 +177,16 @@ book/ - `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)` - `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)` - `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)` -- `TODO(evidence-needed: purged walk-forward CV instead of single train/valid split on the exp-26 reference)` - `TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)` - `TODO(evidence-needed: live round under weekly-rebalance construction with risk-limit spec, to confirm safety-net behavior at higher deployed capital)` +- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)` +- `TODO(evidence-needed: a causal (no-lookahead) regime-change detector that selects the 2026 window before the fact — none of the five guards did)` ### Settled open questions (no longer active) - ~~`weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`. - ~~`long-horizon label (10d/22d) paired with a low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`. - ~~`out-of-universe (non-ETF) validation of the compact stochastic feature set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`. -- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`. \ No newline at end of file +- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`. +- ~~`do the headline results survive walk-forward re-training?`~~ — **ANSWERED (negatively):** exp 52/53/54 re-ran weekly, moments, ndrop2, and m2-sharpe22 across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable; all prior years negative or flat. Edge = 2025–2026 regime artifact. `EVIDENCE#043–045`. +- ~~`is there a pre-deployment guard that isolates the profitable regime?`~~ — **ANSWERED (negatively):** feature-PSI, label-regime PSI, streaming IC (`ic_min_rankic`), adaptive short-window, and staleness guards all refuted. `EVIDENCE#043–047`. \ No newline at end of file diff --git a/book/chapters/00-intro.md b/book/chapters/00-intro.md index 1c68a2a..2906e71 100644 --- a/book/chapters/00-intro.md +++ b/book/chapters/00-intro.md @@ -40,8 +40,9 @@ The most important proof of this discipline is the clean-lake reset, which this - Every claim is tagged `PROVEN` (traced experiment/round), `HYPOTHESIS` (unreproduced), or `REFERENCED` (external source). `EVIDENCE.md` maps each tag to the run, branch, and round behind it. - Chapters 03–10 follow the research arc: what was tested, what was proved, what was refuted, and what moved performance. Refuted runs are cited as evidence too — knowing what *doesn't* work is how the desk avoided paying for it twice. -- Chapter 11 is the reality check: live execution against the research claims. -- Chapter 12 is the synthesis: the scoreboard of what actually improved performance and why. +- Chapter 11 is the walk-forward reality check: do the headline results survive re-training out-of-window, and can any guard isolate the profitable regime? +- Chapter 12 is the reality check on execution: live results against the research claims. +- Chapter 13 is the synthesis: the scoreboard of what actually improved performance and why. ## Open questions diff --git a/book/chapters/01-metrics-and-market-structure.md b/book/chapters/01-metrics-and-market-structure.md index 073e5a0..37f3ab4 100644 --- a/book/chapters/01-metrics-and-market-structure.md +++ b/book/chapters/01-metrics-and-market-structure.md @@ -78,7 +78,7 @@ Two facts about risk matter throughout the book. First, the measurement floor: o Error is the book's discipline: how wrong was the prediction, in a way that can be measured against a null? The canonical metrics on the clean lake are IC, ICIR, Rank IC, Rank ICIR (the rank-based signal quality) and net IR, net return, L/S Sharpe, max drawdown (the portfolio outcome). Pre-reset runs (exp 8–18) recorded a different schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ir_with_cost`) — the two schemas are never compared directly in this book `(EVIDENCE.md: metric-schema note)`. -Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); live results are reconciled with slippage and cost, not taken from the backtest (ch. 11). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26, round 3)`. +Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); a single-window backtest is re-validated walk-forward before shipping (ch. 11); live results are reconciled with slippage and cost, not taken from the backtest (ch. 12). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → walk-forward (does it survive re-training?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26/52–56, round 3)`. ## Probability and significance @@ -101,7 +101,7 @@ The cycle this book runs on: **measure → hypothesize → pre-register → isol 3. **Pre-register** — the claim and its acceptance metric (IC/RankIC above reference, net IR above reference) are fixed before execution, to block post-hoc cherry-picking across the 31+ experiments. 4. **Isolate** — one variable changes per run; the reference book and its metrics are the control (exp 26 → 29/30/31). 5. **Prove or refute** — on the clean lake only. A reproduced improvement becomes PROVEN; a single un-reproduced run stays HYPOTHESIS (exp 30); a degradation is REFUTED and — critically — is recorded as a win for the discipline (exp 29, exp 31 stopped wrong directions). -6. **Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 11). +6. **Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 12). Every metric in this chapter sits on this loop. The drift metric produced the momentum hypothesis and the reversal hypothesis; only one survived isolation. The error metrics are the loop's judge. The decay and risk metrics are why ch. 03 and ch. 09 exist at all. The rest of the book is the working-out of this cycle, claim by claim, with each claim traceable to `EVIDENCE.md` and a recorded run. diff --git a/book/chapters/02-research-loop.md b/book/chapters/02-research-loop.md index ce82fee..c61ca8d 100644 --- a/book/chapters/02-research-loop.md +++ b/book/chapters/02-research-loop.md @@ -8,7 +8,7 @@ The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. T 1. **Lake** — bars and features live in one hive-partitioned lake (market/timeframe/symbol, `family=ta|sp`), with coverage and calendar metadata. It is the single source of bar/feature truth. The clean-lake rebuild proved the stakes: when the lake was rebuilt, the reference signal collapsed (IC 0.0354 → 0.0019) because the old lake's data quality had silently inflated results `(PROVEN → exp 21)`. A claim built on the lake is only as good as the lake. 2. **Experiment** — every run is a traced experiment: a git branch (`exp/N-…`), an MLflow run with recorded config/params/metrics, and hypothesis/evaluation notes recorded before and after the run. The traceability loop was used on exp 8–31; the branch, run, and notes are the reproducible unit `(PROVEN → the traced experiment store; see `rd_exp_*` tools and `EVIDENCE.md`)`. -3. **Live** — a proved book advances to a round window (targets → intents → decisions → orders → fills), and is reconciled (slippage bps, cost, funnel) `(PROVEN → round 3; tac-rd-book trail)`. Live beats backtest: a claim about trading performance must trace to a round, not to a backtest (ch. 11). +3. **Live** — a proved book advances to a round window (targets → intents → decisions → orders → fills), and is reconciled (slippage bps, cost, funnel) `(PROVEN → round 3; tac-rd-book trail)`. Live beats backtest: a claim about trading performance must trace to a round, not to a backtest (ch. 12). ## Why traceability is the methodology @@ -29,6 +29,6 @@ Most additions failed. The loop's value is not that it produced winners — it i ## From here -Ch. 03 applies the loop to the book's first worked question (does the signal clear costs?), ch. 06 to the clean-lake reset, and ch. 11 to the live round that closes the loop with reconciliation. +Ch. 03 applies the loop to the book's first worked question (does the signal clear costs?), ch. 06 to the clean-lake reset, ch. 11 to walk-forward re-validation of the campaign's headline results, and ch. 12 to the live round that closes the loop with reconciliation. Open questions: purge/walk-forward CV instead of single train/valid split `TODO(evidence-needed: purged CV on the exp-26 reference)`, and a PSI-based drift-aware retraining gate `(HYPOTHESIS → chat-ideas.md)`. \ No newline at end of file diff --git a/book/chapters/11-walk-forward-and-guards.md b/book/chapters/11-walk-forward-and-guards.md new file mode 100644 index 0000000..7b8bf70 --- /dev/null +++ b/book/chapters/11-walk-forward-and-guards.md @@ -0,0 +1,87 @@ +# Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact + +Status: drafting. Claim inventory: see `README.md` ch. 11. + +This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it. + +## The walk-forward re-validation (exp 52, 53, 54) + +Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (`42,7,2026,99,123`, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel): + +### 3×3 sweep (exp 52) — the three best configs across 2024/2025/2026 + +Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost. + +| Config | 2024 (test) | 2025 (test) | 2026 (test) | +|--------|-------------|-------------|-------------| +| **A** weekly, n_drop=1 | −18.1% (IR −1.39, maxDD −22.7%) | −4.2% (IR −0.52, maxDD −10.3%) | **+12.5% (IR 1.25, maxDD −4.1%)** | +| **B** moments, n_drop=1 | −16.1% (IR −1.91, maxDD −20.0%) | −8.9% (IR −1.00, maxDD −9.8%) | **+9.2% (IR 0.94, maxDD −6.5%)** | +| **C** base, n_drop=2 | −18.2% (IR −1.97, maxDD −23.3%) | −3.6% (IR −0.51, maxDD −6.3%) | **−1.4% (IR −0.13, maxDD −9.8%)** | + +`PROVEN — EVIDENCE#043 → exp 52`. Read the table carefully: + +- **The 2026 window is the only profitable one**, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026. +- **A and C share identical predictions** — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026. +- **The harness is reproducible**: run 2 (Config A, 2025) exactly replicated exp 45 (`e5ac7a5d`, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (`eb38588c`, +12.51%, IR 1.24). + +### m2-sharpe22 3-window (exp 53) + +The m2-sharpe22 reference (exp 33/Q01, `c7c12228`) re-run on the same 2024/2025/2026 scheme: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. `PROVEN — EVIDENCE#044 → exp 53`. The edge is recent-window-only. + +### Label-regime transfer (exp 54) — 2021 and 2023 + +The 2025 feature-drift check had already shown a naive feature-PSI gate does **not** predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the *label/return regime*: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run: + +- **2023**: −26.0% (IR −2.04, maxDD −30.8%) +- **2021**: −22.9% (IR −2.26, maxDD −27.0%) + +`PROVEN — EVIDENCE#045 → exp 54`. Both closest label-regime matches are as bad as the 2024 tail. **No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year.** 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups. + +## The five guard candidates — all refuted + +With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed: + +| # | Guard | Test | Result | +|---|-------|------|--------| +| 1 | **Feature-PSI gate** | halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) | REFUTED — 2026 has the *highest* drift yet the *best* result; CSRankNorm'd ranks make raw drift scale-invariant. | +| 2 | **Label-regime gate** | trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) | REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up. | +| 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. | +| 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. | +| 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. | + +Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`. + +## The account-level truth + +The blotter's daily `account` field is the authoritative measure (the `return` field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: **2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%**. This reconciles with the recorded metrics — 2026 `excess_return_with_cost` annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. `PROVEN — EVIDENCE#047 → exp 56` (account curves from the exp 53/54 runs' blotter artifacts). + +## The synthesis + +- **The headline results were window-specific.** Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. `PROVEN — EVIDENCE#043/044`. +- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. `PROVEN — EVIDENCE#043–047`. +- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024. +- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe. + +## Desk rules distilled from this chapter + +1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy. +2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison. +3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel. +4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account. +5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back. + +## Open questions + +- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)` +- `TODO(evidence-needed: a regime-change detector that is causal (no lookahead) and demonstrably selects the 2026 window before the fact — none of the five guards did)` + +## Evidence cited in this chapter + +| Tag | Source | +|-----|--------| +| `EVIDENCE#043` | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` | +| `EVIDENCE#044` | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference run `c7c12228` (exp 33, Q01) | +| `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) | +| `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` | +| `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` | +| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds | \ No newline at end of file diff --git a/book/chapters/12-synthesis.md b/book/chapters/13-synthesis.md similarity index 80% rename from book/chapters/12-synthesis.md rename to book/chapters/13-synthesis.md index 1e78421..ae8ed0c 100644 --- a/book/chapters/12-synthesis.md +++ b/book/chapters/13-synthesis.md @@ -1,6 +1,6 @@ -# Chapter 12 — Synthesis: How Proved Truth Compounds +# Chapter 13 — Synthesis: How Proved Truth Compounds -Status: drafting. Claim inventory: see `README.md` ch. 12. +Status: drafting. Claim inventory: see `README.md` ch. 13. This chapter is the scoreboard. It collects everything the campaign proved, in order of what actually moved performance — and why the Q-campaign's 9 refutations were as informative as its 2 passes. @@ -36,7 +36,8 @@ The refuted runs were as valuable as the passes: the label-horizon result (22d l 3. **Test features one at a time** against the reference (ch. 07); prune, don't add (ch. 04). 4. **Use a small ensemble** (5 seeds) and stop there (ch. 05). 5. **A/B risk limits** before shipping; keep them as a safety net (ch. 10). -6. **Reconcile live** — the funnel and slippage are the only claims that count (ch. 00/11). +6. **Reconcile live** — the funnel and slippage are the only claims that count (ch. 00/12). +7. **Re-validate walk-forward before shipping** — a single-window edge is a regime artifact until it survives re-training out-of-window (ch. 11). ## Open questions @@ -49,6 +50,8 @@ The refuted runs were as valuable as the passes: the label-horizon result (22d l - ~~`reproduce exp 39 weekly rebalance on a second window, then a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`. - ~~`long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`. - ~~`out-of-universe (non-ETF) validation of the compact stochastic set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`. +- ~~`do the campaign's headline results (Q01 m2-sharpe22, Q07 weekly) survive walk-forward re-training?`~~ — **ANSWERED (negatively):** exp 52/53/54 re-ran every headline config across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable (+6.5% m2, +12.5% weekly); all prior years are negative or flat. **The edge is a 2025–2026 regime artifact.** `EVIDENCE#043–045`. +- ~~`is there a pre-deployment guard that isolates the profitable regime?`~~ — **ANSWERED (negatively):** five guard candidates (feature-PSI, label-regime PSI, streaming IC `ic_min_rankic`, adaptive short-window, window-staleness) were pre-registered and all refuted; none selects the good years in hindsight. `EVIDENCE#043–047`; see ch. 11. ## Evidence cited in this chapter