diff --git a/book/chapters/11-walk-forward-and-guards.md b/book/chapters/11-walk-forward-and-guards.md index b4141f6..bf1896b 100644 --- a/book/chapters/11-walk-forward-and-guards.md +++ b/book/chapters/11-walk-forward-and-guards.md @@ -2,7 +2,7 @@ Status: drafting. Claim inventory: see `README.md` ch. 11. -This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it. +This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests seven guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it. ## The walk-forward re-validation (exp 52, 53, 54) @@ -37,9 +37,9 @@ The 2025 feature-drift check had already shown a naive feature-PSI gate does **n `PROVEN — EVIDENCE#045 → exp 54`. Both closest label-regime matches are as bad as the 2024 tail. **No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year.** 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups. -## The five guard candidates — all refuted +## The seven guard candidates — all refuted -With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed: +With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested seven guards that could plausibly have preserved the good years and cut the bad ones. All seven were pre-registered as hypotheses (traced experiments), all seven failed: | # | Guard | Test | Result | |---|-------|------|--------| @@ -48,9 +48,10 @@ With the walk-forward sweep showing the edge is 2026-window-specific, the desk t | 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. | | 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. | | 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. | -| 6 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. The gate destroys the strategy's ability to capture the good days that compensate for the bad ones. | +| 6 | **Regime gate** (dispersion / vol / HMM) | daily boolean gate based on market state (low vol, HMM posterior, dispersion) | REFUTED — see below; regime gate closes on the wrong days (hmm_0.7 destroys 2026: +25.5% → +10.8%) | +| 7 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. | -Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`. +Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `PROVEN (refuted) — EVIDENCE#050`; guard 7 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`. ## The account-level truth @@ -83,7 +84,7 @@ Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). ## Model search & robustness of the regime gate finding -The regime gate study (Guard 6, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection: +The regime gate study (Guard 7, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection: | Rank | Exp | Test Window | RankICIR | Net Return | Notes | |------|-----|-------------|----------|------------|-------| @@ -182,7 +183,7 @@ The vol gates show the largest trip differential — they open on more days in 2 `hmm_0.7` has the most interesting profile: it **improves** 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but **destroys** 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods. -**Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**. See Guard 7 (signal-quality gate, EVIDENCE#052) for a gate that works. +**Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**. See Guard 7 (signal-quality gate, EVIDENCE#053) for a gate that was tested on this principle — and still failed. ## Open questions