# Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact Status: drafting. Claim inventory: see `README.md` ch. 11. This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it. ## The walk-forward re-validation (exp 52, 53, 54) Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (`42,7,2026,99,123`, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel): ### 3×3 sweep (exp 52) — the three best configs across 2024/2025/2026 Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost. | Config | 2024 (test) | 2025 (test) | 2026 (test) | |--------|-------------|-------------|-------------| | **A** weekly, n_drop=1 | −18.1% (IR −1.39, maxDD −22.7%) | −4.2% (IR −0.52, maxDD −10.3%) | **+12.5% (IR 1.25, maxDD −4.1%)** | | **B** moments, n_drop=1 | −16.1% (IR −1.91, maxDD −20.0%) | −8.9% (IR −1.00, maxDD −9.8%) | **+9.2% (IR 0.94, maxDD −6.5%)** | | **C** base, n_drop=2 | −18.2% (IR −1.97, maxDD −23.3%) | −3.6% (IR −0.51, maxDD −6.3%) | **−1.4% (IR −0.13, maxDD −9.8%)** | `PROVEN — EVIDENCE#043 → exp 52`. Read the table carefully: - **The 2026 window is the only profitable one**, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026. - **A and C share identical predictions** — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026. - **The harness is reproducible**: run 2 (Config A, 2025) exactly replicated exp 45 (`e5ac7a5d`, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (`eb38588c`, +12.51%, IR 1.24). ### m2-sharpe22 3-window (exp 53) The m2-sharpe22 reference (exp 33/Q01, `c7c12228`) re-run on the same 2024/2025/2026 scheme: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. `PROVEN — EVIDENCE#044 → exp 53`. The edge is recent-window-only. ### Label-regime transfer (exp 54) — 2021 and 2023 The 2025 feature-drift check had already shown a naive feature-PSI gate does **not** predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the *label/return regime*: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run: - **2023**: −26.0% (IR −2.04, maxDD −30.8%) - **2021**: −22.9% (IR −2.26, maxDD −27.0%) `PROVEN — EVIDENCE#045 → exp 54`. Both closest label-regime matches are as bad as the 2024 tail. **No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year.** 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups. ## The five guard candidates — all refuted With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed: | # | Guard | Test | Result | |---|-------|------|--------| | 1 | **Feature-PSI gate** | halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) | REFUTED — 2026 has the *highest* drift yet the *best* result; CSRankNorm'd ranks make raw drift scale-invariant. | | 2 | **Label-regime gate** | trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) | REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up. | | 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. | | 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. | | 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. | Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`. ## The account-level truth The blotter's daily `account` field is the authoritative measure (the `return` field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: **2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%**. This reconciles with the recorded metrics — 2026 `excess_return_with_cost` annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. `PROVEN — EVIDENCE#047 → exp 56` (account curves from the exp 53/54 runs' blotter artifacts). ## The synthesis - **The headline results were window-specific.** Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. `PROVEN — EVIDENCE#043/044`. - **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. `PROVEN — EVIDENCE#043–047`. - **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024. - **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe. ## Within-window robustness (perturbation stress test) The 2026 edge is fragile *across* windows but robust *within* the 2026 window. A perturbation grid on Config A's predictions (exp 52, run `9f98ea5c`, same pred.pkl, varying only backtest parameters): | Perturbation | Config | Ann. return | Sharpe | maxDD | |-------------|--------|-------------|--------|-------| | **Baseline** | topk=10, n_drop=1, costs 5/15/$5 | 32.8% | 1.98 | −5.8% | | topk=5 | concentration ↑ | 32.3% | 1.75 | −7.0% | | topk=15 | concentration ↓ | 26.3% | 1.61 | −6.9% | | n_drop=2 | rotation ↑ | 26.7% | 1.59 | −6.5% | | n_drop=3 | rotation ↑↑ | 28.6% | 1.66 | −6.7% | | costs 3× (15/25/$10) | cost stress | 32.8% | 1.97 | −5.8% | | costs 5× (25/35/$15) | cost stress ↑↑ | 32.7% | 1.97 | −5.8% | `PROVEN — EVIDENCE#049` (ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json`). Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. **Within the one good window, the edge is not a parameter-tuning artifact.** The fragility is entirely across windows (regime dependence), not within them. ## Model search & robustness of the regime gate finding The regime gate study (Guard 6, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection: | Rank | Exp | Test Window | RankICIR | Net Return | Notes | |------|-----|-------------|----------|------------|-------| | 1 | 36 | 2026 only | **0.507** | −4.6% | 22d label, single-window | | 2 | 44 | 2026 only | **0.507** | −4.9% | Same pred as #1 | | 3 | 35 | 2026 only | **0.352** | −9.9% | 10d label | | 4 | 51 | 2026 only | **0.352** | +1.2% | Same pred as #3 | | 5 | 58 | 2025 only | **0.289** | −3.4% | Adaptive 2y | | 6 | 11 | 2026 only | **0.276** | +3.1% | Single seed | | 7 | 33 | 2026 only | **0.259** | −0.9% | 10 seeds | | 8 | **52-C** | **2024–2026** | **0.244** | **+12.5%** | Walk-forward, weekly | | 9 | **52-A** | **2024–2026** | **0.244** | **−1.4%** | Walk-forward, TopkDrop | `PROVEN — EVIDENCE#051` (comprehensive `rd_exp_list` query, run metadata). Key observations: 1. **The 22-day label models (exp 36/44) have the highest RankICIR (0.507) but negative returns** — high IC does not guarantee profitable trading. The 22d label predicts longer-horizon moves that don't translate to short-term alpha after costs. 2. **Most high-RankICIR models are single-window (2026 only)** — they lack the multi-year coverage needed for the regime gate study. Walk-forward coverage (2021–2026) is limited to exp 52 (2024–2026) and exp 56 (2021, 2023). 3. **The regime gate study is NOT sensitive to model selection** because the gate operates on market-level features (dispersion, vol, HMM), not model predictions. Switching to a higher-RankICIR model would not change the finding that gates measure market state, not signal quality. 4. **Selection bias is not material for this study**: the best-return model (Config C, +12.5%) also has the best RankICIR (0.244) among walk-forward configs. The RankICIR and returns rankings are concordant. ## Desk rules distilled from this chapter 1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy. 2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison. 3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel. 4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account. 5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back. ### Guard 6: Regime gate (dispersion / vol / HMM) `PROVEN — EVIDENCE#050` If the edge is regime-dependent, the most direct guard is a regime detector that opens on good years and closes on bad years. We test three detector types, each producing a daily boolean (trade / don't trade): | Detector | Logic | |----------|-------| | **dispersion** | CS std of 22-day rolling returns < threshold (low dispersion → calm market → trade) | | **vol** | CS mean of 22-day rolling realized vol within a band (mid-range vol → trade) | | **HMM** | 2-state Gaussian HMM posterior for regime 1 (productive regime) > threshold | Each detector is applied as a daily gate on top of the weekly-rebalance TopkDropout (topk=10, n_drop=1, yesterday's scores). We run 14 configs across 5 walk-forward windows (2021–2026), tracking trip rate (fraction of days gate is open) and gated return. **Trip rates (2026 vs bad years 2021/2023/2024):** | Gate | 2026 trip | Bad-years avg | Differential | |------|-----------|---------------|-------------| | `vol_low_max20` | 92% | 60% | +32pp | | `vol_low_max25` | 63% | 16% | +47pp | | `hmm_0.7` | 37% | 31% | +6pp | | All dispersion | 0% | 0% | 0pp | The vol gates show the largest trip differential — they open on more days in 2026 than in bad years. But the gate **closes on the wrong days**: when the gate is open only 63% of the time (vol_low_max25), the 2026 return collapses from +25.5% to −1.6%. The gate eliminates the profitable days along with the bad ones. **Gated returns:** | Gate | 2026 base | 2026 gated | 2023 base | 2023 gated | 2025 base | 2025 gated | |------|-----------|------------|-----------|------------|-----------|------------| | `vol_low_max20` | +25.5% | +4.3% | −4.8% | −5.3% | +17.8% | +14.5% | | `hmm_0.7` | +25.5% | +10.8% | −4.8% | +0.6% | +17.8% | +26.8% | `hmm_0.7` has the most interesting profile: it **improves** 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but **destroys** 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods. **Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**. `TODO(evidence-needed: a retrospective signal-quality gate — did yesterday's topk signals predict today's returns? — tested out-of-sample)` ## Open questions - `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)` - `TODO(evidence-needed: a retrospective signal-quality gate — did yesterday's topk signals predict today's returns? — tested out-of-sample)` ## Evidence cited in this chapter | Tag | Source | |-----|--------| | `EVIDENCE#043` | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` | | `EVIDENCE#044` | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference run `c7c12228` (exp 33, Q01) | | `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) | | `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` | | `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` | | Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds | | `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` | | `EVIDENCE#050` | Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation `book/scripts/regime_gate_bt.py`, results `book/data/regime_gate/regime_gate_trip_rates.csv` | | `EVIDENCE#051` | Comprehensive model search: all experiments ranked by RankICIR; regime gate study robust to model selection; `rd_exp_list` + `rd_exp_get_run` queries |