- 3 detector types (dispersion/vol/HMM) × 14 configs across 5 years - Dispersion gates: 0% trip rate everywhere (dead) - Vol gates: trip differential +32-47pp but destroy returns in good years - HMM gates: +6pp differential, hmm_0.7 improves 2023/2025 but kills 2026 - Guard candidate regime gate REFUTED (ch 11) - Script: book/scripts/regime_gate_bt.py - Results: book/data/regime_gate/regime_gate_trip_rates.csv
16 KiB
Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact
Status: drafting. Claim inventory: see README.md ch. 11.
This chapter answers the question every desk must ask before shipping a backtest result: does the edge survive re-training on a different window? The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. Every one of them is refuted. The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.
The walk-forward re-validation (exp 52, 53, 54)
Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (42,7,2026,99,123, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel):
3×3 sweep (exp 52) — the three best configs across 2024/2025/2026
Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost.
| Config | 2024 (test) | 2025 (test) | 2026 (test) |
|---|---|---|---|
| A weekly, n_drop=1 | −18.1% (IR −1.39, maxDD −22.7%) | −4.2% (IR −0.52, maxDD −10.3%) | +12.5% (IR 1.25, maxDD −4.1%) |
| B moments, n_drop=1 | −16.1% (IR −1.91, maxDD −20.0%) | −8.9% (IR −1.00, maxDD −9.8%) | +9.2% (IR 0.94, maxDD −6.5%) |
| C base, n_drop=2 | −18.2% (IR −1.97, maxDD −23.3%) | −3.6% (IR −0.51, maxDD −6.3%) | −1.4% (IR −0.13, maxDD −9.8%) |
PROVEN — EVIDENCE#043 → exp 52. Read the table carefully:
- The 2026 window is the only profitable one, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026.
- A and C share identical predictions — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026.
- The harness is reproducible: run 2 (Config A, 2025) exactly replicated exp 45 (
e5ac7a5d, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (eb38588c, +12.51%, IR 1.24).
m2-sharpe22 3-window (exp 53)
The m2-sharpe22 reference (exp 33/Q01, c7c12228) re-run on the same 2024/2025/2026 scheme: 2026 +6.5% (IR 0.623, maxDD −8.0%), 2025 +0.4% (IR 0.05), 2024 −26.4% (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. PROVEN — EVIDENCE#044 → exp 53. The edge is recent-window-only.
Label-regime transfer (exp 54) — 2021 and 2023
The 2025 feature-drift check had already shown a naive feature-PSI gate does not predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the label/return regime: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run:
- 2023: −26.0% (IR −2.04, maxDD −30.8%)
- 2021: −22.9% (IR −2.26, maxDD −27.0%)
PROVEN — EVIDENCE#045 → exp 54. Both closest label-regime matches are as bad as the 2024 tail. No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups.
The five guard candidates — all refuted
With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed:
| # | Guard | Test | Result |
|---|---|---|---|
| 1 | Feature-PSI gate | halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) | REFUTED — 2026 has the highest drift yet the best result; CSRankNorm'd ranks make raw drift scale-invariant. |
| 2 | Label-regime gate | trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) | REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up. |
| 3 | Streaming IC circuit breaker (ic_min_rankic) |
pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days every year, freezing TopkDropout's rotation out of losers; implemented in tac_qlib/contrib/strategy/ic_gate.py, do not deploy live. |
| 4 | Adaptive short-window retrain (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length. |
| 5 | Window-staleness isolation (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in every bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. |
Guards 1–3 are documented across exp 52/53/54 and the ic_gate.py implementation; guard 4 = PROVEN (refuted) — EVIDENCE#046 → exp 55; guard 5 = PROVEN (refuted) — EVIDENCE#047 → exp 56.
The account-level truth
The blotter's daily account field is the authoritative measure (the return field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%. This reconciles with the recorded metrics — 2026 excess_return_with_cost annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. PROVEN — EVIDENCE#047 → exp 56 (account curves from the exp 53/54 runs' blotter artifacts).
The synthesis
- The headline results were window-specific. Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested.
PROVEN — EVIDENCE#043/044. - No guard candidate recovers the edge out-of-sample. Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live.
PROVEN — EVIDENCE#043–047. - Construction still matters inside the good regime. A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
- Live implication: size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
Within-window robustness (perturbation stress test)
The 2026 edge is fragile across windows but robust within the 2026 window. A perturbation grid on Config A's predictions (exp 52, run 9f98ea5c, same pred.pkl, varying only backtest parameters):
| Perturbation | Config | Ann. return | Sharpe | maxDD |
|---|---|---|---|---|
| Baseline | topk=10, n_drop=1, costs 5/15/$5 | 32.8% | 1.98 | −5.8% |
| topk=5 | concentration ↑ | 32.3% | 1.75 | −7.0% |
| topk=15 | concentration ↓ | 26.3% | 1.61 | −6.9% |
| n_drop=2 | rotation ↑ | 26.7% | 1.59 | −6.5% |
| n_drop=3 | rotation ↑↑ | 28.6% | 1.66 | −6.7% |
| costs 3× (15/25/$10) | cost stress | 32.8% | 1.97 | −5.8% |
| costs 5× (25/35/$15) | cost stress ↑↑ | 32.7% | 1.97 | −5.8% |
PROVEN — EVIDENCE#049 (ad-hoc rd_backtest grid on exp 52 pred.pkl, book/data/perturbation/config_a_2026_sensitivity.json).
Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. Within the one good window, the edge is not a parameter-tuning artifact. The fragility is entirely across windows (regime dependence), not within them.
Desk rules distilled from this chapter
- Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
- Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
- Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel.
- Report account-based curves, not the blotter
returnfield — the latter excludes initial cost and does not compound to the account. - When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.
Guard 6: Regime gate (dispersion / vol / HMM)
PROVEN — EVIDENCE#050
If the edge is regime-dependent, the most direct guard is a regime detector that opens on good years and closes on bad years. We test three detector types, each producing a daily boolean (trade / don't trade):
| Detector | Logic |
|---|---|
| dispersion | CS std of 22-day rolling returns < threshold (low dispersion → calm market → trade) |
| vol | CS mean of 22-day rolling realized vol within a band (mid-range vol → trade) |
| HMM | 2-state Gaussian HMM posterior for regime 1 (productive regime) > threshold |
Each detector is applied as a daily gate on top of the weekly-rebalance TopkDropout (topk=10, n_drop=1, yesterday's scores). We run 14 configs across 5 walk-forward windows (2021–2026), tracking trip rate (fraction of days gate is open) and gated return.
Trip rates (2026 vs bad years 2021/2023/2024):
| Gate | 2026 trip | Bad-years avg | Differential |
|---|---|---|---|
vol_low_max20 |
92% | 60% | +32pp |
vol_low_max25 |
63% | 16% | +47pp |
hmm_0.7 |
37% | 31% | +6pp |
| All dispersion | 0% | 0% | 0pp |
The vol gates show the largest trip differential — they open on more days in 2026 than in bad years. But the gate closes on the wrong days: when the gate is open only 63% of the time (vol_low_max25), the 2026 return collapses from +25.5% to −1.6%. The gate eliminates the profitable days along with the bad ones.
Gated returns:
| Gate | 2026 base | 2026 gated | 2023 base | 2023 gated | 2025 base | 2025 gated |
|---|---|---|---|---|---|---|
vol_low_max20 |
+25.5% | +4.3% | −4.8% | −5.3% | +17.8% | +14.5% |
hmm_0.7 |
+25.5% | +10.8% | −4.8% | +0.6% | +17.8% | +26.8% |
hmm_0.7 has the most interesting profile: it improves 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but destroys 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods.
Why none of these gates work: The gate answers "is the market calm right now?" — but the right question is "will today's signal be profitable tomorrow?" These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict signal quality, not market state.
TODO(evidence-needed: a retrospective signal-quality gate — did yesterday's topk signals predict today's returns? — tested out-of-sample)
Open questions
TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)TODO(evidence-needed: a retrospective signal-quality gate — did yesterday's topk signals predict today's returns? — tested out-of-sample)
Evidence cited in this chapter
| Tag | Source |
|---|---|
EVIDENCE#043 |
exp 52, mlflow exp 52 tac-rd-bt-3x3-windows (9 runs: 9f98ea5c A-2026, fe967416 A-2025, 71ed5bfa A-2024; 163c01ce B-2026, 4a85d68e B-2025, 1e49b8e8 B-2024; e3e06a24 C-2026, 353fff8f C-2025, 13a9bbdf C-2024), branch exp/52-walk-forward-re-validation-of-the-3-best |
EVIDENCE#044 |
exp 53, mlflow exp 53 tac-rd-bt-m2-sharpe22-3windows (runs 7464c3e7 2026, 061f558b 2025, b49c6845 2024), branch exp/53-walk-forward-re-validation-of-m2-sharpe2; reference run c7c12228 (exp 33, Q01) |
EVIDENCE#045 |
exp 54, mlflow exp 56 tac-rd-bt-m2-sharpe22-2021-2023 (runs 4e0700dd 2021, 8ca46e55 2023), branch exp/54-walk-forward-transfer-test-m2-sharpe22-o; feature/label-regime PSI study (exp 53 follow-up) |
EVIDENCE#046 |
exp 55, mlflow exp 57/58 tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}, branch exp/55-adaptive-short-window-retrain-test-the-4 |
EVIDENCE#047 |
exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch exp/56-window-staleness-isolation-the-m2-sharpe |
Guard 3 (ic_min_rankic) |
tac_qlib/tac_qlib/contrib/strategy/ic_gate.py (ICGateTopkDropoutStrategy), tac_qlib/tac_qlib/risk_limits.py; trip-rate study on exp 52/53 preds |
EVIDENCE#049 |
Perturbation stress test on Config A 2026 (exp 52, pred 9f98ea5c): topk/n_drop/cost grid, book/data/perturbation/config_a_2026_sensitivity.json |
EVIDENCE#050 |
Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation book/scripts/regime_gate_bt.py, results book/data/regime_gate/regime_gate_trip_rates.csv |