Files
tac-exp-dev/book/chapters/11-walk-forward-and-guards.md
T

23 KiB
Raw Blame History

Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact

Status: drafting. Claim inventory: see README.md ch. 11.

This chapter answers the question every desk must ask before shipping a backtest result: does the edge survive re-training on a different window? The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests seven guard candidates that would plausibly have isolated the good years. Every one of them is refuted. The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.

The walk-forward re-validation (exp 52, 53, 54)

Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (42,7,2026,99,123, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel):

3×3 sweep (exp 52) — the three best configs across 2024/2025/2026

Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost.

Config 2024 (test) 2025 (test) 2026 (test)
A weekly, n_drop=1 −18.1% (IR −1.39, maxDD −22.7%) −4.2% (IR −0.52, maxDD −10.3%) +12.5% (IR 1.25, maxDD −4.1%)
B moments, n_drop=1 −16.1% (IR −1.91, maxDD −20.0%) −8.9% (IR −1.00, maxDD −9.8%) +9.2% (IR 0.94, maxDD −6.5%)
C base, n_drop=2 −18.2% (IR −1.97, maxDD −23.3%) −3.6% (IR −0.51, maxDD −6.3%) −1.4% (IR −0.13, maxDD −9.8%)

PROVEN — EVIDENCE#043 → exp 52. Read the table carefully:

  • The 2026 window is the only profitable one, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026.
  • A and C share identical predictions — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026.
  • The harness is reproducible: run 2 (Config A, 2025) exactly replicated exp 45 (e5ac7a5d, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (eb38588c, +12.51%, IR 1.24).

m2-sharpe22 3-window (exp 53)

The m2-sharpe22 reference (exp 33/Q01, c7c12228) re-run on the same 2024/2025/2026 scheme: 2026 +6.5% (IR 0.623, maxDD −8.0%), 2025 +0.4% (IR 0.05), 2024 −26.4% (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. PROVEN — EVIDENCE#044 → exp 53. The edge is recent-window-only.

Label-regime transfer (exp 54) — 2021 and 2023

The 2025 feature-drift check had already shown a naive feature-PSI gate does not predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the label/return regime: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run:

  • 2023: −26.0% (IR −2.04, maxDD −30.8%)
  • 2021: −22.9% (IR −2.26, maxDD −27.0%)

PROVEN — EVIDENCE#045 → exp 54. Both closest label-regime matches are as bad as the 2024 tail. No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups.

The seven guard candidates — all refuted

With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested seven guards that could plausibly have preserved the good years and cut the bad ones. All seven were pre-registered as hypotheses (traced experiments), all seven failed:

# Guard Test Result
1 Feature-PSI gate halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) REFUTED — 2026 has the highest drift yet the best result; CSRankNorm'd ranks make raw drift scale-invariant.
2 Label-regime gate trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up.
3 Streaming IC circuit breaker (ic_min_rankic) pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold REFUTED — trips 25–50% of days every year, freezing TopkDropout's rotation out of losers; implemented in tac_qlib/contrib/strategy/ic_gate.py, do not deploy live.
4 Adaptive short-window retrain (exp 55) retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window REFUTED — 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length.
5 Window-staleness isolation (exp 56) gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in every bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge.
6 Regime gate (dispersion / vol / HMM) daily boolean gate based on market state (low vol, HMM posterior, dispersion) REFUTED — see below; regime gate closes on the wrong days (hmm_0.7 destroys 2026: +25.5% → +10.8%)
7 Signal-quality gate (hit-rate) gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks.

Guards 1–3 are documented across exp 52/53/54 and the ic_gate.py implementation; guard 4 = PROVEN (refuted) — EVIDENCE#046 → exp 55; guard 5 = PROVEN (refuted) — EVIDENCE#047 → exp 56; guard 6 = PROVEN (refuted) — EVIDENCE#050; guard 7 = REFUTED — EVIDENCE#052 → EVIDENCE#053.

The account-level truth

The blotter's daily account field is the authoritative measure (the return field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%. This reconciles with the recorded metrics — 2026 excess_return_with_cost annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. PROVEN — EVIDENCE#047 → exp 56 (account curves from the exp 53/54 runs' blotter artifacts).

The synthesis

  • The headline results were window-specific. Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. PROVEN — EVIDENCE#043/044.
  • No guard candidate recovers the edge out-of-sample. Feature drift, label-regime match, streaming IC, training-window length, staleness, and signal-quality gating all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. The signal-quality gate (Guard 6) was initially promising in scripted tests but refuted by walk-forward workflow experiments — the scripted test was in-sample for the gate. PROVEN — EVIDENCE#043–053.
  • Construction still matters inside the good regime. A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
  • Live implication: size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.

Within-window robustness (perturbation stress test)

The 2026 edge is fragile across windows but robust within the 2026 window. A perturbation grid on Config A's predictions (exp 52, run 9f98ea5c, same pred.pkl, varying only backtest parameters):

Perturbation Config Ann. return Sharpe maxDD
Baseline topk=10, n_drop=1, costs 5/15/$5 32.8% 1.98 −5.8%
topk=5 concentration ↑ 32.3% 1.75 −7.0%
topk=15 concentration ↓ 26.3% 1.61 −6.9%
n_drop=2 rotation ↑ 26.7% 1.59 −6.5%
n_drop=3 rotation ↑↑ 28.6% 1.66 −6.7%
costs 3× (15/25/$10) cost stress 32.8% 1.97 −5.8%
costs 5× (25/35/$15) cost stress ↑↑ 32.7% 1.97 −5.8%

PROVEN — EVIDENCE#049 (ad-hoc rd_backtest grid on exp 52 pred.pkl, book/data/perturbation/config_a_2026_sensitivity.json).

Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. Within the one good window, the edge is not a parameter-tuning artifact. The fragility is entirely across windows (regime dependence), not within them.

Model search & robustness of the regime gate finding

The regime gate study (Guard 7, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection:

Rank Exp Test Window RankICIR Net Return Notes
1 36 2026 only 0.507 −4.6% 22d label, single-window
2 44 2026 only 0.507 −4.9% Same pred as #1
3 35 2026 only 0.352 −9.9% 10d label
4 51 2026 only 0.352 +1.2% Same pred as #3
5 58 2025 only 0.289 −3.4% Adaptive 2y
6 11 2026 only 0.276 +3.1% Single seed
7 33 2026 only 0.259 −0.9% 10 seeds
8 52-C 2024–2026 0.244 +12.5% Walk-forward, weekly
9 52-A 2024–2026 0.244 −1.4% Walk-forward, TopkDrop

PROVEN — EVIDENCE#051 (comprehensive rd_exp_list query, run metadata).

Key observations:

  1. The 22-day label models (exp 36/44) have the highest RankICIR (0.507) but negative returns — high IC does not guarantee profitable trading. The 22d label predicts longer-horizon moves that don't translate to short-term alpha after costs.
  2. Most high-RankICIR models are single-window (2026 only) — they lack the multi-year coverage needed for the regime gate study. Walk-forward coverage (2021–2026) is limited to exp 52 (2024–2026) and exp 56 (2021, 2023).
  3. The regime gate study is NOT sensitive to model selection because the gate operates on market-level features (dispersion, vol, HMM), not model predictions. Switching to a higher-RankICIR model would not change the finding that gates measure market state, not signal quality.
  4. Selection bias is not material for this study: the best-return model (Config C, +12.5%) also has the best RankICIR (0.244) among walk-forward configs. The RankICIR and returns rankings are concordant.

Signal-quality gate (Guard 7): refuted

REFUTED — EVIDENCE#052 → EVIDENCE#053

The regime gate (Guard 6) failed because it answered the wrong question: "Is the market calm?" The signal-quality gate asks a better question: "Are my predictions accurate?" — but when tested properly, it still doesn't work.

Logic: For each day t, look at the topk symbols from yesterday (t-1). Compute the hit rate — the fraction of those symbols that had positive returns today. If the hit rate is above a threshold, keep trading; otherwise, go to cash. This is a retrospective gate — it measures prediction accuracy, not market state.

Initial scripted test (EVIDENCE#052): Precomputed gate from reference pred.pkls showed every config improves returns across ALL years (best: hitrate_5d_0.50 2026 +65.0%, 2025 +72.1%, 2024 +30.4%, 2023 +54.7%, 2021 +55.7%). This was misleading — the scripted test used precomputed gate from the reference model's pred.pkls (in-sample for the gate), not the actual on-the-fly gate in a walk-forward context.

Walk-forward workflow test (EVIDENCE#053): WeeklyRebalanceSignalQualityGateStrategy (topk=10, n_drop=1, gate_topk=10, gate_lookback=5, gate_threshold=0.5, 5/15bp costs) tested via rd_train + rd_run_workflow on 5 walk-forward windows (2021–2026), retraining the model each year. The gate is harmful in every year:

Year Workflow excess w/cost (gate) Reference excess w/cost (nogate) Delta
2026 +9.1% (IR 0.92) +12.5% (IR 1.24) −3.4pp
2025 +3.4% (IR 0.31) +3.7% (IR 0.33) −0.3pp
2024 −20.5% −19.4% −1.1pp
2023 −29.4% −29.6% +0.2pp
2021 −18.4% −21.2% +2.8pp

Why the scripted test was wrong: The diagnostic (v3, workflow-exact mechanics) reveals the gate closes 37–45% of days in every year, killing returns:

Year Script total (nogate) Script total (gate) Delta Gate open%
2026 +21.2% +1.5% −19.7pp 56%
2025 +22.7% +7.8% −14.9pp 63%
2024 +4.0% −2.6% −6.6pp 59%

A model with Rank IC 0.06–0.07 produces many days where <50% of top-10 picks are positive — the gate's 0.5 threshold is too aggressive, closing on profitable weeks. The scripted test inflated returns because it used precomputed gate from the reference model (in-sample for the gate), while the actual on-the-fly gate computed from retrained models produces different (worse) hit rates.

Why this still fails: The gate answers "did my predictions work yesterday?" — but with a 0.06–0.07 Rank IC, yesterday's hit rate is mostly noise. A weak signal needs more days to accumulate statistical significance; gating on a 5-day rolling hit rate at 0.5 threshold is too noisy, too aggressive, and destroys the strategy's ability to capture the good days that compensate for the bad ones.

Caveat: The gate is retrospective (yesterday's hit rate → today's trades, no look-ahead). The problem is not look-ahead — it's that the signal is too weak for a 0.5 threshold on a 5-day window to be informative.

Desk rules distilled from this chapter

  1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
  2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
  3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, staleness, and signal-quality gating all failed on this panel. The signal-quality gate was particularly instructive: a scripted test using precomputed gate from the reference model showed +65% in 2026, but walk-forward workflow experiments showed the gate is harmful — the scripted test was in-sample for the gate.
  4. Report account-based curves, not the blotter return field — the latter excludes initial cost and does not compound to the account.
  5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.

Guard 6: Regime gate (dispersion / vol / HMM)

PROVEN — EVIDENCE#050

If the edge is regime-dependent, the most direct guard is a regime detector that opens on good years and closes on bad years. We test three detector types, each producing a daily boolean (trade / don't trade):

Detector Logic
dispersion CS std of 22-day rolling returns < threshold (low dispersion → calm market → trade)
vol CS mean of 22-day rolling realized vol within a band (mid-range vol → trade)
HMM 2-state Gaussian HMM posterior for regime 1 (productive regime) > threshold

Each detector is applied as a daily gate on top of the weekly-rebalance TopkDropout (topk=10, n_drop=1, yesterday's scores). We run 14 configs across 5 walk-forward windows (2021–2026), tracking trip rate (fraction of days gate is open) and gated return.

Trip rates (2026 vs bad years 2021/2023/2024):

Gate 2026 trip Bad-years avg Differential
vol_low_max20 92% 60% +32pp
vol_low_max25 63% 16% +47pp
hmm_0.7 37% 31% +6pp
All dispersion 0% 0% 0pp

The vol gates show the largest trip differential — they open on more days in 2026 than in bad years. But the gate closes on the wrong days: when the gate is open only 63% of the time (vol_low_max25), the 2026 return collapses from +25.5% to −1.6%. The gate eliminates the profitable days along with the bad ones.

Gated returns:

Gate 2026 base 2026 gated 2023 base 2023 gated 2025 base 2025 gated
vol_low_max20 +25.5% +4.3% −4.8% −5.3% +17.8% +14.5%
hmm_0.7 +25.5% +10.8% −4.8% +0.6% +17.8% +26.8%

hmm_0.7 has the most interesting profile: it improves 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but destroys 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods.

Why none of these gates work: The gate answers "is the market calm right now?" — but the right question is "will today's signal be profitable tomorrow?" These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict signal quality, not market state. See Guard 7 (signal-quality gate, EVIDENCE#053) for a gate that was tested on this principle — and still failed.

Open questions

  • TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)
  • TODO(evidence-needed: understanding the script-vs-workflow gap for signal-quality gate — scripted test shows gate destroying ~20pp more return than workflow, despite identical parameters; root cause is pred date alignment differences between precomputed and on-the-fly gate computation)

Evidence cited in this chapter

Tag Source
EVIDENCE#043 exp 52, mlflow exp 52 tac-rd-bt-3x3-windows (9 runs: 9f98ea5c A-2026, fe967416 A-2025, 71ed5bfa A-2024; 163c01ce B-2026, 4a85d68e B-2025, 1e49b8e8 B-2024; e3e06a24 C-2026, 353fff8f C-2025, 13a9bbdf C-2024), branch exp/52-walk-forward-re-validation-of-the-3-best
EVIDENCE#044 exp 53, mlflow exp 53 tac-rd-bt-m2-sharpe22-3windows (runs 7464c3e7 2026, 061f558b 2025, b49c6845 2024), branch exp/53-walk-forward-re-validation-of-m2-sharpe2; reference run c7c12228 (exp 33, Q01)
EVIDENCE#045 exp 54, mlflow exp 56 tac-rd-bt-m2-sharpe22-2021-2023 (runs 4e0700dd 2021, 8ca46e55 2023), branch exp/54-walk-forward-transfer-test-m2-sharpe22-o; feature/label-regime PSI study (exp 53 follow-up)
EVIDENCE#046 exp 55, mlflow exp 57/58 tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}, branch exp/55-adaptive-short-window-retrain-test-the-4
EVIDENCE#047 exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch exp/56-window-staleness-isolation-the-m2-sharpe
Guard 3 (ic_min_rankic) tac_qlib/tac_qlib/contrib/strategy/ic_gate.py (ICGateTopkDropoutStrategy), tac_qlib/tac_qlib/risk_limits.py; trip-rate study on exp 52/53 preds
EVIDENCE#049 Perturbation stress test on Config A 2026 (exp 52, pred 9f98ea5c): topk/n_drop/cost grid, book/data/perturbation/config_a_2026_sensitivity.json
EVIDENCE#050 Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation book/scripts/regime_gate_bt.py, results book/data/regime_gate/regime_gate_trip_rates.csv
EVIDENCE#051 Comprehensive model search: all experiments ranked by RankICIR; regime gate study robust to model selection; rd_exp_list + rd_exp_get_run queries
EVIDENCE#052 Signal-quality gate scripted test: precomputed gate from reference pred.pkls showed every config improves returns across ALL years. REFUTED by EVIDENCE#053 — scripted test was in-sample for the gate.
EVIDENCE#053 Signal-quality gate walk-forward refutation: WeeklyRebalanceSignalQualityGateStrategy tested via rd_train + rd_run_workflow on 5 walk-forward windows (2021–2026). Gate harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp). Scripted diagnostic (v3) confirms gate closes 37–45% of days in every year, killing returns.