Files
tac-exp-dev/book/chapters/11-walk-forward-and-guards.md
T

11 KiB
Raw Blame History

Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact

Status: drafting. Claim inventory: see README.md ch. 11.

This chapter answers the question every desk must ask before shipping a backtest result: does the edge survive re-training on a different window? The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. Every one of them is refuted. The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.

The walk-forward re-validation (exp 52, 53, 54)

Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (42,7,2026,99,123, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel):

3×3 sweep (exp 52) — the three best configs across 2024/2025/2026

Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost.

Config 2024 (test) 2025 (test) 2026 (test)
A weekly, n_drop=1 −18.1% (IR −1.39, maxDD −22.7%) −4.2% (IR −0.52, maxDD −10.3%) +12.5% (IR 1.25, maxDD −4.1%)
B moments, n_drop=1 −16.1% (IR −1.91, maxDD −20.0%) −8.9% (IR −1.00, maxDD −9.8%) +9.2% (IR 0.94, maxDD −6.5%)
C base, n_drop=2 −18.2% (IR −1.97, maxDD −23.3%) −3.6% (IR −0.51, maxDD −6.3%) −1.4% (IR −0.13, maxDD −9.8%)

PROVEN — EVIDENCE#043 → exp 52. Read the table carefully:

  • The 2026 window is the only profitable one, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026.
  • A and C share identical predictions — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026.
  • The harness is reproducible: run 2 (Config A, 2025) exactly replicated exp 45 (e5ac7a5d, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (eb38588c, +12.51%, IR 1.24).

m2-sharpe22 3-window (exp 53)

The m2-sharpe22 reference (exp 33/Q01, c7c12228) re-run on the same 2024/2025/2026 scheme: 2026 +6.5% (IR 0.623, maxDD −8.0%), 2025 +0.4% (IR 0.05), 2024 −26.4% (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. PROVEN — EVIDENCE#044 → exp 53. The edge is recent-window-only.

Label-regime transfer (exp 54) — 2021 and 2023

The 2025 feature-drift check had already shown a naive feature-PSI gate does not predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the label/return regime: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run:

  • 2023: −26.0% (IR −2.04, maxDD −30.8%)
  • 2021: −22.9% (IR −2.26, maxDD −27.0%)

PROVEN — EVIDENCE#045 → exp 54. Both closest label-regime matches are as bad as the 2024 tail. No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups.

The five guard candidates — all refuted

With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed:

# Guard Test Result
1 Feature-PSI gate halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) REFUTED — 2026 has the highest drift yet the best result; CSRankNorm'd ranks make raw drift scale-invariant.
2 Label-regime gate trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up.
3 Streaming IC circuit breaker (ic_min_rankic) pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold REFUTED — trips 25–50% of days every year, freezing TopkDropout's rotation out of losers; implemented in tac_qlib/contrib/strategy/ic_gate.py, do not deploy live.
4 Adaptive short-window retrain (exp 55) retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window REFUTED — 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length.
5 Window-staleness isolation (exp 56) gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in every bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge.

Guards 1–3 are documented across exp 52/53/54 and the ic_gate.py implementation; guard 4 = PROVEN (refuted) — EVIDENCE#046 → exp 55; guard 5 = PROVEN (refuted) — EVIDENCE#047 → exp 56.

The account-level truth

The blotter's daily account field is the authoritative measure (the return field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%. This reconciles with the recorded metrics — 2026 excess_return_with_cost annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. PROVEN — EVIDENCE#047 → exp 56 (account curves from the exp 53/54 runs' blotter artifacts).

The synthesis

  • The headline results were window-specific. Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. PROVEN — EVIDENCE#043/044.
  • No guard candidate recovers the edge out-of-sample. Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. PROVEN — EVIDENCE#043–047.
  • Construction still matters inside the good regime. A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
  • Live implication: size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.

Desk rules distilled from this chapter

  1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
  2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
  3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel.
  4. Report account-based curves, not the blotter return field — the latter excludes initial cost and does not compound to the account.
  5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.

Open questions

  • TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)
  • TODO(evidence-needed: a regime-change detector that is causal (no lookahead) and demonstrably selects the 2026 window before the fact — none of the five guards did)

Evidence cited in this chapter

Tag Source
EVIDENCE#043 exp 52, mlflow exp 52 tac-rd-bt-3x3-windows (9 runs: 9f98ea5c A-2026, fe967416 A-2025, 71ed5bfa A-2024; 163c01ce B-2026, 4a85d68e B-2025, 1e49b8e8 B-2024; e3e06a24 C-2026, 353fff8f C-2025, 13a9bbdf C-2024), branch exp/52-walk-forward-re-validation-of-the-3-best
EVIDENCE#044 exp 53, mlflow exp 53 tac-rd-bt-m2-sharpe22-3windows (runs 7464c3e7 2026, 061f558b 2025, b49c6845 2024), branch exp/53-walk-forward-re-validation-of-m2-sharpe2; reference run c7c12228 (exp 33, Q01)
EVIDENCE#045 exp 54, mlflow exp 56 tac-rd-bt-m2-sharpe22-2021-2023 (runs 4e0700dd 2021, 8ca46e55 2023), branch exp/54-walk-forward-transfer-test-m2-sharpe22-o; feature/label-regime PSI study (exp 53 follow-up)
EVIDENCE#046 exp 55, mlflow exp 57/58 tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}, branch exp/55-adaptive-short-window-retrain-test-the-4
EVIDENCE#047 exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch exp/56-window-staleness-isolation-the-m2-sharpe
Guard 3 (ic_min_rankic) tac_qlib/tac_qlib/contrib/strategy/ic_gate.py (ICGateTopkDropoutStrategy), tac_qlib/tac_qlib/risk_limits.py; trip-rate study on exp 52/53 preds