book: add ch 11 walk-forward + 5 refuted guards (exp 52-56); EVIDENCE#043-047; rename 12→13-synthesis
This commit is contained in:
@@ -0,0 +1,87 @@
|
||||
# Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 11.
|
||||
|
||||
This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.
|
||||
|
||||
## The walk-forward re-validation (exp 52, 53, 54)
|
||||
|
||||
Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (`42,7,2026,99,123`, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel):
|
||||
|
||||
### 3×3 sweep (exp 52) — the three best configs across 2024/2025/2026
|
||||
|
||||
Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost.
|
||||
|
||||
| Config | 2024 (test) | 2025 (test) | 2026 (test) |
|
||||
|--------|-------------|-------------|-------------|
|
||||
| **A** weekly, n_drop=1 | −18.1% (IR −1.39, maxDD −22.7%) | −4.2% (IR −0.52, maxDD −10.3%) | **+12.5% (IR 1.25, maxDD −4.1%)** |
|
||||
| **B** moments, n_drop=1 | −16.1% (IR −1.91, maxDD −20.0%) | −8.9% (IR −1.00, maxDD −9.8%) | **+9.2% (IR 0.94, maxDD −6.5%)** |
|
||||
| **C** base, n_drop=2 | −18.2% (IR −1.97, maxDD −23.3%) | −3.6% (IR −0.51, maxDD −6.3%) | **−1.4% (IR −0.13, maxDD −9.8%)** |
|
||||
|
||||
`PROVEN — EVIDENCE#043 → exp 52`. Read the table carefully:
|
||||
|
||||
- **The 2026 window is the only profitable one**, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026.
|
||||
- **A and C share identical predictions** — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026.
|
||||
- **The harness is reproducible**: run 2 (Config A, 2025) exactly replicated exp 45 (`e5ac7a5d`, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (`eb38588c`, +12.51%, IR 1.24).
|
||||
|
||||
### m2-sharpe22 3-window (exp 53)
|
||||
|
||||
The m2-sharpe22 reference (exp 33/Q01, `c7c12228`) re-run on the same 2024/2025/2026 scheme: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. `PROVEN — EVIDENCE#044 → exp 53`. The edge is recent-window-only.
|
||||
|
||||
### Label-regime transfer (exp 54) — 2021 and 2023
|
||||
|
||||
The 2025 feature-drift check had already shown a naive feature-PSI gate does **not** predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the *label/return regime*: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run:
|
||||
|
||||
- **2023**: −26.0% (IR −2.04, maxDD −30.8%)
|
||||
- **2021**: −22.9% (IR −2.26, maxDD −27.0%)
|
||||
|
||||
`PROVEN — EVIDENCE#045 → exp 54`. Both closest label-regime matches are as bad as the 2024 tail. **No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year.** 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups.
|
||||
|
||||
## The five guard candidates — all refuted
|
||||
|
||||
With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed:
|
||||
|
||||
| # | Guard | Test | Result |
|
||||
|---|-------|------|--------|
|
||||
| 1 | **Feature-PSI gate** | halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) | REFUTED — 2026 has the *highest* drift yet the *best* result; CSRankNorm'd ranks make raw drift scale-invariant. |
|
||||
| 2 | **Label-regime gate** | trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) | REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up. |
|
||||
| 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. |
|
||||
| 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. |
|
||||
| 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. |
|
||||
|
||||
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`.
|
||||
|
||||
## The account-level truth
|
||||
|
||||
The blotter's daily `account` field is the authoritative measure (the `return` field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: **2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%**. This reconciles with the recorded metrics — 2026 `excess_return_with_cost` annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. `PROVEN — EVIDENCE#047 → exp 56` (account curves from the exp 53/54 runs' blotter artifacts).
|
||||
|
||||
## The synthesis
|
||||
|
||||
- **The headline results were window-specific.** Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. `PROVEN — EVIDENCE#043/044`.
|
||||
- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. `PROVEN — EVIDENCE#043–047`.
|
||||
- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
|
||||
- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
|
||||
|
||||
## Desk rules distilled from this chapter
|
||||
|
||||
1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
|
||||
2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
|
||||
3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel.
|
||||
4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account.
|
||||
5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.
|
||||
|
||||
## Open questions
|
||||
|
||||
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
|
||||
- `TODO(evidence-needed: a regime-change detector that is causal (no lookahead) and demonstrably selects the 2026 window before the fact — none of the five guards did)`
|
||||
|
||||
## Evidence cited in this chapter
|
||||
|
||||
| Tag | Source |
|
||||
|-----|--------|
|
||||
| `EVIDENCE#043` | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` |
|
||||
| `EVIDENCE#044` | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference run `c7c12228` (exp 33, Q01) |
|
||||
| `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) |
|
||||
| `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` |
|
||||
| `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` |
|
||||
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
|
||||
Reference in New Issue
Block a user