ch11: fix guard numbering — regime gate is Guard 6, SQ gate is Guard 7

This commit is contained in:
zhaoli
2026-08-21 02:44:08 +00:00
parent 6e3c82408c
commit eefc26ce8f
+8 -7
View File
@@ -2,7 +2,7 @@
Status: drafting. Claim inventory: see `README.md` ch. 11.
This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests five guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.
This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests seven guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.
## The walk-forward re-validation (exp 52, 53, 54)
@@ -37,9 +37,9 @@ The 2025 feature-drift check had already shown a naive feature-PSI gate does **n
`PROVEN — EVIDENCE#045 → exp 54`. Both closest label-regime matches are as bad as the 2024 tail. **No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year.** 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups.
## The five guard candidates — all refuted
## The seven guard candidates — all refuted
With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested five guards that could plausibly have preserved the good years and cut the bad ones. All five were pre-registered as hypotheses (traced experiments), all five failed:
With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested seven guards that could plausibly have preserved the good years and cut the bad ones. All seven were pre-registered as hypotheses (traced experiments), all seven failed:
| # | Guard | Test | Result |
|---|-------|------|--------|
@@ -48,9 +48,10 @@ With the walk-forward sweep showing the edge is 2026-window-specific, the desk t
| 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. |
| 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. |
| 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. |
| 6 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. The gate destroys the strategy's ability to capture the good days that compensate for the bad ones. |
| 6 | **Regime gate** (dispersion / vol / HMM) | daily boolean gate based on market state (low vol, HMM posterior, dispersion) | REFUTED — see below; regime gate closes on the wrong days (hmm_0.7 destroys 2026: +25.5% → +10.8%) |
| 7 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. |
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`.
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `PROVEN (refuted) — EVIDENCE#050`; guard 7 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`.
## The account-level truth
@@ -83,7 +84,7 @@ Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp).
## Model search & robustness of the regime gate finding
The regime gate study (Guard 6, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection:
The regime gate study (Guard 7, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection:
| Rank | Exp | Test Window | RankICIR | Net Return | Notes |
|------|-----|-------------|----------|------------|-------|
@@ -182,7 +183,7 @@ The vol gates show the largest trip differential — they open on more days in 2
`hmm_0.7` has the most interesting profile: it **improves** 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but **destroys** 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods.
**Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**. See Guard 7 (signal-quality gate, EVIDENCE#052) for a gate that works.
**Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**. See Guard 7 (signal-quality gate, EVIDENCE#053) for a gate that was tested on this principle — and still failed.
## Open questions