Add regime gate walk-forward test (EVIDENCE#050)

- 3 detector types (dispersion/vol/HMM) × 14 configs across 5 years
- Dispersion gates: 0% trip rate everywhere (dead)
- Vol gates: trip differential +32-47pp but destroy returns in good years
- HMM gates: +6pp differential, hmm_0.7 improves 2023/2025 but kills 2026
- Guard candidate regime gate REFUTED (ch 11)
- Script: book/scripts/regime_gate_bt.py
- Results: book/data/regime_gate/regime_gate_trip_rates.csv
This commit is contained in:
zhaoli
2026-08-20 22:14:31 +00:00
parent 4831a2770c
commit 718b048df9
6 changed files with 1636 additions and 2 deletions
+41 -2
View File
@@ -88,10 +88,48 @@ Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp).
4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account.
5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.
### Guard 6: Regime gate (dispersion / vol / HMM)
`PROVEN — EVIDENCE#050`
If the edge is regime-dependent, the most direct guard is a regime detector that opens on good years and closes on bad years. We test three detector types, each producing a daily boolean (trade / don't trade):
| Detector | Logic |
|----------|-------|
| **dispersion** | CS std of 22-day rolling returns < threshold (low dispersion → calm market → trade) |
| **vol** | CS mean of 22-day rolling realized vol within a band (mid-range vol → trade) |
| **HMM** | 2-state Gaussian HMM posterior for regime 1 (productive regime) > threshold |
Each detector is applied as a daily gate on top of the weekly-rebalance TopkDropout (topk=10, n_drop=1, yesterday's scores). We run 14 configs across 5 walk-forward windows (2021–2026), tracking trip rate (fraction of days gate is open) and gated return.
**Trip rates (2026 vs bad years 2021/2023/2024):**
| Gate | 2026 trip | Bad-years avg | Differential |
|------|-----------|---------------|-------------|
| `vol_low_max20` | 92% | 60% | +32pp |
| `vol_low_max25` | 63% | 16% | +47pp |
| `hmm_0.7` | 37% | 31% | +6pp |
| All dispersion | 0% | 0% | 0pp |
The vol gates show the largest trip differential — they open on more days in 2026 than in bad years. But the gate **closes on the wrong days**: when the gate is open only 63% of the time (vol_low_max25), the 2026 return collapses from +25.5% to −1.6%. The gate eliminates the profitable days along with the bad ones.
**Gated returns:**
| Gate | 2026 base | 2026 gated | 2023 base | 2023 gated | 2025 base | 2025 gated |
|------|-----------|------------|-----------|------------|-----------|------------|
| `vol_low_max20` | +25.5% | +4.3% | −4.8% | −5.3% | +17.8% | +14.5% |
| `hmm_0.7` | +25.5% | +10.8% | −4.8% | +0.6% | +17.8% | +26.8% |
`hmm_0.7` has the most interesting profile: it **improves** 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but **destroys** 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods.
**Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**.
`TODO(evidence-needed: a retrospective signal-quality gate — did yesterday's topk signals predict today's returns? — tested out-of-sample)`
## Open questions
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
- `TODO(evidence-needed: a regime-change detector that is causal (no lookahead) and demonstrably selects the 2026 window before the fact — none of the five guards did)`
- `TODO(evidence-needed: a retrospective signal-quality gate — did yesterday's topk signals predict today's returns? — tested out-of-sample)`
## Evidence cited in this chapter
@@ -103,4 +141,5 @@ Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp).
| `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` |
| `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` |
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |
| `EVIDENCE#050` | Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation `book/scripts/regime_gate_bt.py`, results `book/data/regime_gate/regime_gate_trip_rates.csv` |