ch11: add EVIDENCE#051 — comprehensive model search confirms regime gate robustness to model selection
This commit is contained in:
@@ -80,6 +80,12 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
|
||||
| EVIDENCE#049 | Perturbation stress test on Config A 2026 (exp 52, pred from run `9f98ea5c`): same signal, varying topk (5/10/15), n_drop (1/2/3), costs (base/high/5×base). **topk**: 10 optimal (32.8% raw, Sharpe 1.98); 5 loses ~0.5pp, 15 loses ~6.5pp. **n_drop**: 1 optimal; 2 loses ~6pp, 3 loses ~4pp. **costs**: immaterial — 5× cost increase (25bp/35bp/$15) drops return only 0.17pp (32.84%→32.67%). maxDD stable −5.8% to −7.0% across all perturbations. **Within the 2026 window the edge is robust to parameter perturbation.** The problem remains that it does not exist in other windows (ch 11). | ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json` | yes — within-window robustness confirmed |
|
||||
| EVIDENCE#050 | Regime gate walk-forward test across 5 years (2021–2026): three detector types (dispersion, vol, HMM) × 14 configs. **Dispersion gates**: 0% trip rate everywhere — CS std of 22d returns never crosses any threshold. **Vol gates** (best: `vol_low_max20`): opens 92% in 2026 vs 60% in bad years (+32pp differential), but 2026 gated return collapses from +25.5% to +4.3% — the gate closes on profitable days. **HMM gates** (best: `hmm_0.7`): opens 37% in 2026 vs 31% in bad years (+6pp differential), 2026 return drops from +25.5% to +10.8%. No detector type achieves the goal of selective protection: tripping more in bad years while preserving good-year returns. The gate measures current market state, not whether yesterday's signals will predict today's returns. | scripted simulation: `book/scripts/regime_gate_bt.py`, results `book/data/regime_gate/regime_gate_trip_rates.csv`, pred.pkl from exp 52 (2024–2026) and exp 56 (2021, 2023) | yes — guard candidate regime gate REFUTED |
|
||||
|
||||
## Model search & robustness (exp 52 context)
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| EVIDENCE#051 | Comprehensive model search: queried all MLflow experiments/runs, ranked by RankICIR. Top models: exp 36/44 (label22d, RankICIR 0.507, single-window 2026 only), exp 35/51 (label10d, RankICIR 0.352, single-window), exp 58 (adaptive-2y, RankICIR 0.289), exp 11 (single-seed, RankICIR 0.276). Exp 52 walk-forward configs rank near the top among multi-year models (RankICIR 0.244). The 22-day label models have highest IC but negative returns (−4.6%) — high IC does not guarantee profitable trading. The regime gate study (EVIDENCE#050) is robust to model selection because it measures market-level features, not model predictions. Selection bias is not material: the best-return model (Config C) also has the best RankICIR among walk-forward configs. | `rd_exp_list` query across all MLflow experiments, run metadata from `rd_exp_get_run` for exp 11/33/36/58/52 | yes — robustness check |
|
||||
|
||||
## External references (book/references/)
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|
||||
Reference in New Issue
Block a user