ch11: signal-quality gate REFUTED — walk-forward workflow shows gate harmful (exp 61-67, EVIDENCE#053)
This commit is contained in:
+2
-1
@@ -85,7 +85,8 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| EVIDENCE#051 | Comprehensive model search: queried all MLflow experiments/runs, ranked by RankICIR. Top models: exp 36/44 (label22d, RankICIR 0.507, single-window 2026 only), exp 35/51 (label10d, RankICIR 0.352, single-window), exp 58 (adaptive-2y, RankICIR 0.289), exp 11 (single-seed, RankICIR 0.276). Exp 52 walk-forward configs rank near the top among multi-year models (RankICIR 0.244). The 22-day label models have highest IC but negative returns (−4.6%) — high IC does not guarantee profitable trading. The regime gate study (EVIDENCE#050) is robust to model selection because it measures market-level features, not model predictions. Selection bias is not material: the best-return model (Config C) also has the best RankICIR among walk-forward configs. | `rd_exp_list` query across all MLflow experiments, run metadata from `rd_exp_get_run` for exp 11/33/36/58/52 | yes — robustness check |
|
||||
| EVIDENCE#052 | Signal-quality gate (hit-rate based on topk predictions): gates trades based on whether the model's recent topk predictions were correct. **Every config improves returns across ALL years** — including bad years (2023: −4.8% → +54.7%, 2024: +8.2% → +30.4%). Best config (`hitrate_5d_0.50`): 2026 +65.0% (base +25.5%), 2025 +72.1% (base +17.8%), 2024 +30.4% (base +8.2%), 2023 +54.7% (base −4.8%), 2021 +55.7% (base +18.4%). Gate trips ~40–50% of days. The regime gate (EVIDENCE#050) failed because it asked "is the market calm?" — the signal-quality gate asks "are my predictions accurate?" and succeeds. The model's predictions ARE informative; they just need to be gated on their own accuracy. | scripted simulation: `book/scripts/signal_quality_gate_bt.py`, results `book/data/signal_quality_gate/signal_quality_gate_results.csv`, pred.pkl from exp 52 (2024–2026) and exp 56 (2021, 2023) | yes — signal-quality gate PROVEN |
|
||||
| EVIDENCE#052 | **REFUTED by EVIDENCE#053.** Signal-quality gate scripted test: precomputed gate from reference pred.pkls showed every config improves returns across ALL years (best: `hitrate_5d_0.50` 2026 +65.0%, 2025 +72.1%, 2024 +30.4%, 2023 +54.7%, 2021 +55.7%). **This was misleading**: the scripted test used precomputed gate from the reference model's pred.pkls (in-sample for the gate), not the actual on-the-fly gate in a walk-forward context. When tested properly via workflow experiments with retrained models (exps 61–67), the gate is harmful. | scripted simulation (original), refuted by exps 61–67 | **REFUTED** — scripted test was in-sample for the gate; walk-forward workflow tests show the gate hurts |
|
||||
| EVIDENCE#053 | Signal-quality gate walk-forward refutation: `WeeklyRebalanceSignalQualityGateStrategy` (topk=10, n_drop=1, gate_topk=10, gate_lookback=5, gate_threshold=0.5, 5/15bp costs) tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026). **The gate is harmful in every year.** Workflow excess-with-cost: 2026 +9.1% (IR 0.92) vs reference +12.5% (IR 1.24, exp 38); 2025 +3.4% (IR 0.31); 2024 −20.5%; 2023 −29.4%; 2021 −18.4%. Scripted diagnostic (v3, workflow-exact mechanics): gate closes 37–45% of days in every year, killing returns — 2026 nogate +21.2% total → gate +1.5% total (−19.7pp); 2025 +22.7% → +7.8% (−14.9pp). The gate's hit-rate threshold (0.5) is too aggressive: a model with Rank IC 0.06–0.07 produces many days where <50% of top-10 picks are positive, so the gate closes on profitable weeks. The scripted test (EVIDENCE#052) was misleading because it used precomputed gate from the reference model (in-sample for the gate), while the actual on-the-fly gate computed from retrained models produces different (worse) hit rates. **Guard 7 (signal-quality gate) is REFUTED.** | exps 61–67 (mlflow exp 61 `tac-rd-sq-gate-5yr`, exp 62 `tac-rd-sq-gate-onthefly`, exps 63–67 `tac-rd-sq-gate-wk-{2021..2026}`); scripted diagnostic `book/scripts/diagnose_script_vs_workflow_v3.py`, results `book/data/diag_script_vs_wf/diagnosis_v3.json`; strategy `tac_qlib/contrib/strategy/weekly_sq_gate.py` | yes — guard 7 REFUTED |
|
||||
|
||||
## External references (book/references/)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user