ch11: signal-quality gate REFUTED — walk-forward workflow shows gate harmful (exp 61-67, EVIDENCE#053)

This commit is contained in:
zhaoli
2026-08-21 02:35:50 +00:00
parent 1603a063ba
commit 6e3c82408c
26 changed files with 4934 additions and 27 deletions
+28 -26
View File
@@ -48,8 +48,9 @@ With the walk-forward sweep showing the edge is 2026-window-specific, the desk t
| 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. |
| 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. |
| 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. |
| 6 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. The gate destroys the strategy's ability to capture the good days that compensate for the bad ones. |
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`.
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`.
## The account-level truth
@@ -58,7 +59,7 @@ The blotter's daily `account` field is the authoritative measure (the `return` f
## The synthesis
- **The headline results were window-specific.** Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. `PROVEN — EVIDENCE#043/044`.
- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. `PROVEN — EVIDENCE#043–047`.
- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, staleness, and signal-quality gating all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. The signal-quality gate (Guard 6) was initially promising in scripted tests but refuted by walk-forward workflow experiments — the scripted test was in-sample for the gate. `PROVEN — EVIDENCE#043–053`.
- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
@@ -105,44 +106,45 @@ Key observations:
3. **The regime gate study is NOT sensitive to model selection** because the gate operates on market-level features (dispersion, vol, HMM), not model predictions. Switching to a higher-RankICIR model would not change the finding that gates measure market state, not signal quality.
4. **Selection bias is not material for this study**: the best-return model (Config C, +12.5%) also has the best RankICIR (0.244) among walk-forward configs. The RankICIR and returns rankings are concordant.
## Signal-quality gate (Guard 7): the gate that works
## Signal-quality gate (Guard 7): refuted
`PROVEN — EVIDENCE#052`
`REFUTED — EVIDENCE#052 → EVIDENCE#053`
The regime gate (Guard 6) failed because it answered the wrong question: *"Is the market calm?"* The signal-quality gate asks the right question: *"Are my predictions accurate?"*
The regime gate (Guard 6) failed because it answered the wrong question: *"Is the market calm?"* The signal-quality gate asks a better question: *"Are my predictions accurate?"* — but when tested properly, it still doesn't work.
**Logic:** For each day t, look at the topk symbols from yesterday (t-1). Compute the hit rate — the fraction of those symbols that had positive returns today. If the hit rate is above a threshold, keep trading; otherwise, go to cash. This is a retrospective gate — it measures prediction accuracy, not market state.
**Results across 5 walk-forward windows (2021–2026):**
**Initial scripted test (EVIDENCE#052):** Precomputed gate from reference pred.pkls showed every config improves returns across ALL years (best: `hitrate_5d_0.50` 2026 +65.0%, 2025 +72.1%, 2024 +30.4%, 2023 +54.7%, 2021 +55.7%). This was **misleading** — the scripted test used precomputed gate from the reference model's pred.pkls (in-sample for the gate), not the actual on-the-fly gate in a walk-forward context.
| Gate | 2026 base | 2026 gated | 2025 base | 2025 gated | 2024 base | 2024 gated | 2023 base | 2023 gated | 2021 base | 2021 gated |
|------|-----------|------------|-----------|------------|-----------|------------|-----------|------------|-----------|------------|
| `hitrate_5d_0.50` | +25.5% | **+65.0%** | +17.8% | **+72.1%** | +8.2% | **+30.4%** | −4.8% | **+54.7%** | +18.4% | **+55.7%** |
| `hitrate_5d_0.60` | +25.5% | +48.9% | +17.8% | +48.6% | +8.2% | +26.5% | −4.8% | +55.6% | +18.4% | +35.2% |
| `hitrate_10d_0.50` | +25.5% | +33.6% | +17.8% | +43.6% | +8.2% | +27.7% | −4.8% | +46.0% | +18.4% | +46.5% |
| `hitrate_20d_0.50` | +25.5% | +34.4% | +17.8% | +34.1% | +8.2% | +21.1% | −4.8% | +36.6% | +18.4% | +29.0% |
**Walk-forward workflow test (EVIDENCE#053):** `WeeklyRebalanceSignalQualityGateStrategy` (topk=10, n_drop=1, gate_topk=10, gate_lookback=5, gate_threshold=0.5, 5/15bp costs) tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026), retraining the model each year. **The gate is harmful in every year:**
`PROVEN — EVIDENCE#052` (scripted simulation: `book/scripts/signal_quality_gate_bt.py`, results `book/data/signal_quality_gate/signal_quality_gate_results.csv`).
| Year | Workflow excess w/cost (gate) | Reference excess w/cost (nogate) | Delta |
|------|------------------------------|----------------------------------|-------|
| 2026 | +9.1% (IR 0.92) | +12.5% (IR 1.24) | **−3.4pp** |
| 2025 | +3.4% (IR 0.31) | +3.7% (IR 0.33) | **−0.3pp** |
| 2024 | −20.5% | −19.4% | **−1.1pp** |
| 2023 | −29.4% | −29.6% | +0.2pp |
| 2021 | −18.4% | −21.2% | +2.8pp |
Key observations:
**Why the scripted test was wrong:** The diagnostic (v3, workflow-exact mechanics) reveals the gate closes 37–45% of days in every year, killing returns:
1. **Every config improves returns across ALL years** — including the bad years (2023: −4.8% → +54.7%, 2024: +8.2% → +30.4%). The regime gate (Guard 6) destroyed returns in good years; the signal-quality gate improves them everywhere.
| Year | Script total (nogate) | Script total (gate) | Delta | Gate open% |
|------|----------------------|--------------------|-------|-----------|
| 2026 | +21.2% | +1.5% | −19.7pp | 56% |
| 2025 | +22.7% | +7.8% | −14.9pp | 63% |
| 2024 | +4.0% | −2.6% | −6.6pp | 59% |
2. **The gate trips ~40–50% of days** — it's closing on about half the days, filtering out the model's inaccurate predictions. This is the opposite of the regime gate, which closed on the wrong days.
A model with Rank IC 0.06–0.07 produces many days where <50% of top-10 picks are positive — the gate's 0.5 threshold is too aggressive, closing on profitable weeks. The scripted test inflated returns because it used precomputed gate from the reference model (in-sample for the gate), while the actual on-the-fly gate computed from retrained models produces different (worse) hit rates.
3. **The 5-day lookback with 0.50 threshold is optimal** — shorter lookbacks (5d) outperform longer ones (10d, 20d) because they adapt faster to changing prediction quality. The 0.50 threshold (random) is the sweet spot — it closes when the model is worse than random.
**Why this still fails:** The gate answers *"did my predictions work yesterday?"* — but with a 0.06–0.07 Rank IC, yesterday's hit rate is mostly noise. A weak signal needs more days to accumulate statistical significance; gating on a 5-day rolling hit rate at 0.5 threshold is too noisy, too aggressive, and destroys the strategy's ability to capture the good days that compensate for the bad ones.
4. **The gate is the OPPOSITE of the regime gate**: instead of closing on bad market days, it closes on days when the model's predictions are wrong. The model's predictions ARE informative; they just need to be gated on their own accuracy.
**Why this works:** The regime gate answered *"Is the market calm?"* — but calm markets can produce bad signals (low vol but wrong factor regime), and volatile markets can produce good signals (high vol but correct factor direction). The signal-quality gate answers *"Did my predictions work yesterday?"* — which directly predicts whether they'll work today.
**Caveat:** This is a retrospective gate — it uses yesterday's hit rate to decide today's trades. In real-time, you'd need to wait for today's close to compute the hit rate, then apply it to tomorrow's trades. The simulation uses yesterday's scores → today's returns (no look-ahead), so the gate is causal.
**Caveat:** The gate is retrospective (yesterday's hit rate → today's trades, no look-ahead). The problem is not look-ahead — it's that the signal is too weak for a 0.5 threshold on a 5-day window to be informative.
## Desk rules distilled from this chapter
1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel.
3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, staleness, and signal-quality gating all failed on this panel. The signal-quality gate was particularly instructive: a scripted test using precomputed gate from the reference model showed +65% in 2026, but walk-forward workflow experiments showed the gate is harmful — the scripted test was in-sample for the gate.
4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account.
5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.
@@ -185,8 +187,7 @@ The vol gates show the largest trip differential — they open on more days in 2
## Open questions
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
- `TODO(evidence-needed: signal-quality gate tested on out-of-sample data — the current test uses the same pred.pkl for gate computation and trading, which is in-sample for the gate itself)`
- `TODO(evidence-needed: signal-quality gate combined with the regime gate — does layering both gates improve results further?)`
- `TODO(evidence-needed: understanding the script-vs-workflow gap for signal-quality gate — scripted test shows gate destroying ~20pp more return than workflow, despite identical parameters; root cause is pred date alignment differences between precomputed and on-the-fly gate computation)`
## Evidence cited in this chapter
@@ -201,4 +202,5 @@ The vol gates show the largest trip differential — they open on more days in 2
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |
| `EVIDENCE#050` | Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation `book/scripts/regime_gate_bt.py`, results `book/data/regime_gate/regime_gate_trip_rates.csv` |
| `EVIDENCE#051` | Comprehensive model search: all experiments ranked by RankICIR; regime gate study robust to model selection; `rd_exp_list` + `rd_exp_get_run` queries |
| `EVIDENCE#052` | Signal-quality gate (hit-rate based on topk predictions): every config improves returns across ALL years; scripted simulation `book/scripts/signal_quality_gate_bt.py`, results `book/data/signal_quality_gate/signal_quality_gate_results.csv` |
| `EVIDENCE#052` | Signal-quality gate scripted test: precomputed gate from reference pred.pkls showed every config improves returns across ALL years. **REFUTED by EVIDENCE#053** — scripted test was in-sample for the gate. |
| `EVIDENCE#053` | Signal-quality gate walk-forward refutation: `WeeklyRebalanceSignalQualityGateStrategy` tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026). Gate harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp). Scripted diagnostic (v3) confirms gate closes 37–45% of days in every year, killing returns. |