ch11: signal-quality gate REFUTED — walk-forward workflow shows gate harmful (exp 61-67, EVIDENCE#053)
This commit is contained in:
@@ -48,8 +48,9 @@ With the walk-forward sweep showing the edge is 2026-window-specific, the desk t
|
||||
| 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. |
|
||||
| 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. |
|
||||
| 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. |
|
||||
| 6 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. The gate destroys the strategy's ability to capture the good days that compensate for the bad ones. |
|
||||
|
||||
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`.
|
||||
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`.
|
||||
|
||||
## The account-level truth
|
||||
|
||||
@@ -58,7 +59,7 @@ The blotter's daily `account` field is the authoritative measure (the `return` f
|
||||
## The synthesis
|
||||
|
||||
- **The headline results were window-specific.** Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. `PROVEN — EVIDENCE#043/044`.
|
||||
- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, and staleness all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. `PROVEN — EVIDENCE#043–047`.
|
||||
- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, staleness, and signal-quality gating all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. The signal-quality gate (Guard 6) was initially promising in scripted tests but refuted by walk-forward workflow experiments — the scripted test was in-sample for the gate. `PROVEN — EVIDENCE#043–053`.
|
||||
- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
|
||||
- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
|
||||
|
||||
@@ -105,44 +106,45 @@ Key observations:
|
||||
3. **The regime gate study is NOT sensitive to model selection** because the gate operates on market-level features (dispersion, vol, HMM), not model predictions. Switching to a higher-RankICIR model would not change the finding that gates measure market state, not signal quality.
|
||||
4. **Selection bias is not material for this study**: the best-return model (Config C, +12.5%) also has the best RankICIR (0.244) among walk-forward configs. The RankICIR and returns rankings are concordant.
|
||||
|
||||
## Signal-quality gate (Guard 7): the gate that works
|
||||
## Signal-quality gate (Guard 7): refuted
|
||||
|
||||
`PROVEN — EVIDENCE#052`
|
||||
`REFUTED — EVIDENCE#052 → EVIDENCE#053`
|
||||
|
||||
The regime gate (Guard 6) failed because it answered the wrong question: *"Is the market calm?"* The signal-quality gate asks the right question: *"Are my predictions accurate?"*
|
||||
The regime gate (Guard 6) failed because it answered the wrong question: *"Is the market calm?"* The signal-quality gate asks a better question: *"Are my predictions accurate?"* — but when tested properly, it still doesn't work.
|
||||
|
||||
**Logic:** For each day t, look at the topk symbols from yesterday (t-1). Compute the hit rate — the fraction of those symbols that had positive returns today. If the hit rate is above a threshold, keep trading; otherwise, go to cash. This is a retrospective gate — it measures prediction accuracy, not market state.
|
||||
|
||||
**Results across 5 walk-forward windows (2021–2026):**
|
||||
**Initial scripted test (EVIDENCE#052):** Precomputed gate from reference pred.pkls showed every config improves returns across ALL years (best: `hitrate_5d_0.50` 2026 +65.0%, 2025 +72.1%, 2024 +30.4%, 2023 +54.7%, 2021 +55.7%). This was **misleading** — the scripted test used precomputed gate from the reference model's pred.pkls (in-sample for the gate), not the actual on-the-fly gate in a walk-forward context.
|
||||
|
||||
| Gate | 2026 base | 2026 gated | 2025 base | 2025 gated | 2024 base | 2024 gated | 2023 base | 2023 gated | 2021 base | 2021 gated |
|
||||
|------|-----------|------------|-----------|------------|-----------|------------|-----------|------------|-----------|------------|
|
||||
| `hitrate_5d_0.50` | +25.5% | **+65.0%** | +17.8% | **+72.1%** | +8.2% | **+30.4%** | −4.8% | **+54.7%** | +18.4% | **+55.7%** |
|
||||
| `hitrate_5d_0.60` | +25.5% | +48.9% | +17.8% | +48.6% | +8.2% | +26.5% | −4.8% | +55.6% | +18.4% | +35.2% |
|
||||
| `hitrate_10d_0.50` | +25.5% | +33.6% | +17.8% | +43.6% | +8.2% | +27.7% | −4.8% | +46.0% | +18.4% | +46.5% |
|
||||
| `hitrate_20d_0.50` | +25.5% | +34.4% | +17.8% | +34.1% | +8.2% | +21.1% | −4.8% | +36.6% | +18.4% | +29.0% |
|
||||
**Walk-forward workflow test (EVIDENCE#053):** `WeeklyRebalanceSignalQualityGateStrategy` (topk=10, n_drop=1, gate_topk=10, gate_lookback=5, gate_threshold=0.5, 5/15bp costs) tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026), retraining the model each year. **The gate is harmful in every year:**
|
||||
|
||||
`PROVEN — EVIDENCE#052` (scripted simulation: `book/scripts/signal_quality_gate_bt.py`, results `book/data/signal_quality_gate/signal_quality_gate_results.csv`).
|
||||
| Year | Workflow excess w/cost (gate) | Reference excess w/cost (nogate) | Delta |
|
||||
|------|------------------------------|----------------------------------|-------|
|
||||
| 2026 | +9.1% (IR 0.92) | +12.5% (IR 1.24) | **−3.4pp** |
|
||||
| 2025 | +3.4% (IR 0.31) | +3.7% (IR 0.33) | **−0.3pp** |
|
||||
| 2024 | −20.5% | −19.4% | **−1.1pp** |
|
||||
| 2023 | −29.4% | −29.6% | +0.2pp |
|
||||
| 2021 | −18.4% | −21.2% | +2.8pp |
|
||||
|
||||
Key observations:
|
||||
**Why the scripted test was wrong:** The diagnostic (v3, workflow-exact mechanics) reveals the gate closes 37–45% of days in every year, killing returns:
|
||||
|
||||
1. **Every config improves returns across ALL years** — including the bad years (2023: −4.8% → +54.7%, 2024: +8.2% → +30.4%). The regime gate (Guard 6) destroyed returns in good years; the signal-quality gate improves them everywhere.
|
||||
| Year | Script total (nogate) | Script total (gate) | Delta | Gate open% |
|
||||
|------|----------------------|--------------------|-------|-----------|
|
||||
| 2026 | +21.2% | +1.5% | −19.7pp | 56% |
|
||||
| 2025 | +22.7% | +7.8% | −14.9pp | 63% |
|
||||
| 2024 | +4.0% | −2.6% | −6.6pp | 59% |
|
||||
|
||||
2. **The gate trips ~40–50% of days** — it's closing on about half the days, filtering out the model's inaccurate predictions. This is the opposite of the regime gate, which closed on the wrong days.
|
||||
A model with Rank IC 0.06–0.07 produces many days where <50% of top-10 picks are positive — the gate's 0.5 threshold is too aggressive, closing on profitable weeks. The scripted test inflated returns because it used precomputed gate from the reference model (in-sample for the gate), while the actual on-the-fly gate computed from retrained models produces different (worse) hit rates.
|
||||
|
||||
3. **The 5-day lookback with 0.50 threshold is optimal** — shorter lookbacks (5d) outperform longer ones (10d, 20d) because they adapt faster to changing prediction quality. The 0.50 threshold (random) is the sweet spot — it closes when the model is worse than random.
|
||||
**Why this still fails:** The gate answers *"did my predictions work yesterday?"* — but with a 0.06–0.07 Rank IC, yesterday's hit rate is mostly noise. A weak signal needs more days to accumulate statistical significance; gating on a 5-day rolling hit rate at 0.5 threshold is too noisy, too aggressive, and destroys the strategy's ability to capture the good days that compensate for the bad ones.
|
||||
|
||||
4. **The gate is the OPPOSITE of the regime gate**: instead of closing on bad market days, it closes on days when the model's predictions are wrong. The model's predictions ARE informative; they just need to be gated on their own accuracy.
|
||||
|
||||
**Why this works:** The regime gate answered *"Is the market calm?"* — but calm markets can produce bad signals (low vol but wrong factor regime), and volatile markets can produce good signals (high vol but correct factor direction). The signal-quality gate answers *"Did my predictions work yesterday?"* — which directly predicts whether they'll work today.
|
||||
|
||||
**Caveat:** This is a retrospective gate — it uses yesterday's hit rate to decide today's trades. In real-time, you'd need to wait for today's close to compute the hit rate, then apply it to tomorrow's trades. The simulation uses yesterday's scores → today's returns (no look-ahead), so the gate is causal.
|
||||
**Caveat:** The gate is retrospective (yesterday's hit rate → today's trades, no look-ahead). The problem is not look-ahead — it's that the signal is too weak for a 0.5 threshold on a 5-day window to be informative.
|
||||
|
||||
## Desk rules distilled from this chapter
|
||||
|
||||
1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
|
||||
2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
|
||||
3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, and staleness all failed on this panel.
|
||||
3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, staleness, and signal-quality gating all failed on this panel. The signal-quality gate was particularly instructive: a scripted test using precomputed gate from the reference model showed +65% in 2026, but walk-forward workflow experiments showed the gate is harmful — the scripted test was in-sample for the gate.
|
||||
4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account.
|
||||
5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.
|
||||
|
||||
@@ -185,8 +187,7 @@ The vol gates show the largest trip differential — they open on more days in 2
|
||||
## Open questions
|
||||
|
||||
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
|
||||
- `TODO(evidence-needed: signal-quality gate tested on out-of-sample data — the current test uses the same pred.pkl for gate computation and trading, which is in-sample for the gate itself)`
|
||||
- `TODO(evidence-needed: signal-quality gate combined with the regime gate — does layering both gates improve results further?)`
|
||||
- `TODO(evidence-needed: understanding the script-vs-workflow gap for signal-quality gate — scripted test shows gate destroying ~20pp more return than workflow, despite identical parameters; root cause is pred date alignment differences between precomputed and on-the-fly gate computation)`
|
||||
|
||||
## Evidence cited in this chapter
|
||||
|
||||
@@ -201,4 +202,5 @@ The vol gates show the largest trip differential — they open on more days in 2
|
||||
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |
|
||||
| `EVIDENCE#050` | Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation `book/scripts/regime_gate_bt.py`, results `book/data/regime_gate/regime_gate_trip_rates.csv` |
|
||||
| `EVIDENCE#051` | Comprehensive model search: all experiments ranked by RankICIR; regime gate study robust to model selection; `rd_exp_list` + `rd_exp_get_run` queries |
|
||||
| `EVIDENCE#052` | Signal-quality gate (hit-rate based on topk predictions): every config improves returns across ALL years; scripted simulation `book/scripts/signal_quality_gate_bt.py`, results `book/data/signal_quality_gate/signal_quality_gate_results.csv` |
|
||||
| `EVIDENCE#052` | Signal-quality gate scripted test: precomputed gate from reference pred.pkls showed every config improves returns across ALL years. **REFUTED by EVIDENCE#053** — scripted test was in-sample for the gate. |
|
||||
| `EVIDENCE#053` | Signal-quality gate walk-forward refutation: `WeeklyRebalanceSignalQualityGateStrategy` tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026). Gate harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp). Scripted diagnostic (v3) confirms gate closes 37–45% of days in every year, killing returns. |
|
||||
Reference in New Issue
Block a user