207 lines
23 KiB
Markdown
207 lines
23 KiB
Markdown
# Chapter 11 — Walk-Forward Re-validation and Guard Candidates: The Edge Is a Regime Artifact
|
||
|
||
Status: drafting. Claim inventory: see `README.md` ch. 11.
|
||
|
||
This chapter answers the question every desk must ask before shipping a backtest result: **does the edge survive re-training on a different window?** The TradeAC campaign's headline results — weekly rebalance (exp 39, Q07), realized-moments features (exp 48, Q17), and the m2-sharpe22 reference (exp 33, Q01) — were all measured on a single 2026 window. This chapter re-runs them walk-forward across 2024–2026 (and 2021/2023 for the label-regime matches), then tests seven guard candidates that would plausibly have isolated the good years. **Every one of them is refuted.** The edge is a 2025–2026 regime artifact; no pre-deployment measurable gate selects it.
|
||
|
||
## The walk-forward re-validation (exp 52, 53, 54)
|
||
|
||
Three independent walk-forward sweeps, all on the clean lake, all with the same 5-seed RankIC ensemble (`42,7,2026,99,123`, 5d label, SPY benchmark, 5bp/15bp/$5 costs, 50-ETF panel):
|
||
|
||
### 3×3 sweep (exp 52) — the three best configs across 2024/2025/2026
|
||
|
||
Nine runs (3 configs × 3 windows), train/valid shifted per window to avoid overlap. Config A = weekly-rebalance TopkDropout topk10/n_drop1 (exp 39, Q07); Config B = TopkDropout topk10/n_drop1 with realized-moments features (exp 48, Q17); Config C = TopkDropout topk10/n_drop2 base features (exp 26 reference). Excess = annualized return over SPY, net of cost.
|
||
|
||
| Config | 2024 (test) | 2025 (test) | 2026 (test) |
|
||
|--------|-------------|-------------|-------------|
|
||
| **A** weekly, n_drop=1 | −18.1% (IR −1.39, maxDD −22.7%) | −4.2% (IR −0.52, maxDD −10.3%) | **+12.5% (IR 1.25, maxDD −4.1%)** |
|
||
| **B** moments, n_drop=1 | −16.1% (IR −1.91, maxDD −20.0%) | −8.9% (IR −1.00, maxDD −9.8%) | **+9.2% (IR 0.94, maxDD −6.5%)** |
|
||
| **C** base, n_drop=2 | −18.2% (IR −1.97, maxDD −23.3%) | −3.6% (IR −0.51, maxDD −6.3%) | **−1.4% (IR −0.13, maxDD −9.8%)** |
|
||
|
||
`PROVEN — EVIDENCE#043 → exp 52`. Read the table carefully:
|
||
|
||
- **The 2026 window is the only profitable one**, and only for A (+12.5%) and B (+9.2%); C goes negative even in 2026.
|
||
- **A and C share identical predictions** — same model, same features, byte-identical IC/RankIC in every window (e.g. 2026 IC 0.0494, RankIC 0.0637 for both). The strategy layer alone (weekly recompute vs daily n_drop2) differentiates the outcome. This is the cleanest possible demonstration that construction, not signal, separated A from C in 2026.
|
||
- **The harness is reproducible**: run 2 (Config A, 2025) exactly replicated exp 45 (`e5ac7a5d`, net −4.21%, IR −0.52) and Config A's 2026 run replicated exp 39 (`eb38588c`, +12.51%, IR 1.24).
|
||
|
||
### m2-sharpe22 3-window (exp 53)
|
||
|
||
The m2-sharpe22 reference (exp 33/Q01, `c7c12228`) re-run on the same 2024/2025/2026 scheme: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — same edge, same window, same config. `PROVEN — EVIDENCE#044 → exp 53`. The edge is recent-window-only.
|
||
|
||
### Label-regime transfer (exp 54) — 2021 and 2023
|
||
|
||
The 2025 feature-drift check had already shown a naive feature-PSI gate does **not** predict walk-forward performance — 2026 has the highest feature drift yet the best result (the model consumes CSRankNorm'd ranks, so raw feature drift is scale-invariant noise). Exp 54 instead tested the *label/return regime*: high cross-sectional 5d-label dispersion → good ranking year (2026 disp 0.0302, +6.5%); fat right tail / high skew → topk blowup (2024 skew +29, −26.4%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two untested closest matches were run:
|
||
|
||
- **2023**: −26.0% (IR −2.04, maxDD −30.8%)
|
||
- **2021**: −22.9% (IR −2.26, maxDD −27.0%)
|
||
|
||
`PROVEN — EVIDENCE#045 → exp 54`. Both closest label-regime matches are as bad as the 2024 tail. **No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year.** 2023 had decent dispersion but negative skew (−4.7) and still lost 26%; label dispersion alone does not protect against blowups.
|
||
|
||
## The seven guard candidates — all refuted
|
||
|
||
With the walk-forward sweep showing the edge is 2026-window-specific, the desk tested seven guards that could plausibly have preserved the good years and cut the bad ones. All seven were pre-registered as hypotheses (traced experiments), all seven failed:
|
||
|
||
| # | Guard | Test | Result |
|
||
|---|-------|------|--------|
|
||
| 1 | **Feature-PSI gate** | halt when the live feature distribution drifts from the training distribution (exp 52/53 feature-drift study) | REFUTED — 2026 has the *highest* drift yet the *best* result; CSRankNorm'd ranks make raw drift scale-invariant. |
|
||
| 2 | **Label-regime gate** | trade only when the live label regime matches the profitable 2026 regime (PSI on 5d-label dispersion/skew/vol) | REFUTED — closest matches (2023, 2021) both ≈ −26%/−23%; 2023 had decent dispersion and still blew up. |
|
||
| 3 | **Streaming IC circuit breaker** (`ic_min_rankic`) | pause new buys while trailing realized RankIC (computed causally from lake bars) is below a threshold | REFUTED — trips 25–50% of days *every year*, freezing TopkDropout's rotation out of losers; implemented in `tac_qlib/contrib/strategy/ic_gate.py`, do not deploy live. |
|
||
| 4 | **Adaptive short-window retrain** (exp 55) | retrain on rolling 1y/2y windows instead of the growing 2016→prev-Aug window | REFUTED — 1y and 2y put **every** test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for *every* window length. |
|
||
| 5 | **Window-staleness isolation** (exp 56) | gate on days-since-training-cutoff; the hypothesis was that the edge concentrates in fresh (low-staleness) predictions and bad years bleed when the model is stale | REFUTED — pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in **every** bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%): the *freshest* bucket is the *most* negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug; Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). Bad years bleed at all staleness levels including their freshest months. No staleness threshold isolates the edge. |
|
||
| 6 | **Regime gate** (dispersion / vol / HMM) | daily boolean gate based on market state (low vol, HMM posterior, dispersion) | REFUTED — see below; regime gate closes on the wrong days (hmm_0.7 destroys 2026: +25.5% → +10.8%) |
|
||
| 7 | **Signal-quality gate** (hit-rate) | gate on whether the model's recent topk predictions were correct (5-day rolling hit rate > 0.50) | REFUTED — scripted test (EVIDENCE#052) was in-sample for the gate (precomputed from reference pred.pkls); walk-forward workflow tests (EVIDENCE#053) show the gate is harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp), 2024 −20.5% vs −19.4% (−1.1pp). A model with Rank IC 0.06–0.07 produces too many days where <50% of top-10 picks are positive — the 0.5 threshold is too aggressive, closing on profitable weeks. |
|
||
|
||
Guards 1–3 are documented across exp 52/53/54 and the `ic_gate.py` implementation; guard 4 = `PROVEN (refuted) — EVIDENCE#046 → exp 55`; guard 5 = `PROVEN (refuted) — EVIDENCE#047 → exp 56`; guard 6 = `PROVEN (refuted) — EVIDENCE#050`; guard 7 = `REFUTED — EVIDENCE#052 → EVIDENCE#053`.
|
||
|
||
## The account-level truth
|
||
|
||
The blotter's daily `account` field is the authoritative measure (the `return` field excludes initial cost and does not compound to the final account). Cumulative excess vs SPY, account-based: **2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38%**. This reconciles with the recorded metrics — 2026 `excess_return_with_cost` annualized +6.5% (IR 0.62; without cost +11.4%, IR 1.09) — the same sign and order of magnitude on a shorter window. `PROVEN — EVIDENCE#047 → exp 56` (account curves from the exp 53/54 runs' blotter artifacts).
|
||
|
||
## The synthesis
|
||
|
||
- **The headline results were window-specific.** Weekly rebalance (+12.51%, IR 1.24) and m2-sharpe22 (+6.5%, IR 0.62) are 2026-only. Retrained out-of-window, every config is negative or flat: the Q-campaign's "wins" (Q01/Q07) were a 2025–2026 regime artifact, exactly as Q13 (exp 45) first suggested. `PROVEN — EVIDENCE#043/044`.
|
||
- **No guard candidate recovers the edge out-of-sample.** Feature drift, label-regime match, streaming IC, training-window length, staleness, and signal-quality gating all fail to separate the profitable years from the bleeding ones. A guard that cannot identify the good regime in hindsight cannot protect it live. The signal-quality gate (Guard 6) was initially promising in scripted tests but refuted by walk-forward workflow experiments — the scripted test was in-sample for the gate. `PROVEN — EVIDENCE#043–053`.
|
||
- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
|
||
- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
|
||
|
||
## Within-window robustness (perturbation stress test)
|
||
|
||
The 2026 edge is fragile *across* windows but robust *within* the 2026 window. A perturbation grid on Config A's predictions (exp 52, run `9f98ea5c`, same pred.pkl, varying only backtest parameters):
|
||
|
||
| Perturbation | Config | Ann. return | Sharpe | maxDD |
|
||
|-------------|--------|-------------|--------|-------|
|
||
| **Baseline** | topk=10, n_drop=1, costs 5/15/$5 | 32.8% | 1.98 | −5.8% |
|
||
| topk=5 | concentration ↑ | 32.3% | 1.75 | −7.0% |
|
||
| topk=15 | concentration ↓ | 26.3% | 1.61 | −6.9% |
|
||
| n_drop=2 | rotation ↑ | 26.7% | 1.59 | −6.5% |
|
||
| n_drop=3 | rotation ↑↑ | 28.6% | 1.66 | −6.7% |
|
||
| costs 3× (15/25/$10) | cost stress | 32.8% | 1.97 | −5.8% |
|
||
| costs 5× (25/35/$15) | cost stress ↑↑ | 32.7% | 1.97 | −5.8% |
|
||
|
||
`PROVEN — EVIDENCE#049` (ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json`).
|
||
|
||
Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. **Within the one good window, the edge is not a parameter-tuning artifact.** The fragility is entirely across windows (regime dependence), not within them.
|
||
|
||
## Model search & robustness of the regime gate finding
|
||
|
||
The regime gate study (Guard 7, EVIDENCE#050) used pred.pkl files from exp 52 (Configs A/C, 2024–2026) and exp 56 (2021, 2023). A comprehensive query of all MLflow experiments confirms the regime gate finding is robust to model selection:
|
||
|
||
| Rank | Exp | Test Window | RankICIR | Net Return | Notes |
|
||
|------|-----|-------------|----------|------------|-------|
|
||
| 1 | 36 | 2026 only | **0.507** | −4.6% | 22d label, single-window |
|
||
| 2 | 44 | 2026 only | **0.507** | −4.9% | Same pred as #1 |
|
||
| 3 | 35 | 2026 only | **0.352** | −9.9% | 10d label |
|
||
| 4 | 51 | 2026 only | **0.352** | +1.2% | Same pred as #3 |
|
||
| 5 | 58 | 2025 only | **0.289** | −3.4% | Adaptive 2y |
|
||
| 6 | 11 | 2026 only | **0.276** | +3.1% | Single seed |
|
||
| 7 | 33 | 2026 only | **0.259** | −0.9% | 10 seeds |
|
||
| 8 | **52-C** | **2024–2026** | **0.244** | **+12.5%** | Walk-forward, weekly |
|
||
| 9 | **52-A** | **2024–2026** | **0.244** | **−1.4%** | Walk-forward, TopkDrop |
|
||
|
||
`PROVEN — EVIDENCE#051` (comprehensive `rd_exp_list` query, run metadata).
|
||
|
||
Key observations:
|
||
|
||
1. **The 22-day label models (exp 36/44) have the highest RankICIR (0.507) but negative returns** — high IC does not guarantee profitable trading. The 22d label predicts longer-horizon moves that don't translate to short-term alpha after costs.
|
||
2. **Most high-RankICIR models are single-window (2026 only)** — they lack the multi-year coverage needed for the regime gate study. Walk-forward coverage (2021–2026) is limited to exp 52 (2024–2026) and exp 56 (2021, 2023).
|
||
3. **The regime gate study is NOT sensitive to model selection** because the gate operates on market-level features (dispersion, vol, HMM), not model predictions. Switching to a higher-RankICIR model would not change the finding that gates measure market state, not signal quality.
|
||
4. **Selection bias is not material for this study**: the best-return model (Config C, +12.5%) also has the best RankICIR (0.244) among walk-forward configs. The RankICIR and returns rankings are concordant.
|
||
|
||
## Signal-quality gate (Guard 7): refuted
|
||
|
||
`REFUTED — EVIDENCE#052 → EVIDENCE#053`
|
||
|
||
The regime gate (Guard 6) failed because it answered the wrong question: *"Is the market calm?"* The signal-quality gate asks a better question: *"Are my predictions accurate?"* — but when tested properly, it still doesn't work.
|
||
|
||
**Logic:** For each day t, look at the topk symbols from yesterday (t-1). Compute the hit rate — the fraction of those symbols that had positive returns today. If the hit rate is above a threshold, keep trading; otherwise, go to cash. This is a retrospective gate — it measures prediction accuracy, not market state.
|
||
|
||
**Initial scripted test (EVIDENCE#052):** Precomputed gate from reference pred.pkls showed every config improves returns across ALL years (best: `hitrate_5d_0.50` 2026 +65.0%, 2025 +72.1%, 2024 +30.4%, 2023 +54.7%, 2021 +55.7%). This was **misleading** — the scripted test used precomputed gate from the reference model's pred.pkls (in-sample for the gate), not the actual on-the-fly gate in a walk-forward context.
|
||
|
||
**Walk-forward workflow test (EVIDENCE#053):** `WeeklyRebalanceSignalQualityGateStrategy` (topk=10, n_drop=1, gate_topk=10, gate_lookback=5, gate_threshold=0.5, 5/15bp costs) tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026), retraining the model each year. **The gate is harmful in every year:**
|
||
|
||
| Year | Workflow excess w/cost (gate) | Reference excess w/cost (nogate) | Delta |
|
||
|------|------------------------------|----------------------------------|-------|
|
||
| 2026 | +9.1% (IR 0.92) | +12.5% (IR 1.24) | **−3.4pp** |
|
||
| 2025 | +3.4% (IR 0.31) | +3.7% (IR 0.33) | **−0.3pp** |
|
||
| 2024 | −20.5% | −19.4% | **−1.1pp** |
|
||
| 2023 | −29.4% | −29.6% | +0.2pp |
|
||
| 2021 | −18.4% | −21.2% | +2.8pp |
|
||
|
||
**Why the scripted test was wrong:** The diagnostic (v3, workflow-exact mechanics) reveals the gate closes 37–45% of days in every year, killing returns:
|
||
|
||
| Year | Script total (nogate) | Script total (gate) | Delta | Gate open% |
|
||
|------|----------------------|--------------------|-------|-----------|
|
||
| 2026 | +21.2% | +1.5% | −19.7pp | 56% |
|
||
| 2025 | +22.7% | +7.8% | −14.9pp | 63% |
|
||
| 2024 | +4.0% | −2.6% | −6.6pp | 59% |
|
||
|
||
A model with Rank IC 0.06–0.07 produces many days where <50% of top-10 picks are positive — the gate's 0.5 threshold is too aggressive, closing on profitable weeks. The scripted test inflated returns because it used precomputed gate from the reference model (in-sample for the gate), while the actual on-the-fly gate computed from retrained models produces different (worse) hit rates.
|
||
|
||
**Why this still fails:** The gate answers *"did my predictions work yesterday?"* — but with a 0.06–0.07 Rank IC, yesterday's hit rate is mostly noise. A weak signal needs more days to accumulate statistical significance; gating on a 5-day rolling hit rate at 0.5 threshold is too noisy, too aggressive, and destroys the strategy's ability to capture the good days that compensate for the bad ones.
|
||
|
||
**Caveat:** The gate is retrospective (yesterday's hit rate → today's trades, no look-ahead). The problem is not look-ahead — it's that the signal is too weak for a 0.5 threshold on a 5-day window to be informative.
|
||
|
||
## Desk rules distilled from this chapter
|
||
|
||
1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
|
||
2. Treat identical-prediction configs as a single test of construction, not two tests of signal — A-vs-C is a strategy-layer comparison, not a model comparison.
|
||
3. Do not ship a guard that cannot select the good regime in hindsight. Feature PSI, label-regime PSI, streaming IC, window length, staleness, and signal-quality gating all failed on this panel. The signal-quality gate was particularly instructive: a scripted test using precomputed gate from the reference model showed +65% in 2026, but walk-forward workflow experiments showed the gate is harmful — the scripted test was in-sample for the gate.
|
||
4. Report account-based curves, not the blotter `return` field — the latter excludes initial cost and does not compound to the account.
|
||
5. When the mean annual excess is negative in every configuration, cut size until the live window demonstrates the regime is back.
|
||
|
||
### Guard 6: Regime gate (dispersion / vol / HMM)
|
||
|
||
`PROVEN — EVIDENCE#050`
|
||
|
||
If the edge is regime-dependent, the most direct guard is a regime detector that opens on good years and closes on bad years. We test three detector types, each producing a daily boolean (trade / don't trade):
|
||
|
||
| Detector | Logic |
|
||
|----------|-------|
|
||
| **dispersion** | CS std of 22-day rolling returns < threshold (low dispersion → calm market → trade) |
|
||
| **vol** | CS mean of 22-day rolling realized vol within a band (mid-range vol → trade) |
|
||
| **HMM** | 2-state Gaussian HMM posterior for regime 1 (productive regime) > threshold |
|
||
|
||
Each detector is applied as a daily gate on top of the weekly-rebalance TopkDropout (topk=10, n_drop=1, yesterday's scores). We run 14 configs across 5 walk-forward windows (2021–2026), tracking trip rate (fraction of days gate is open) and gated return.
|
||
|
||
**Trip rates (2026 vs bad years 2021/2023/2024):**
|
||
|
||
| Gate | 2026 trip | Bad-years avg | Differential |
|
||
|------|-----------|---------------|-------------|
|
||
| `vol_low_max20` | 92% | 60% | +32pp |
|
||
| `vol_low_max25` | 63% | 16% | +47pp |
|
||
| `hmm_0.7` | 37% | 31% | +6pp |
|
||
| All dispersion | 0% | 0% | 0pp |
|
||
|
||
The vol gates show the largest trip differential — they open on more days in 2026 than in bad years. But the gate **closes on the wrong days**: when the gate is open only 63% of the time (vol_low_max25), the 2026 return collapses from +25.5% to −1.6%. The gate eliminates the profitable days along with the bad ones.
|
||
|
||
**Gated returns:**
|
||
|
||
| Gate | 2026 base | 2026 gated | 2023 base | 2023 gated | 2025 base | 2025 gated |
|
||
|------|-----------|------------|-----------|------------|-----------|------------|
|
||
| `vol_low_max20` | +25.5% | +4.3% | −4.8% | −5.3% | +17.8% | +14.5% |
|
||
| `hmm_0.7` | +25.5% | +10.8% | −4.8% | +0.6% | +17.8% | +26.8% |
|
||
|
||
`hmm_0.7` has the most interesting profile: it **improves** 2023 (−4.8% → +0.6%) and 2025 (+17.8% → +26.8%), but **destroys** 2026 (+25.5% → +10.8%). The gate's Sharpe is inflated (1.78 in 2021) because it spends most of its time in cash — the Sharpe measures "active days only" and ignores the flat periods.
|
||
|
||
**Why none of these gates work:** The gate answers *"is the market calm right now?"* — but the right question is *"will today's signal be profitable tomorrow?"* These are different questions. A calm market can produce bad signals (low vol but wrong factor regime), and a volatile market can produce good signals (high vol but correct factor direction). The gate needs to predict **signal quality**, not **market state**. See Guard 7 (signal-quality gate, EVIDENCE#053) for a gate that was tested on this principle — and still failed.
|
||
|
||
## Open questions
|
||
|
||
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
|
||
- `TODO(evidence-needed: understanding the script-vs-workflow gap for signal-quality gate — scripted test shows gate destroying ~20pp more return than workflow, despite identical parameters; root cause is pred date alignment differences between precomputed and on-the-fly gate computation)`
|
||
|
||
## Evidence cited in this chapter
|
||
|
||
| Tag | Source |
|
||
|-----|--------|
|
||
| `EVIDENCE#043` | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` |
|
||
| `EVIDENCE#044` | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference run `c7c12228` (exp 33, Q01) |
|
||
| `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) |
|
||
| `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` |
|
||
| `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` |
|
||
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
|
||
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |
|
||
| `EVIDENCE#050` | Regime gate walk-forward test (2021–2026): 3 detector types × 14 configs; scripted simulation `book/scripts/regime_gate_bt.py`, results `book/data/regime_gate/regime_gate_trip_rates.csv` |
|
||
| `EVIDENCE#051` | Comprehensive model search: all experiments ranked by RankICIR; regime gate study robust to model selection; `rd_exp_list` + `rd_exp_get_run` queries |
|
||
| `EVIDENCE#052` | Signal-quality gate scripted test: precomputed gate from reference pred.pkls showed every config improves returns across ALL years. **REFUTED by EVIDENCE#053** — scripted test was in-sample for the gate. |
|
||
| `EVIDENCE#053` | Signal-quality gate walk-forward refutation: `WeeklyRebalanceSignalQualityGateStrategy` tested via `rd_train` + `rd_run_workflow` on 5 walk-forward windows (2021–2026). Gate harmful in every year: 2026 +9.1% vs +12.5% reference (−3.4pp), 2025 +3.4% vs +3.7% (−0.3pp). Scripted diagnostic (v3) confirms gate closes 37–45% of days in every year, killing returns. | |