book: add perturbation stress test (EVIDENCE#049) — Config A 2026 robust to topk/n_drop/cost within window

This commit is contained in:
zhaoli
2026-08-20 21:44:04 +00:00
parent 18b61bbb8f
commit 4831a2770c
3 changed files with 58 additions and 1 deletions
+1
View File
@@ -77,6 +77,7 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
| ID | Claim | Source | Verified? |
|----|-------|--------|-----------|
| EVIDENCE#048 | Streaming IC circuit-breaker (`ic_min_rankic`, `ICGateTopkDropoutStrategy` in `tac_qlib/contrib/strategy/ic_gate.py`) trip-rate study: with thresholds 0.02–0.06, the gate trips on 25–50% of days in every year (2021–2026), freezing TopkDropout's rotation out of losers. A gate that trips every year cannot separate good years from bad. Do not deploy live. | ad-hoc scripted study on exp 52/53 pred/label artifacts, `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py`, `tac_qlib/tac_qlib/risk_limits.py` | yes — guard candidate 3 REFUTED |
| EVIDENCE#049 | Perturbation stress test on Config A 2026 (exp 52, pred from run `9f98ea5c`): same signal, varying topk (5/10/15), n_drop (1/2/3), costs (base/high/5×base). **topk**: 10 optimal (32.8% raw, Sharpe 1.98); 5 loses ~0.5pp, 15 loses ~6.5pp. **n_drop**: 1 optimal; 2 loses ~6pp, 3 loses ~4pp. **costs**: immaterial — 5× cost increase (25bp/35bp/$15) drops return only 0.17pp (32.84%→32.67%). maxDD stable −5.8% to −7.0% across all perturbations. **Within the 2026 window the edge is robust to parameter perturbation.** The problem remains that it does not exist in other windows (ch 11). | ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json` | yes — within-window robustness confirmed |
## External references (book/references/)
+20 -1
View File
@@ -62,6 +62,24 @@ The blotter's daily `account` field is the authoritative measure (the `return` f
- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
## Within-window robustness (perturbation stress test)
The 2026 edge is fragile *across* windows but robust *within* the 2026 window. A perturbation grid on Config A's predictions (exp 52, run `9f98ea5c`, same pred.pkl, varying only backtest parameters):
| Perturbation | Config | Ann. return | Sharpe | maxDD |
|-------------|--------|-------------|--------|-------|
| **Baseline** | topk=10, n_drop=1, costs 5/15/$5 | 32.8% | 1.98 | −5.8% |
| topk=5 | concentration ↑ | 32.3% | 1.75 | −7.0% |
| topk=15 | concentration ↓ | 26.3% | 1.61 | −6.9% |
| n_drop=2 | rotation ↑ | 26.7% | 1.59 | −6.5% |
| n_drop=3 | rotation ↑↑ | 28.6% | 1.66 | −6.7% |
| costs 3× (15/25/$10) | cost stress | 32.8% | 1.97 | −5.8% |
| costs 5× (25/35/$15) | cost stress ↑↑ | 32.7% | 1.97 | −5.8% |
`PROVEN — EVIDENCE#049` (ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json`).
Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. **Within the one good window, the edge is not a parameter-tuning artifact.** The fragility is entirely across windows (regime dependence), not within them.
## Desk rules distilled from this chapter
1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
@@ -84,4 +102,5 @@ The blotter's daily `account` field is the authoritative measure (the `return` f
| `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) |
| `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` |
| `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` |
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |
@@ -0,0 +1,37 @@
{
"description": "Perturbation stress test on Config A (exp 52, run 9f98ea5c) — 2026 window (2026-01-04 to 2026-08-19). Same pred.pkl, varying backtest parameters. Raw returns (not excess over SPY).",
"pred_source": "exp 52, run 9f98ea5c550a409f87b56a6cd8fee343",
"test_window": ["2026-01-04", "2026-08-19"],
"trading_days": 157,
"topk_sensitivity": {
"description": "vary topk, n_drop=1, costs=base (5bp/15bp/$5)",
"results": [
{"topk": 5, "n_drop": 1, "ann_return": 0.3230, "sharpe": 1.745, "maxDD": -0.0695},
{"topk": 10, "n_drop": 1, "ann_return": 0.3284, "sharpe": 1.978, "maxDD": -0.0581},
{"topk": 15, "n_drop": 1, "ann_return": 0.2630, "sharpe": 1.611, "maxDD": -0.0689}
]
},
"ndrop_sensitivity": {
"description": "vary n_drop, topk=10, costs=base",
"results": [
{"topk": 10, "n_drop": 1, "ann_return": 0.3284, "sharpe": 1.978, "maxDD": -0.0581},
{"topk": 10, "n_drop": 2, "ann_return": 0.2674, "sharpe": 1.586, "maxDD": -0.0652},
{"topk": 10, "n_drop": 3, "ann_return": 0.2861, "sharpe": 1.658, "maxDD": -0.0667}
]
},
"cost_sensitivity": {
"description": "vary costs, topk=10, n_drop=1",
"results": [
{"open_cost": 0.0005, "close_cost": 0.0015, "min_cost": 5, "ann_return": 0.3284, "sharpe": 1.978, "maxDD": -0.0581},
{"open_cost": 0.0015, "close_cost": 0.0025, "min_cost": 10, "ann_return": 0.3275, "sharpe": 1.974, "maxDD": -0.0580},
{"open_cost": 0.0025, "close_cost": 0.0035, "min_cost": 15, "ann_return": 0.3267, "sharpe": 1.971, "maxDD": -0.0580}
]
},
"findings": {
"topk": "topk=10 is optimal. topk=5 loses ~0.5pp (concentration risk), topk=15 loses ~6.5pp (signal dilution). Edge is moderate-sensitivity to topk.",
"n_drop": "n_drop=1 is best. n_drop=2 loses ~6pp, n_drop=3 loses ~4pp. More rotation hurts in this window.",
"costs": "Almost irrelevant. Even at 5x base costs (25bp/35bp/$15), return drops only 0.17pp (32.84% → 32.67%). Low turnover + large gross edge makes cost assumptions immaterial.",
"maxDD": "Stable across all perturbations: range -5.8% to -7.0%. No blowup risk from parameter changes.",
"overall": "The 2026 edge is robust WITHIN the window. The problem is it doesn't exist in other windows (ch 11)."
}
}