book: add perturbation stress test (EVIDENCE#049) — Config A 2026 robust to topk/n_drop/cost within window

This commit is contained in:
zhaoli
2026-08-20 21:44:04 +00:00
parent 18b61bbb8f
commit 4831a2770c
3 changed files with 58 additions and 1 deletions
+20 -1
View File
@@ -62,6 +62,24 @@ The blotter's daily `account` field is the authoritative measure (the `return` f
- **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024.
- **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe.
## Within-window robustness (perturbation stress test)
The 2026 edge is fragile *across* windows but robust *within* the 2026 window. A perturbation grid on Config A's predictions (exp 52, run `9f98ea5c`, same pred.pkl, varying only backtest parameters):
| Perturbation | Config | Ann. return | Sharpe | maxDD |
|-------------|--------|-------------|--------|-------|
| **Baseline** | topk=10, n_drop=1, costs 5/15/$5 | 32.8% | 1.98 | −5.8% |
| topk=5 | concentration ↑ | 32.3% | 1.75 | −7.0% |
| topk=15 | concentration ↓ | 26.3% | 1.61 | −6.9% |
| n_drop=2 | rotation ↑ | 26.7% | 1.59 | −6.5% |
| n_drop=3 | rotation ↑↑ | 28.6% | 1.66 | −6.7% |
| costs 3× (15/25/$10) | cost stress | 32.8% | 1.97 | −5.8% |
| costs 5× (25/35/$15) | cost stress ↑↑ | 32.7% | 1.97 | −5.8% |
`PROVEN — EVIDENCE#049` (ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json`).
Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. **Within the one good window, the edge is not a parameter-tuning artifact.** The fragility is entirely across windows (regime dependence), not within them.
## Desk rules distilled from this chapter
1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy.
@@ -84,4 +102,5 @@ The blotter's daily `account` field is the authoritative measure (the `return` f
| `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) |
| `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` |
| `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` |
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds |
| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` |