From 4831a2770c1b7eeb984d1c160c1068adf0224ccd Mon Sep 17 00:00:00 2001 From: zhaoli Date: Thu, 20 Aug 2026 21:44:04 +0000 Subject: [PATCH] =?UTF-8?q?book:=20add=20perturbation=20stress=20test=20(E?= =?UTF-8?q?VIDENCE#049)=20=E2=80=94=20Config=20A=202026=20robust=20to=20to?= =?UTF-8?q?pk/n=5Fdrop/cost=20within=20window?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- book/EVIDENCE.md | 1 + book/chapters/11-walk-forward-and-guards.md | 21 ++++++++++- .../config_a_2026_sensitivity.json | 37 +++++++++++++++++++ 3 files changed, 58 insertions(+), 1 deletion(-) create mode 100644 book/data/perturbation/config_a_2026_sensitivity.json diff --git a/book/EVIDENCE.md b/book/EVIDENCE.md index 2514b1f..8f09d7a 100644 --- a/book/EVIDENCE.md +++ b/book/EVIDENCE.md @@ -77,6 +77,7 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit | ID | Claim | Source | Verified? | |----|-------|--------|-----------| | EVIDENCE#048 | Streaming IC circuit-breaker (`ic_min_rankic`, `ICGateTopkDropoutStrategy` in `tac_qlib/contrib/strategy/ic_gate.py`) trip-rate study: with thresholds 0.02–0.06, the gate trips on 25–50% of days in every year (2021–2026), freezing TopkDropout's rotation out of losers. A gate that trips every year cannot separate good years from bad. Do not deploy live. | ad-hoc scripted study on exp 52/53 pred/label artifacts, `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py`, `tac_qlib/tac_qlib/risk_limits.py` | yes — guard candidate 3 REFUTED | +| EVIDENCE#049 | Perturbation stress test on Config A 2026 (exp 52, pred from run `9f98ea5c`): same signal, varying topk (5/10/15), n_drop (1/2/3), costs (base/high/5×base). **topk**: 10 optimal (32.8% raw, Sharpe 1.98); 5 loses ~0.5pp, 15 loses ~6.5pp. **n_drop**: 1 optimal; 2 loses ~6pp, 3 loses ~4pp. **costs**: immaterial — 5× cost increase (25bp/35bp/$15) drops return only 0.17pp (32.84%→32.67%). maxDD stable −5.8% to −7.0% across all perturbations. **Within the 2026 window the edge is robust to parameter perturbation.** The problem remains that it does not exist in other windows (ch 11). | ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json` | yes — within-window robustness confirmed | ## External references (book/references/) diff --git a/book/chapters/11-walk-forward-and-guards.md b/book/chapters/11-walk-forward-and-guards.md index 7b8bf70..a40b536 100644 --- a/book/chapters/11-walk-forward-and-guards.md +++ b/book/chapters/11-walk-forward-and-guards.md @@ -62,6 +62,24 @@ The blotter's daily `account` field is the authoritative measure (the `return` f - **Construction still matters inside the good regime.** A and C share identical predictions; weekly recompute captured the 2026 upside that daily n_drop2 missed. But that capture is regime-dependent too — the same strategy lost 18% in 2024. - **Live implication:** size for the mean, not the tail. The mean annual excess across every window length is ≈ −13%. Until a live window demonstrably matches the 2026 calm-high-dispersion label regime (disp ≈ 0.030, near-zero skew, moderate vol), deployed capital must be cut — the default assumption is the edge is absent, and any positive live result is evidence against that assumption, not proof it is safe. +## Within-window robustness (perturbation stress test) + +The 2026 edge is fragile *across* windows but robust *within* the 2026 window. A perturbation grid on Config A's predictions (exp 52, run `9f98ea5c`, same pred.pkl, varying only backtest parameters): + +| Perturbation | Config | Ann. return | Sharpe | maxDD | +|-------------|--------|-------------|--------|-------| +| **Baseline** | topk=10, n_drop=1, costs 5/15/$5 | 32.8% | 1.98 | −5.8% | +| topk=5 | concentration ↑ | 32.3% | 1.75 | −7.0% | +| topk=15 | concentration ↓ | 26.3% | 1.61 | −6.9% | +| n_drop=2 | rotation ↑ | 26.7% | 1.59 | −6.5% | +| n_drop=3 | rotation ↑↑ | 28.6% | 1.66 | −6.7% | +| costs 3× (15/25/$10) | cost stress | 32.8% | 1.97 | −5.8% | +| costs 5× (25/35/$15) | cost stress ↑↑ | 32.7% | 1.97 | −5.8% | + +`PROVEN — EVIDENCE#049` (ad-hoc rd_backtest grid on exp 52 pred.pkl, `book/data/perturbation/config_a_2026_sensitivity.json`). + +Key takeaways: topk=10 is the sweet spot (topk=15 dilutes the signal by ~6.5pp). n_drop=1 is best; more rotation hurts. Costs are almost immaterial — even 5× base costs drop return by only 0.17pp, because the strategy is low-turnover and the gross edge is large. maxDD is stable (−5.8% to −7.0%) across all perturbations. **Within the one good window, the edge is not a parameter-tuning artifact.** The fragility is entirely across windows (regime dependence), not within them. + ## Desk rules distilled from this chapter 1. Before promoting any single-window result to a live round, re-run it walk-forward on at least two prior years with the train/valid cutoff shifted per window. If the edge does not survive, it is a regime artifact, not a strategy. @@ -84,4 +102,5 @@ The blotter's daily `account` field is the authoritative measure (the `return` f | `EVIDENCE#045` | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) | | `EVIDENCE#046` | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` | | `EVIDENCE#047` | exp 56, staleness analysis on the exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` | -| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds | \ No newline at end of file +| Guard 3 (`ic_min_rankic`) | `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py` (ICGateTopkDropoutStrategy), `tac_qlib/tac_qlib/risk_limits.py`; trip-rate study on exp 52/53 preds | +| `EVIDENCE#049` | Perturbation stress test on Config A 2026 (exp 52, pred `9f98ea5c`): topk/n_drop/cost grid, `book/data/perturbation/config_a_2026_sensitivity.json` | \ No newline at end of file diff --git a/book/data/perturbation/config_a_2026_sensitivity.json b/book/data/perturbation/config_a_2026_sensitivity.json new file mode 100644 index 0000000..a36bf89 --- /dev/null +++ b/book/data/perturbation/config_a_2026_sensitivity.json @@ -0,0 +1,37 @@ +{ + "description": "Perturbation stress test on Config A (exp 52, run 9f98ea5c) — 2026 window (2026-01-04 to 2026-08-19). Same pred.pkl, varying backtest parameters. Raw returns (not excess over SPY).", + "pred_source": "exp 52, run 9f98ea5c550a409f87b56a6cd8fee343", + "test_window": ["2026-01-04", "2026-08-19"], + "trading_days": 157, + "topk_sensitivity": { + "description": "vary topk, n_drop=1, costs=base (5bp/15bp/$5)", + "results": [ + {"topk": 5, "n_drop": 1, "ann_return": 0.3230, "sharpe": 1.745, "maxDD": -0.0695}, + {"topk": 10, "n_drop": 1, "ann_return": 0.3284, "sharpe": 1.978, "maxDD": -0.0581}, + {"topk": 15, "n_drop": 1, "ann_return": 0.2630, "sharpe": 1.611, "maxDD": -0.0689} + ] + }, + "ndrop_sensitivity": { + "description": "vary n_drop, topk=10, costs=base", + "results": [ + {"topk": 10, "n_drop": 1, "ann_return": 0.3284, "sharpe": 1.978, "maxDD": -0.0581}, + {"topk": 10, "n_drop": 2, "ann_return": 0.2674, "sharpe": 1.586, "maxDD": -0.0652}, + {"topk": 10, "n_drop": 3, "ann_return": 0.2861, "sharpe": 1.658, "maxDD": -0.0667} + ] + }, + "cost_sensitivity": { + "description": "vary costs, topk=10, n_drop=1", + "results": [ + {"open_cost": 0.0005, "close_cost": 0.0015, "min_cost": 5, "ann_return": 0.3284, "sharpe": 1.978, "maxDD": -0.0581}, + {"open_cost": 0.0015, "close_cost": 0.0025, "min_cost": 10, "ann_return": 0.3275, "sharpe": 1.974, "maxDD": -0.0580}, + {"open_cost": 0.0025, "close_cost": 0.0035, "min_cost": 15, "ann_return": 0.3267, "sharpe": 1.971, "maxDD": -0.0580} + ] + }, + "findings": { + "topk": "topk=10 is optimal. topk=5 loses ~0.5pp (concentration risk), topk=15 loses ~6.5pp (signal dilution). Edge is moderate-sensitivity to topk.", + "n_drop": "n_drop=1 is best. n_drop=2 loses ~6pp, n_drop=3 loses ~4pp. More rotation hurts in this window.", + "costs": "Almost irrelevant. Even at 5x base costs (25bp/35bp/$15), return drops only 0.17pp (32.84% → 32.67%). Low turnover + large gross edge makes cost assumptions immaterial.", + "maxDD": "Stable across all perturbations: range -5.8% to -7.0%. No blowup risk from parameter changes.", + "overall": "The 2026 edge is robust WITHIN the window. The problem is it doesn't exist in other windows (ch 11)." + } +}