- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches - CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims - README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign - new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated - Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
54 lines
4.3 KiB
Markdown
54 lines
4.3 KiB
Markdown
# Chapter 10 — Risk Limits and Gates: Safety Net, Not Alpha
|
||
|
||
Status: drafting. Claim inventory: see `README.md` ch. 10.
|
||
|
||
Every desk wants to believe risk controls are a performance lever. On the TradeAC clean lake the evidence says otherwise: **risk limits are a safety net, and overlay gates mostly churn.** This chapter separates the two claims — what a liquidity floor does (defund) and what a regime gate does (churn) — on the post-reset signal.
|
||
|
||
## The post-reset A/B (Q08)
|
||
|
||
The reference pred (exp 26, `21afc6af…`) was run through `rd_risk_calibrate` with the live spec `{liquidity_floor_adv: $5M, size_cap_pct: 0.12, concentration_cap_pct: 0.95, drawdown_pause_pct: 0.10}` versus no limits, same window (2026-01-04→2026-08-10), topk10/n_drop1, SPY, $1M:
|
||
|
||
| Spec | ann return | IR | maxDD |
|
||
|------|-----------|-----|-------|
|
||
| baseline (no limits) | +27.50% | **1.5804** | −6.91% |
|
||
| candidate (5M floor + caps) | +2.20% | **1.5121** | **−0.65%** |
|
||
|
||
`PROVEN — EVIDENCE#029 → exp 40`. Read the columns carefully:
|
||
|
||
- **The floor binds.** The $5M ADV floor drops DBA, DBC, ESPO, FDN, REM, TAN, UNG, XAR (8 of 50 names) — it does real work on this panel.
|
||
- **No IR edge.** Candidate IR 1.5121 < baseline 1.5804. Gating does not improve the risk-adjusted return; the floor removes small-AVD names but the surviving book has no better rank.
|
||
- **The drawdown cut is pure defunding.** size_cap 0.12 × concentration_cap 0.95 folds the effective risk_degree to ≈ 0.0095 — about **$9.5k deployed of a $1M book**. maxDD falls to −0.65% because there is almost nothing at risk, not because risk was managed well.
|
||
|
||
The pre-clean-lake claim that "$5M liquidity floor improves IR 0.81→0.98" (exp 18) is **not reproduced** on the clean-lake signal. That number stays idea material `(EVIDENCE#008 → exp 18, pre-clean-lake)`. `PROVEN (refutation) — EVIDENCE#029 → exp 40`.
|
||
|
||
## The regime gate (Q10)
|
||
|
||
A HMM regime overlay (`sp_hmm_p_regime1 ≥ 0.5` entry gate, `RegimeGateDropoutStrategy`) on the same daily signal:
|
||
|
||
- net −4.26%, IR −0.382, maxDD −7.38% — meets the drawdown leg (7.38% < 7.69%) but far below the net-IR acceptance;
|
||
- the gate churned 276 trades in ~150 days; ~6.3pp of cost erased the +2.02% gross;
|
||
- signal metrics byte-identical to the reference (IC 0.0502, RankIC 0.0660).
|
||
|
||
`PROVEN — EVIDENCE#031 → exp 42`. A regime gate that flips exposure on a regime posterior priced into the features already just adds turnover. This clean-lake re-test refutes the "gates are a free drawdown cut" idea carried from exp 20 (pre-clean-lake, byte-identical no-ops there) `(EVIDENCE#009)`.
|
||
|
||
## The synthesis
|
||
|
||
- **Risk limits**: keep them as a live harness (the round-3 live round used the same spec and the funnel held — EVIDENCE#020), but never market them as alpha. On this signal they defund, not improve. `PROVEN — EVIDENCE#029/020`.
|
||
- **Gates**: regime/momentum overlays on top of features the model already sees add turnover, not edge. `PROVEN — EVIDENCE#031/009`.
|
||
- **Where risk does earn its keep**: as a *cap on damage*, not a return source. The drawdown pause and floor are the reason the live round stays disciplined; their value is the tail, not the mean. `REFERENCED (risk-management practice) + PROVEN (round-3 funnel held under the spec)`.
|
||
|
||
## Desk rules distilled from this chapter
|
||
|
||
1. A/B any risk-limit spec against no-limits on the same pred before shipping it; if IR does not improve, it is defunding.
|
||
2. Report deployed capital alongside maxDD — a smaller drawdown with 100x less risk is not a risk win.
|
||
3. Prefer limits that bind rarely but cap hard (liquidity floor, drawdown pause) over gates that churn every day (regime overlay).
|
||
4. `TODO(evidence-needed: a live round under the weekly-rebalance construction with the risk-limit spec, to confirm the safety-net behavior at higher deployed capital)`
|
||
|
||
## Evidence cited in this chapter
|
||
|
||
| Tag | Source |
|
||
|-----|--------|
|
||
| `EVIDENCE#029` | exp 40 (Q08), MLflow run `4667984187…` (exp `tac-rd-q08-risklimit`, id 43), branch `exp/40-q08-risk-limit-ab-on-exp-26-reference-si`, `book/data/evidence/q08-risklimit/risk_calibration.json` |
|
||
| `EVIDENCE#031` | exp 42 (Q10), run `436acd01…`, branch `exp/42-q10-hmm-regime-overlay-entry-gate-on-sph` |
|
||
| `EVIDENCE#020` | round 3, trace 27, branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac` |
|
||
| `EVIDENCE#009/008` | exp 20/18 (pre-clean-lake, idea material) | |