Files
tac-exp-dev/book/chapters/10-risk-limits-and-gates.md
T
zhaoli 06fb1e8ee9 book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2026-08-20 01:23:21 +00:00

54 lines
4.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 10 — Risk Limits and Gates: Safety Net, Not Alpha
Status: drafting. Claim inventory: see `README.md` ch. 10.
Every desk wants to believe risk controls are a performance lever. On the TradeAC clean lake the evidence says otherwise: **risk limits are a safety net, and overlay gates mostly churn.** This chapter separates the two claims — what a liquidity floor does (defund) and what a regime gate does (churn) — on the post-reset signal.
## The post-reset A/B (Q08)
The reference pred (exp 26, `21afc6af…`) was run through `rd_risk_calibrate` with the live spec `{liquidity_floor_adv: $5M, size_cap_pct: 0.12, concentration_cap_pct: 0.95, drawdown_pause_pct: 0.10}` versus no limits, same window (2026-01-04→2026-08-10), topk10/n_drop1, SPY, $1M:
| Spec | ann return | IR | maxDD |
|------|-----------|-----|-------|
| baseline (no limits) | +27.50% | **1.5804** | −6.91% |
| candidate (5M floor + caps) | +2.20% | **1.5121** | **−0.65%** |
`PROVEN — EVIDENCE#029 → exp 40`. Read the columns carefully:
- **The floor binds.** The $5M ADV floor drops DBA, DBC, ESPO, FDN, REM, TAN, UNG, XAR (8 of 50 names) — it does real work on this panel.
- **No IR edge.** Candidate IR 1.5121 < baseline 1.5804. Gating does not improve the risk-adjusted return; the floor removes small-AVD names but the surviving book has no better rank.
- **The drawdown cut is pure defunding.** size_cap 0.12 × concentration_cap 0.95 folds the effective risk_degree to ≈ 0.0095 — about **$9.5k deployed of a $1M book**. maxDD falls to −0.65% because there is almost nothing at risk, not because risk was managed well.
The pre-clean-lake claim that "$5M liquidity floor improves IR 0.81→0.98" (exp 18) is **not reproduced** on the clean-lake signal. That number stays idea material `(EVIDENCE#008 → exp 18, pre-clean-lake)`. `PROVEN (refutation) — EVIDENCE#029 → exp 40`.
## The regime gate (Q10)
A HMM regime overlay (`sp_hmm_p_regime1 ≥ 0.5` entry gate, `RegimeGateDropoutStrategy`) on the same daily signal:
- net −4.26%, IR −0.382, maxDD −7.38% — meets the drawdown leg (7.38% < 7.69%) but far below the net-IR acceptance;
- the gate churned 276 trades in ~150 days; ~6.3pp of cost erased the +2.02% gross;
- signal metrics byte-identical to the reference (IC 0.0502, RankIC 0.0660).
`PROVEN — EVIDENCE#031 → exp 42`. A regime gate that flips exposure on a regime posterior priced into the features already just adds turnover. This clean-lake re-test refutes the "gates are a free drawdown cut" idea carried from exp 20 (pre-clean-lake, byte-identical no-ops there) `(EVIDENCE#009)`.
## The synthesis
- **Risk limits**: keep them as a live harness (the round-3 live round used the same spec and the funnel held — EVIDENCE#020), but never market them as alpha. On this signal they defund, not improve. `PROVEN — EVIDENCE#029/020`.
- **Gates**: regime/momentum overlays on top of features the model already sees add turnover, not edge. `PROVEN — EVIDENCE#031/009`.
- **Where risk does earn its keep**: as a *cap on damage*, not a return source. The drawdown pause and floor are the reason the live round stays disciplined; their value is the tail, not the mean. `REFERENCED (risk-management practice) + PROVEN (round-3 funnel held under the spec)`.
## Desk rules distilled from this chapter
1. A/B any risk-limit spec against no-limits on the same pred before shipping it; if IR does not improve, it is defunding.
2. Report deployed capital alongside maxDD — a smaller drawdown with 100x less risk is not a risk win.
3. Prefer limits that bind rarely but cap hard (liquidity floor, drawdown pause) over gates that churn every day (regime overlay).
4. `TODO(evidence-needed: a live round under the weekly-rebalance construction with the risk-limit spec, to confirm the safety-net behavior at higher deployed capital)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#029` | exp 40 (Q08), MLflow run `4667984187…` (exp `tac-rd-q08-risklimit`, id 43), branch `exp/40-q08-risk-limit-ab-on-exp-26-reference-si`, `book/data/evidence/q08-risklimit/risk_calibration.json` |
| `EVIDENCE#031` | exp 42 (Q10), run `436acd01…`, branch `exp/42-q10-hmm-regime-overlay-entry-gate-on-sph` |
| `EVIDENCE#020` | round 3, trace 27, branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac` |
| `EVIDENCE#009/008` | exp 20/18 (pre-clean-lake, idea material) |