book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters

- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
This commit is contained in:
zhaoli
2026-08-20 01:23:21 +00:00
parent 436692a620
commit 06fb1e8ee9
15 changed files with 847 additions and 36 deletions
@@ -0,0 +1,35 @@
# Q08 — Risk-limit A/B re-validation (trace 40)
**Status:** DONE (verdict: REFUTED as an IR edge; safety-net value retained)
## Input
- Reference signal: exp-26 pred, run `21afc6afdb674a399b59dd76c97628ce` (mlflow exp 25)
- Window: 2026-01-04 → 2026-08-10, Topk10 n_drop1, SPY benchmark, $1M, 5/15bp/$5
- Tool: `rd_risk_calibrate` (A/B + sensitivity grid). Full JSON: `risk_calibration.json`
## Candidate spec (round-3 live spec)
`{"liquidity_floor_adv": 5000000, "size_cap_pct": 0.12, "concentration_cap_pct": 0.95, "drawdown_pause_pct": 0.10}`
## Results (net, with cost)
| Config | IR | Ann. return | Max DD |
|---|---|---|---|
| baseline (no limits) | 1.5804 | +27.50% | −6.91% |
| **candidate (5M floor + caps)** | **1.5121** | +2.20% | **−0.65%** |
| liquidity $10M | 1.5457 | +2.25% | −0.64% |
## Findings
- **Floor binds, not a no-op**: $5M liquidity floor dropped 8 symbols —
`DBA, DBC, ESPO, FDN, REM, TAN, UNG, XAR`.
- **No IR edge from the gate**: candidate IR (1.512) is BELOW baseline (1.580).
The exp-18 direction (floor IR 0.81→0.98) does NOT reproduce on the clean-lake
reference signal.
- **Drawdown cut is pure defunding**: size_cap 0.12 × concentration 0.95 fold
the effective risk_degree to ~0.0095 → ~$9.5k deployed of $1M (~100x less).
Sensitivity grid shows both caps are no-ops (conc 20–50% identical,
size_cap 5–20% identical); only the liquidity floor moves returns, marginally.
- **Conclusion**: keep the live spec as a safety net; there is no risk-limit
gate IR edge to harvest when the signal is the bottleneck (exp-20 pattern).
## Artifacts on this branch
- `evidence/q08-risklimit/risk_calibration.json` — full calibration dump
- `queue/designs/q08_risk_limit_ab.md` — the pre-registered design doc