Files
tac-exp-dev/evidence/q08-risklimit/RESULT.md

1.7 KiB
Raw Permalink Blame History

Q08 — Risk-limit A/B re-validation (trace 40)

Status: DONE (verdict: REFUTED as an IR edge; safety-net value retained)

Input

  • Reference signal: exp-26 pred, run 21afc6afdb674a399b59dd76c97628ce (mlflow exp 25)
  • Window: 2026-01-04 → 2026-08-10, Topk10 n_drop1, SPY benchmark, $1M, 5/15bp/$5
  • Tool: rd_risk_calibrate (A/B + sensitivity grid). Full JSON: risk_calibration.json

Candidate spec (round-3 live spec)

{"liquidity_floor_adv": 5000000, "size_cap_pct": 0.12, "concentration_cap_pct": 0.95, "drawdown_pause_pct": 0.10}

Results (net, with cost)

Config IR Ann. return Max DD
baseline (no limits) 1.5804 +27.50% −6.91%
candidate (5M floor + caps) 1.5121 +2.20% −0.65%
liquidity $10M 1.5457 +2.25% −0.64%

Findings

  • Floor binds, not a no-op: $5M liquidity floor dropped 8 symbols — DBA, DBC, ESPO, FDN, REM, TAN, UNG, XAR.
  • No IR edge from the gate: candidate IR (1.512) is BELOW baseline (1.580). The exp-18 direction (floor IR 0.81→0.98) does NOT reproduce on the clean-lake reference signal.
  • Drawdown cut is pure defunding: size_cap 0.12 × concentration 0.95 fold the effective risk_degree to ~0.0095 → ~$9.5k deployed of $1M (~100x less). Sensitivity grid shows both caps are no-ops (conc 20–50% identical, size_cap 5–20% identical); only the liquidity floor moves returns, marginally.
  • Conclusion: keep the live spec as a safety net; there is no risk-limit gate IR edge to harvest when the signal is the bottleneck (exp-20 pattern).

Artifacts on this branch

  • evidence/q08-risklimit/risk_calibration.json — full calibration dump
  • queue/designs/q08_risk_limit_ab.md — the pre-registered design doc