Files
tac-exp-dev/queue/designs/q08_risk_limit_ab.md

1.9 KiB

QUEUE-08 — Risk-limit A/B re-validation: $5M liquidity floor on the exp-26 reference

Status: QUEUED · Priority: P1 · Effort: tool-only (no new code)

Hypothesis (prove)

The $5M liquidity floor improves net IR and cuts drawdown on the post-reset reference signal (pre-reset exp 18, EVIDENCE#008: net IR 0.81→0.98, cumDD 7.93%→5.44%), while size/concentration caps hurt by cutting deployed capital. Needs re-validation on the exp-26 lineage because exp 18 is pre-clean-lake and not comparable (EVIDENCE#009/010). Source: book/CLAIMS.md open question + book/README.md TODO(evidence-needed: reconciliation of exp 18 risk-limit spec on the post-reset reference signal).

Change vs exp-26 reference (ONE variable)

  • Reference: the saved exp-26 prediction (run 21afc6af…, mlflow exp 25).
  • A/B via rd_risk_calibrate (runs limit-vs-no-limit A/B + sensitivity grid over size_cap_pct, concentration_cap_pct, liquidity_floor_adv) and/or rd_backtest with risk_limits on the SAME saved pred.pkl:
    • baseline: no limits (this must reproduce the exp-26 net +2.13% / IR 0.21);
    • candidate: {"liquidity_floor_adv": 5000000, "size_cap_pct": 0.12, "concentration_cap_pct": 0.95, "drawdown_pause_pct": 0.10} (round-3 spec).
  • Pick the spec (B2 calibration) that keeps live ≈ backtest.

Acceptance

  • Candidate spec: net_IR > 0.21 AND net_max_drawdown < 7.69% vs no-limit on the same pred. Size/concentration caps expected to REDUCE deployed capital (record the direction as confirmation of exp 18).
  • If the floor is a no-op (gates don't bind at this signal) → report that gates are no-ops when the signal is the bottleneck (exp 20 pattern) as a PROVEN clean-lake result.

Execution prerequisites

  • None (uses saved pred + rd_risk_calibrate/rd_backtest). Trace the A/B as an experiment; record the spec chosen for the next live round.