- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches - CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims - README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign - new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated - Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
3.7 KiB
3.7 KiB
Chapter 12 — Synthesis: How Proved Truth Compounds
Status: drafting. Claim inventory: see README.md ch. 12.
This chapter is the scoreboard. It collects everything the campaign proved, in order of what actually moved performance — and why the Q-campaign's 9 refutations were as informative as its 2 passes.
The scoreboard (clean lake, exp 21–43)
| Lever | Evidence | Net effect |
|---|---|---|
| Data quality | exp 21 (collapse), 22–24 (fix + revalidation) | The single largest swing: pre-reset +7.77% became −20.6% on the same config, then reappeared as a real signal. Everything before the reset is void. |
| Cost relief / construction | exp 26 (n_drop 2→1): −3.21%→+2.13%; exp 39 (weekly, Q07): +12.51%, IR 1.24, maxDD −4.13% | The dominant positive lever. Same signal, cadence changed. |
| Feature pruning | exp 9, 24, 25, 31, 33 (Q01), 43 (Q11) | Generic pruned set > model-specific; one warranted addition (sp_sharpe_22, Q01); standalone reversal refuted (Q11). |
| Ensemble | exp 28 (2<5 seeds), exp 34 (Q02: 10 seeds) | Real but bounded: breadth saturates ~5 seeds; extra seeds don't clear costs. |
| Risk limits | exp 40 (Q08) | Safety net only: floor binds, no IR edge, DD relief is defunding. |
| Gates | exp 20, 42 (Q10) | Refuted: regime overlay churns, adds cost, no edge. |
| Sizing | exp 38 (Q06) | Weak: half-Kelly mildly positive, below bar. |
| Long-short | exp 41 (Q09) | Refuted by turnover: pre-cost edge +6.6% destroyed by $96.7k cost. |
PROVEN — EVIDENCE#010–032.
What the Q-campaign settled
Eleven pre-registered runs, two passes:
- Q01 PASS — M2's risk-adjusted 22d Sharpe-drift feature reproduces on the compact set (net +6.53%, IR 0.62). Promotes exp-30's lone result from HYPOTHESIS to PROVEN.
EVIDENCE#022. - Q07 PASS — weekly recompute is the campaign's best construction (net +12.51%, IR 1.24). The forward path.
EVIDENCE#028. - Nine refutations — seed breadth (Q02), wider book (Q03), long labels under daily churn (Q04/Q05), Kelly sizing (Q06), risk-limit-as-alpha (Q08), long-short (Q09), regime gate (Q10), standalone reversal (Q11). Each closed a direction the desk had been considering, at one-run cost each.
EVIDENCE#023–027, 029–032.
The refuted runs were as valuable as the passes: the label-horizon result (22d label, IC 0.097, yet net negative) is exactly the kind of counterintuitive fact a desk must not re-learn. REFERENCED (falsification) + PROVEN (recorded negatives).
The order of operations a reader should copy
- Fix data first — re-validate the lake before any run (ch. 06).
- Attack turnover before signal — cadence and dropped-name policy are the proven levers (ch. 09).
- Test features one at a time against the reference (ch. 07); prune, don't add (ch. 04).
- Use a small ensemble (5 seeds) and stop there (ch. 05).
- A/B risk limits before shipping; keep them as a safety net (ch. 10).
- Reconcile live — the funnel and slippage are the only claims that count (ch. 00/11).
Open questions
TODO(evidence-needed: reproduce exp 39 weekly rebalance on a second window, then a live round)TODO(evidence-needed: long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction)TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic set)TODO(evidence-needed: realized-cost reconciliation of the weekly construction against the 5bp/15bp/$5 model once live)
Evidence cited in this chapter
Composite of EVIDENCE#010–032; see the per-chapter evidence tables (ch. 04–10) and EVIDENCE.md for run/branch level citations.