58 lines
5.3 KiB
Markdown
58 lines
5.3 KiB
Markdown
# Chapter 13 — Synthesis: How Proved Truth Compounds
|
||
|
||
Status: drafting. Claim inventory: see `README.md` ch. 13.
|
||
|
||
This chapter is the scoreboard. It collects everything the campaign proved, in order of what actually moved performance — and why the Q-campaign's 9 refutations were as informative as its 2 passes.
|
||
|
||
## The scoreboard (clean lake, exp 21–43)
|
||
|
||
| Lever | Evidence | Net effect |
|
||
|-------|----------|-----------|
|
||
| **Data quality** | exp 21 (collapse), 22–24 (fix + revalidation) | The single largest swing: pre-reset +7.77% became −20.6% on the same config, then reappeared as a real signal. Everything before the reset is void. |
|
||
| **Cost relief / construction** | exp 26 (n_drop 2→1): −3.21%→+2.13%; exp 39 (weekly, Q07): **+12.51%, IR 1.24, maxDD −4.13%** | The dominant *positive* lever. Same signal, cadence changed. |
|
||
| **Feature pruning** | exp 9, 24, 25, 31, 33 (Q01), 43 (Q11) | Generic pruned set > model-specific; one warranted addition (sp_sharpe_22, Q01); standalone reversal refuted (Q11). |
|
||
| **Ensemble** | exp 28 (2<5 seeds), exp 34 (Q02: 10 seeds) | Real but bounded: breadth saturates ~5 seeds; extra seeds don't clear costs. |
|
||
| **Risk limits** | exp 40 (Q08) | Safety net only: floor binds, no IR edge, DD relief is defunding. |
|
||
| **Gates** | exp 20, 42 (Q10) | Refuted: regime overlay churns, adds cost, no edge. |
|
||
| **Sizing** | exp 38 (Q06) | Weak: half-Kelly mildly positive, below bar. |
|
||
| **Long-short** | exp 41 (Q09) | Refuted by turnover: pre-cost edge +6.6% destroyed by $96.7k cost. |
|
||
|
||
`PROVEN — EVIDENCE#010–032`.
|
||
|
||
## What the Q-campaign settled
|
||
|
||
Eleven pre-registered runs, two passes:
|
||
|
||
1. **Q01 PASS** — M2's risk-adjusted 22d Sharpe-drift feature reproduces on the compact set (net +6.53%, IR 0.62). Promotes exp-30's lone result from HYPOTHESIS to PROVEN. `EVIDENCE#022`.
|
||
2. **Q07 PASS** — weekly recompute is the campaign's best construction (net +12.51%, IR 1.24). The forward path. `EVIDENCE#028`.
|
||
3. **Nine refutations** — seed breadth (Q02), wider book (Q03), long labels under daily churn (Q04/Q05), Kelly sizing (Q06), risk-limit-as-alpha (Q08), long-short (Q09), regime gate (Q10), standalone reversal (Q11). Each closed a direction the desk had been considering, at one-run cost each. `EVIDENCE#023–027, 029–032`.
|
||
|
||
The refuted runs were as valuable as the passes: the label-horizon result (22d label, IC 0.097, yet net negative) is exactly the kind of counterintuitive fact a desk must not re-learn. `REFERENCED (falsification) + PROVEN (recorded negatives)`.
|
||
|
||
## The order of operations a reader should copy
|
||
|
||
1. **Fix data first** — re-validate the lake before any run (ch. 06).
|
||
2. **Attack turnover before signal** — cadence and dropped-name policy are the proven levers (ch. 09).
|
||
3. **Test features one at a time** against the reference (ch. 07); prune, don't add (ch. 04).
|
||
4. **Use a small ensemble** (5 seeds) and stop there (ch. 05).
|
||
5. **A/B risk limits** before shipping; keep them as a safety net (ch. 10).
|
||
6. **Reconcile live** — the funnel and slippage are the only claims that count (ch. 00/12).
|
||
7. **Re-validate walk-forward before shipping** — a single-window edge is a regime artifact until it survives re-training out-of-window (ch. 11).
|
||
|
||
## Open questions
|
||
|
||
- `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
|
||
- `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)`
|
||
- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
|
||
|
||
### Settled open questions
|
||
|
||
- ~~`reproduce exp 39 weekly rebalance on a second window, then a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`.
|
||
- ~~`long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`.
|
||
- ~~`out-of-universe (non-ETF) validation of the compact stochastic set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`.
|
||
- ~~`do the campaign's headline results (Q01 m2-sharpe22, Q07 weekly) survive walk-forward re-training?`~~ — **ANSWERED (negatively):** exp 52/53/54 re-ran every headline config across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable (+6.5% m2, +12.5% weekly); all prior years are negative or flat. **The edge is a 2025–2026 regime artifact.** `EVIDENCE#043–045`.
|
||
- ~~`is there a pre-deployment guard that isolates the profitable regime?`~~ — **ANSWERED (negatively):** five guard candidates (feature-PSI, label-regime PSI, streaming IC `ic_min_rankic`, adaptive short-window, window-staleness) were pre-registered and all refuted; none selects the good years in hindsight. `EVIDENCE#043–047`; see ch. 11.
|
||
|
||
## Evidence cited in this chapter
|
||
|
||
Composite of `EVIDENCE#010–032`; see the per-chapter evidence tables (ch. 04–10) and `EVIDENCE.md` for run/branch level citations. |