Files
book-tac/book/chapters/12-synthesis.md
T

46 lines
4.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 12 — Synthesis: How Proved Truth Compounds
Status: drafting. Claim inventory: see `README.md` ch. 12.
This chapter is the scoreboard: what actually moved performance, what was refuted, and why the book's methodology — not any single experiment — is the durable product of the campaign.
## The scoreboard
Everything PROVEN below is on the clean lake (exp 21+) or a reconciled post-reset round; everything else is labeled what it is.
| Lever | What moved | Status |
|-------|-----------|--------|
| **Data quality** (clean-lake reset) | The single largest event: invalidated all pre-reset results; the reference collapsed and was rebuilt (IC 0.0354→0.0019, then rebuilt to 0.0511) | `PROVEN` — exp 21→22–24 (EVIDENCE#010–013) |
| **Cost/turnover relief** (n_drop 2→1) | The largest *positive* lever: net −3.21%→+2.13% on an identical signal (IR 0.21, MDD −7.69%) | `PROVEN` — exp 26 (EVIDENCE#015) |
| **Feature pruning** (compact generic set) | The biggest series of wins were refutations: OU (exp 25), momentum (exp 29), GARCH (exp 31) all rejected; the compact set stands (RankIC 0.0663) | `PROVEN` — exp 24/25/29/31 (EVIDENCE#013/#014/#017/#019) |
| **Ensemble & seed count** | Variance reduction, not new information; 2 seeds < 5 seeds (net −1.49% vs +2.13%) | `PROVEN` — exp 28 (EVIDENCE#016) |
| **Risk limits** | Executed and non-interfering in round 3 (liquidity floor dropped 8, funnel 10→10→10→9); the floor-beats-caps A/B is pre-reset idea material | `PROVEN` (execution) / `HYPOTHESIS` (A/B) — round 3 + exp 18 |
| **Portfolio construction** | TopkDropout beat stochastic-control on the pre-reset lake; never re-tested post-reset | `HYPOTHESIS` (idea) — exp 13/14 |
| **Live execution** | Funnel held, slippage 4.54 bps, cost ~$45, turnover 0.74 — the first reconciled live number | `PROVEN` — round 3 (EVIDENCE#020) |
## The pattern beneath the scoreboard
Two positive levers (data quality, cost relief), one protective discipline (pruning, whose wins were negatives), one reinforcement (seed count). The pattern: **performance improved by removing lies, removing cost, and removing features — not by adding anything to the signal.** The only surviving clean-lake addition candidate is exp 30's M2 (Sharpe-drift feature), which the book keeps at HYPOTHESIS precisely because it improved one layer and degraded another in a single unreproduced run (ch. 07).
The refuted runs were as valuable as the wins: exp 11, 13, 14, 20, 25, 29, 31 each stopped a wrong direction at the cost of a few runs `(PROVEN — refuted runs recorded in the ledger; REFERENCED — falsification as method)`. A campaign that counts its refutations as output is a campaign that spends its budget learning, not re-learning.
## The methodology that made it compound
None of the scoreboard above is usable without the machinery of ch. 00–02:
1. **The execution trail** (targets→decisions→fills, reconcile) is the spine — it is what let the desk catch the clean-lake collapse and what turns the live round into evidence.
2. **The clean-lake boundary** is the watermark — it is why exp 18's pretty risk numbers are hypotheses and exp 26's thin-but-real numbers are facts.
3. **Isolation and pre-registration** make each verdict attributable (ch. 07).
4. **The two-layer metrics ladder** (rank + portfolio, ch. 01) is why exp 30 is a hypothesis and not a claim.
5. **Live beats backtest** (ch. 11) is the final gate — no metric in this book outranks a reconciled round.
## What the book still does not know
- Whether the 50-ETF panel generalizes — the widest open question `TODO(evidence-needed: out-of-panel universe)`.
- Why the strongest single-feature signal (OU) degrades the model (the OU paradox, ch. 01).
- Whether 5-day reversal trades standalone net of costs (ch. 01, ch. 07).
- Whether M2 reproduces (ch. 07), whether the risk-limit A/B holds on clean data (ch. 10), and whether a second live round confirms the funnel and slippage under a different regime (ch. 11).
## Closing
This book's claims are deliberately thin: a RankIC near 0.066 on 50 names, an IR near 0.2 net, one reconciled live round. That thinness is the point. Every number in it can be re-derived from a recorded run or a re-opened round; every hypothesis is marked as one; every backtest is labeled a backtest. A quant-desk reader can act on the book's method even where its edge is small — and the book expects its own claims to be superseded as the next rounds and experiments land (living document, `AGENTS.md` rule 6).