Files
book-tac/book/chapters/07-isolation-runs.md
T

46 lines
4.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 07 — Isolation Runs: Single-Variable Discipline
Status: drafting. Claim inventory: see `README.md` ch. 07.
The isolation run is the campaign's unit of proof: one variable changes against the fixed reference; the same metrics decide the verdict. This chapter walks the clean-lake sequence (exp 28–31) as the worked example, and states the discipline's rules so a reader can run it themselves.
## Why isolation is the discipline
The research loop (ch. 02) is only as honest as its attribution. A run that changes two things cannot say which one moved the result. The clean-lake campaign therefore fixed the reference — exp 26's n_drop=1 configuration on the compact stochastic set `(PROVEN → exp 26, EVIDENCE#015)` — and ran every subsequent experiment as a single-variable change against it:
| Run | One variable changed | Reference | Verdict |
|-----|----------------------|-----------|---------|
| exp 28 | seed count 5 → 2 | same features/book | REFUTED (2 seeds worse) `(PROVEN → EVIDENCE#016)` |
| exp 29 | + multi-horizon momentum bundle (M1) | same book | REFUTED `(PROVEN → EVIDENCE#017)` |
| exp 30 | + risk-adjusted 22d Sharpe drift (M2) | same book | HYPOTHESIS (mixed) `(PROVEN run, EVIDENCE#018)` |
| exp 31 | + GARCH(1,1) vol-regime trio (M3) | same book | REFUTED `(PROVEN → EVIDENCE#019)` |
Three of four refuted; one mixed. The isolation design is what makes "most additions fail" a *finding* rather than an anecdote.
## The acceptance contract
For an addition to earn its place it must beat the reference on **both** layers of the metrics ladder (ch. 01): the rank layer (IC/RankIC/ICIR/RankICIR) *and* the portfolio layer (net IR, MDD). Exp 30 is the canonical trap: M2 looked strong on the portfolio layer (net +6.53%, IR 0.62 vs +2.13%, IR 0.21) while its rank metrics were *lower* than reference (RankIC 0.0576 vs 0.0663) `(PROVEN → exp 30, EVIDENCE#018)`. Because the two layers disagreed and the run was not reproduced, the book labels it HYPOTHESIS rather than PROVEN. The rule: **a single run that improves one layer and degrades the other is a hypothesis, not a win** `TODO(evidence-needed: reproduction of exp 30 M2 on a second window)`.
## What each refutation taught
- **exp 28 (2 seeds):** the ensemble's value is tied to seed count; halving it is not a harmless cost cut (ch. 05).
- **exp 29 (momentum bundle):** drift-as-feature fails even when the underlying structure exists (ch. 01, ch. 04). The mechanism (name-specific scale, collision with trend features) is a hypothesis.
- **exp 31 (GARCH):** parametric vol modeling adds nothing to a model that already has the realized-vol ladder — the generic ladder is the feature; the parametric overlay is not (ch. 01).
- **exp 30 (M2):** the single interesting non-refutation. Risk-adjusting the drift feature changed the portfolio behavior without improving the rank — an unexplained, unreproduced anomaly worth one more run, not a claim.
## How to run an isolation campaign
1. **Freeze a reference** — a configuration, a book spec, and a metric table that everything is judged against (exp 26 n_drop=1 in this campaign).
2. **Change exactly one thing** per run; record the hypothesis and acceptance metric in the run notes *before* running (ch. 02, pre-registration).
3. **Judge on both layers** of the metrics ladder; a single-layer improvement is a hypothesis.
4. **Expect most runs to fail** — that is the point. A refuted run is a recorded negative that protects the next hypothesis from paying for the same mistake twice.
## The discipline as the value
The campaign's net-of-cost performance barely cleared costs at its best `(PROVEN → exp 26, EVIDENCE#015)`. In that regime, undisciplined feature-hunting is not neutral — it is the largest *expected* destroyer of the edge. Isolation runs converted feature-hunting into a bounded cost: a few runs to prove each family dead, instead of silently degrading the live signal. That is why the book counts exp 29 and exp 31 among its wins (ch. 12).
## Open questions
- `TODO(evidence-needed: M2 reproduction — the only surviving clean-lake improvement candidate)`
- `TODO(evidence-needed: 5d-reversal standalone strategy net of costs, the un-isolated idea from ch. 01)`
- `TODO(evidence-needed: HMM regime overlay (long-only gate) on the exp-26 book — an overlay test, not a feature test)`