# Chapter 07 — Isolation Runs: Single-Variable Discipline Status: drafting. Claim inventory: see `README.md` ch. 07. The isolation run is the campaign's unit of proof: one variable changes against the fixed reference; the same metrics decide the verdict. This chapter walks the clean-lake sequence (exp 28–31) as the worked example, and states the discipline's rules so a reader can run it themselves. ## Why isolation is the discipline The research loop (ch. 02) is only as honest as its attribution. A run that changes two things cannot say which one moved the result. The clean-lake campaign therefore fixed the reference — exp 26's n_drop=1 configuration on the compact stochastic set `(PROVEN → exp 26, EVIDENCE#015)` — and ran every subsequent experiment as a single-variable change against it: | Run | One variable changed | Reference | Verdict | |-----|----------------------|-----------|---------| | exp 28 | seed count 5 → 2 | same features/book | REFUTED (2 seeds worse) `(PROVEN → EVIDENCE#016)` | | exp 29 | + multi-horizon momentum bundle (M1) | same book | REFUTED `(PROVEN → EVIDENCE#017)` | | exp 30 | + risk-adjusted 22d Sharpe drift (M2) | same book | HYPOTHESIS (mixed) `(PROVEN run, EVIDENCE#018)` | | exp 31 | + GARCH(1,1) vol-regime trio (M3) | same book | REFUTED `(PROVEN → EVIDENCE#019)` | Three of four refuted; one mixed. The isolation design is what makes "most additions fail" a *finding* rather than an anecdote. ## The acceptance contract For an addition to earn its place it must beat the reference on **both** layers of the metrics ladder (ch. 01): the rank layer (IC/RankIC/ICIR/RankICIR) *and* the portfolio layer (net IR, MDD). Exp 30 is the canonical trap: M2 looked strong on the portfolio layer (net +6.53%, IR 0.62 vs +2.13%, IR 0.21) while its rank metrics were *lower* than reference (RankIC 0.0576 vs 0.0663) `(PROVEN → exp 30, EVIDENCE#018)`. Because the two layers disagreed and the run was not reproduced, the book labels it HYPOTHESIS rather than PROVEN. The rule: **a single run that improves one layer and degrades the other is a hypothesis, not a win** `TODO(evidence-needed: reproduction of exp 30 M2 on a second window)`. ## What each refutation taught - **exp 28 (2 seeds):** the ensemble's value is tied to seed count; halving it is not a harmless cost cut (ch. 05). - **exp 29 (momentum bundle):** drift-as-feature fails even when the underlying structure exists (ch. 01, ch. 04). The mechanism (name-specific scale, collision with trend features) is a hypothesis. - **exp 31 (GARCH):** parametric vol modeling adds nothing to a model that already has the realized-vol ladder — the generic ladder is the feature; the parametric overlay is not (ch. 01). - **exp 30 (M2):** the single interesting non-refutation. Risk-adjusting the drift feature changed the portfolio behavior without improving the rank — an unexplained, unreproduced anomaly worth one more run, not a claim. ## How to run an isolation campaign 1. **Freeze a reference** — a configuration, a book spec, and a metric table that everything is judged against (exp 26 n_drop=1 in this campaign). 2. **Change exactly one thing** per run; record the hypothesis and acceptance metric in the run notes *before* running (ch. 02, pre-registration). 3. **Judge on both layers** of the metrics ladder; a single-layer improvement is a hypothesis. 4. **Expect most runs to fail** — that is the point. A refuted run is a recorded negative that protects the next hypothesis from paying for the same mistake twice. ## The discipline as the value The campaign's net-of-cost performance barely cleared costs at its best `(PROVEN → exp 26, EVIDENCE#015)`. In that regime, undisciplined feature-hunting is not neutral — it is the largest *expected* destroyer of the edge. Isolation runs converted feature-hunting into a bounded cost: a few runs to prove each family dead, instead of silently degrading the live signal. That is why the book counts exp 29 and exp 31 among its wins (ch. 12). ## Open questions - `TODO(evidence-needed: M2 reproduction — the only surviving clean-lake improvement candidate)` - `TODO(evidence-needed: 5d-reversal standalone strategy net of costs, the un-isolated idea from ch. 01)` - `TODO(evidence-needed: HMM regime overlay (long-only gate) on the exp-26 book — an overlay test, not a feature test)`