From 9b96ee6dbf567d4eb85730a84a0c6885005e1918 Mon Sep 17 00:00:00 2001 From: TradeAC Book Agent Date: Tue, 18 Aug 2026 23:18:29 +0000 Subject: [PATCH] =?UTF-8?q?book:=20ch07=20isolation=20runs=20single-variab?= =?UTF-8?q?le=20discipline=20=E2=80=94=20exp=2026/28-31,=20evidence=20#015?= =?UTF-8?q?-#019?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- book/chapters/07-isolation-runs.md | 46 ++++++++++++++++++++++++++++++ 1 file changed, 46 insertions(+) create mode 100644 book/chapters/07-isolation-runs.md diff --git a/book/chapters/07-isolation-runs.md b/book/chapters/07-isolation-runs.md new file mode 100644 index 0000000..370f14d --- /dev/null +++ b/book/chapters/07-isolation-runs.md @@ -0,0 +1,46 @@ +# Chapter 07 — Isolation Runs: Single-Variable Discipline + +Status: drafting. Claim inventory: see `README.md` ch. 07. + +The isolation run is the campaign's unit of proof: one variable changes against the fixed reference; the same metrics decide the verdict. This chapter walks the clean-lake sequence (exp 28–31) as the worked example, and states the discipline's rules so a reader can run it themselves. + +## Why isolation is the discipline + +The research loop (ch. 02) is only as honest as its attribution. A run that changes two things cannot say which one moved the result. The clean-lake campaign therefore fixed the reference — exp 26's n_drop=1 configuration on the compact stochastic set `(PROVEN → exp 26, EVIDENCE#015)` — and ran every subsequent experiment as a single-variable change against it: + +| Run | One variable changed | Reference | Verdict | +|-----|----------------------|-----------|---------| +| exp 28 | seed count 5 → 2 | same features/book | REFUTED (2 seeds worse) `(PROVEN → EVIDENCE#016)` | +| exp 29 | + multi-horizon momentum bundle (M1) | same book | REFUTED `(PROVEN → EVIDENCE#017)` | +| exp 30 | + risk-adjusted 22d Sharpe drift (M2) | same book | HYPOTHESIS (mixed) `(PROVEN run, EVIDENCE#018)` | +| exp 31 | + GARCH(1,1) vol-regime trio (M3) | same book | REFUTED `(PROVEN → EVIDENCE#019)` | + +Three of four refuted; one mixed. The isolation design is what makes "most additions fail" a *finding* rather than an anecdote. + +## The acceptance contract + +For an addition to earn its place it must beat the reference on **both** layers of the metrics ladder (ch. 01): the rank layer (IC/RankIC/ICIR/RankICIR) *and* the portfolio layer (net IR, MDD). Exp 30 is the canonical trap: M2 looked strong on the portfolio layer (net +6.53%, IR 0.62 vs +2.13%, IR 0.21) while its rank metrics were *lower* than reference (RankIC 0.0576 vs 0.0663) `(PROVEN → exp 30, EVIDENCE#018)`. Because the two layers disagreed and the run was not reproduced, the book labels it HYPOTHESIS rather than PROVEN. The rule: **a single run that improves one layer and degrades the other is a hypothesis, not a win** `TODO(evidence-needed: reproduction of exp 30 M2 on a second window)`. + +## What each refutation taught + +- **exp 28 (2 seeds):** the ensemble's value is tied to seed count; halving it is not a harmless cost cut (ch. 05). +- **exp 29 (momentum bundle):** drift-as-feature fails even when the underlying structure exists (ch. 01, ch. 04). The mechanism (name-specific scale, collision with trend features) is a hypothesis. +- **exp 31 (GARCH):** parametric vol modeling adds nothing to a model that already has the realized-vol ladder — the generic ladder is the feature; the parametric overlay is not (ch. 01). +- **exp 30 (M2):** the single interesting non-refutation. Risk-adjusting the drift feature changed the portfolio behavior without improving the rank — an unexplained, unreproduced anomaly worth one more run, not a claim. + +## How to run an isolation campaign + +1. **Freeze a reference** — a configuration, a book spec, and a metric table that everything is judged against (exp 26 n_drop=1 in this campaign). +2. **Change exactly one thing** per run; record the hypothesis and acceptance metric in the run notes *before* running (ch. 02, pre-registration). +3. **Judge on both layers** of the metrics ladder; a single-layer improvement is a hypothesis. +4. **Expect most runs to fail** — that is the point. A refuted run is a recorded negative that protects the next hypothesis from paying for the same mistake twice. + +## The discipline as the value + +The campaign's net-of-cost performance barely cleared costs at its best `(PROVEN → exp 26, EVIDENCE#015)`. In that regime, undisciplined feature-hunting is not neutral — it is the largest *expected* destroyer of the edge. Isolation runs converted feature-hunting into a bounded cost: a few runs to prove each family dead, instead of silently degrading the live signal. That is why the book counts exp 29 and exp 31 among its wins (ch. 12). + +## Open questions + +- `TODO(evidence-needed: M2 reproduction — the only surviving clean-lake improvement candidate)` +- `TODO(evidence-needed: 5d-reversal standalone strategy net of costs, the un-isolated idea from ch. 01)` +- `TODO(evidence-needed: HMM regime overlay (long-only gate) on the exp-26 book — an overlay test, not a feature test)` \ No newline at end of file