Files
tac-exp-dev/book/chapters/07-isolation-runs.md
T
zhaoli 06fb1e8ee9 book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2026-08-20 01:23:21 +00:00

58 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 07 — Isolation Runs: Single-Variable Discipline
Status: drafting. Claim inventory: see `README.md` ch. 07.
The research loop's discipline (ch. 02) is that one variable changes per run. This chapter runs that discipline across the clean-lake campaign and the Q-series (exp 33–43), and shows that most additions fail — the discipline is the value, not the win rate.
## The design
The reference is the compact stochastic set, 5-seed RankIC ensemble, topk=10, n_drop=1, 5-day forward label, 50-ETF panel, train 2016-01-04→2025-09-01 / valid →2026-01-03 / test 2026-01-04→2026-08-10, account $1M, benchmark SPY, cost 5bp open / 15bp close / $5 min. Acceptance bar (from the reference): **net_IR ≥ 0.21, net_ann ≥ +2.13%, maxDD ≤ 7.69%** `(PROVEN — exp 24/26 baseline)`. Every Q-run changed exactly one thing against this reference.
| Q | One variable changed | Verdict |
|---|---------------------|---------|
| Q01 (exp 33) | features +1: `sp_sharpe_22` | PASS — reproduces M2 |
| Q02 (exp 34) | seeds 5→10 (parallel 10) | FAIL — signal up, book negative |
| Q03 (exp 35) | topk 10→20 | FAIL — no edge, lower vol |
| Q04 (exp 36) | label 5d→10d | FAIL — IC up, net collapses |
| Q05 (exp 37) | label 5d→22d | FAIL — best IC, flat gross |
| Q06 (exp 38) | sizing → fractional Kelly (cap 0.5) | FAIL — below bar |
| Q07 (exp 39) | rebalance daily→weekly | PASS — campaign best |
| Q08 (exp 40) | risk-limit gates on | FAIL as alpha (safety net) |
| Q09 (exp 41) | construction → long-short | FAIL — turnover kills |
| Q10 (exp 42) | regime entry gate on | FAIL — churns |
| Q11 (exp 43) | features → single `sp_trend_slope_5` | FAIL — no reversal learned |
`PROVEN — EVIDENCE#022–032 → exp 33–43, all pre-registered in trace start + workflow YAML before each run`.
## What isolation bought
Because each Q-run changed one thing, the verdicts attribute cleanly:
- **Feature axis (Q01, Q11):** adding the risk-adjusted Sharpe-drift feature is reproducible and positive `(EVIDENCE#022 → exp 33, Q01)`; stripping to a single mean-reversion feature is not learnable — the model trained *positive* IC (+0.0023), meaning there is no standalone reversal to find in the pooled cross-section `(EVIDENCE#032 → exp 43, Q11)`. The "mean reversion is the stable single-feature edge" hypothesis is **refuted** on the clean lake.
- **Label axis (Q04, Q05):** longer forward-return labels monotonically *improve* the signal — 10d: IC 0.0925 / RankIC 0.0960; 22d: IC 0.0970 / RankIC 0.1165 — yet net-of-cost performance *worsens* (10d: −9.92% IR −1.15; 22d: −4.60% IR −0.59). `PROVEN — EVIDENCE#025/026 → exp 36/37`. Horizon signal and daily-turnover construction are incompatible.
- **Model axis (Q02):** 10 seeds raise rank breadth (RankIC 0.0671, L/S Sharpe 4.58) but the book stays negative net (−0.93%) — the added breadth never crosses the cost barrier. `PROVEN — EVIDENCE#023 → exp 34`.
- **Construction axis (Q03, Q06, Q07, Q09):** see ch. 08 — weekly recompute (Q07) is the only change that clears the bar by a wide margin.
- **Risk/gate axis (Q08, Q10):** see ch. 10 — both met at most a drawdown leg; neither adds alpha.
## The discipline is the output
Only 2 of 11 Q-runs passed. That is not a failure of the campaign — it is the mechanism doing its job. Each FAIL closed a candidate direction at the cost of one run, and the two PASSes (Q01 reproducing the M2 feature, Q07 the weekly construction) are the campaign's forward path. The campaign's refuted runs were as valuable as its wins: knowing that a 22d label has an IC of 0.097 *and still loses money daily* is exactly the kind of fact a desk must not learn twice. `REFERENCED (falsification) + PROVEN (recorded negatives) — EVIDENCE#022–032`.
## Desk rules distilled from this chapter
1. Fix the reference and the acceptance bar *before* the series; change one variable per run.
2. Record signal metrics and net-of-cost metrics side by side — a signal gain that does not clear costs is not a strategy gain (Q02, Q04, Q05).
3. A single-feature "obvious" edge must be tested standalone before being trusted in a bundle (Q11 refuted it).
4. `TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#022` | exp 33 (Q01), run `c7c12228…`, branch `exp/33-q01-m2-reproduction-add-spsharpe22-to-th` |
| `EVIDENCE#023` | exp 34 (Q02), run `ce49e4e0…`, branch `exp/34-q02-seed10-10-seed-rankicensemble-vs-ref` |
| `EVIDENCE#024` | exp 35 (Q03), run `2a844c02…`, branch `exp/35-q03-topk20-widen-topkdropout-portfolio-f` |
| `EVIDENCE#025` | exp 36 (Q04), run `ef211826…`, branch `exp/36-q04-label10d-10d-forward-return-label-vs` |
| `EVIDENCE#026` | exp 37 (Q05), run `daad5042…`, branch `exp/37-q05-label22d-22d-forward-return-label-vs` |
| `EVIDENCE#032` | exp 43 (Q11), run `e859adfe…`, branch `exp/43-q11-standalone-5d-reversal-single-featur` |
| reference | exp 24/26 (compact set / n_drop=1), runs `fe469a19…`/`21afc6af…` |