Files
tac-exp-dev/book/chapters/07-isolation-runs.md
T
zhaoli 06fb1e8ee9 book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2026-08-20 01:23:21 +00:00

5.0 KiB
Raw Blame History

Chapter 07 — Isolation Runs: Single-Variable Discipline

Status: drafting. Claim inventory: see README.md ch. 07.

The research loop's discipline (ch. 02) is that one variable changes per run. This chapter runs that discipline across the clean-lake campaign and the Q-series (exp 33–43), and shows that most additions fail — the discipline is the value, not the win rate.

The design

The reference is the compact stochastic set, 5-seed RankIC ensemble, topk=10, n_drop=1, 5-day forward label, 50-ETF panel, train 2016-01-04→2025-09-01 / valid →2026-01-03 / test 2026-01-04→2026-08-10, account $1M, benchmark SPY, cost 5bp open / 15bp close / $5 min. Acceptance bar (from the reference): net_IR ≥ 0.21, net_ann ≥ +2.13%, maxDD ≤ 7.69% (PROVEN — exp 24/26 baseline). Every Q-run changed exactly one thing against this reference.

Q One variable changed Verdict
Q01 (exp 33) features +1: sp_sharpe_22 PASS — reproduces M2
Q02 (exp 34) seeds 5→10 (parallel 10) FAIL — signal up, book negative
Q03 (exp 35) topk 10→20 FAIL — no edge, lower vol
Q04 (exp 36) label 5d→10d FAIL — IC up, net collapses
Q05 (exp 37) label 5d→22d FAIL — best IC, flat gross
Q06 (exp 38) sizing → fractional Kelly (cap 0.5) FAIL — below bar
Q07 (exp 39) rebalance daily→weekly PASS — campaign best
Q08 (exp 40) risk-limit gates on FAIL as alpha (safety net)
Q09 (exp 41) construction → long-short FAIL — turnover kills
Q10 (exp 42) regime entry gate on FAIL — churns
Q11 (exp 43) features → single sp_trend_slope_5 FAIL — no reversal learned

PROVEN — EVIDENCE#022–032 → exp 33–43, all pre-registered in trace start + workflow YAML before each run.

What isolation bought

Because each Q-run changed one thing, the verdicts attribute cleanly:

  • Feature axis (Q01, Q11): adding the risk-adjusted Sharpe-drift feature is reproducible and positive (EVIDENCE#022 → exp 33, Q01); stripping to a single mean-reversion feature is not learnable — the model trained positive IC (+0.0023), meaning there is no standalone reversal to find in the pooled cross-section (EVIDENCE#032 → exp 43, Q11). The "mean reversion is the stable single-feature edge" hypothesis is refuted on the clean lake.
  • Label axis (Q04, Q05): longer forward-return labels monotonically improve the signal — 10d: IC 0.0925 / RankIC 0.0960; 22d: IC 0.0970 / RankIC 0.1165 — yet net-of-cost performance worsens (10d: −9.92% IR −1.15; 22d: −4.60% IR −0.59). PROVEN — EVIDENCE#025/026 → exp 36/37. Horizon signal and daily-turnover construction are incompatible.
  • Model axis (Q02): 10 seeds raise rank breadth (RankIC 0.0671, L/S Sharpe 4.58) but the book stays negative net (−0.93%) — the added breadth never crosses the cost barrier. PROVEN — EVIDENCE#023 → exp 34.
  • Construction axis (Q03, Q06, Q07, Q09): see ch. 08 — weekly recompute (Q07) is the only change that clears the bar by a wide margin.
  • Risk/gate axis (Q08, Q10): see ch. 10 — both met at most a drawdown leg; neither adds alpha.

The discipline is the output

Only 2 of 11 Q-runs passed. That is not a failure of the campaign — it is the mechanism doing its job. Each FAIL closed a candidate direction at the cost of one run, and the two PASSes (Q01 reproducing the M2 feature, Q07 the weekly construction) are the campaign's forward path. The campaign's refuted runs were as valuable as its wins: knowing that a 22d label has an IC of 0.097 and still loses money daily is exactly the kind of fact a desk must not learn twice. REFERENCED (falsification) + PROVEN (recorded negatives) — EVIDENCE#022–032.

Desk rules distilled from this chapter

  1. Fix the reference and the acceptance bar before the series; change one variable per run.
  2. Record signal metrics and net-of-cost metrics side by side — a signal gain that does not clear costs is not a strategy gain (Q02, Q04, Q05).
  3. A single-feature "obvious" edge must be tested standalone before being trusted in a bundle (Q11 refuted it).
  4. TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)

Evidence cited in this chapter

Tag Source
EVIDENCE#022 exp 33 (Q01), run c7c12228…, branch exp/33-q01-m2-reproduction-add-spsharpe22-to-th
EVIDENCE#023 exp 34 (Q02), run ce49e4e0…, branch exp/34-q02-seed10-10-seed-rankicensemble-vs-ref
EVIDENCE#024 exp 35 (Q03), run 2a844c02…, branch exp/35-q03-topk20-widen-topkdropout-portfolio-f
EVIDENCE#025 exp 36 (Q04), run ef211826…, branch exp/36-q04-label10d-10d-forward-return-label-vs
EVIDENCE#026 exp 37 (Q05), run daad5042…, branch exp/37-q05-label22d-22d-forward-return-label-vs
EVIDENCE#032 exp 43 (Q11), run e859adfe…, branch exp/43-q11-standalone-5d-reversal-single-featur
reference exp 24/26 (compact set / n_drop=1), runs fe469a19…/21afc6af…