book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters

- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
This commit is contained in:
zhaoli
2026-08-20 01:23:21 +00:00
parent 436692a620
commit 06fb1e8ee9
15 changed files with 847 additions and 36 deletions
+1 -1
View File
@@ -32,7 +32,7 @@ None of these numbers — slippage in bps, cost as a fraction of gross, the rati
## Backtests are historical, not promises
Throughout this book, backtest metrics carry a warning label, not a hiding place: universe, date window, and whether the hypothesis was pre-registered before the run. This matters because TradeAC ran 31+ experiments; with that many draws, some positive results will be luck. The book is explicit about which runs were pre-registered (e.g. isolation runs exp 28–31) and which were exploratory. `REFERENCED` — multiple-testing/cherry-picking risk is standard research practice; see `references/` as it accrues.
Throughout this book, backtest metrics carry a warning label, not a hiding place: universe, date window, and whether the hypothesis was pre-registered before the run. This matters because TradeAC ran 40+ experiments; with that many draws, some positive results will be luck. The book is explicit about which runs were pre-registered (e.g. isolation runs exp 28–31 and the Q-campaign exp 33–43) and which were exploratory. `REFERENCED` — multiple-testing/cherry-picking risk is standard research practice; see `references/` as it accrues.
The most important proof of this discipline is the clean-lake reset, which this book treats as a turning point rather than a footnote: the pre-reset reference signal did **not** reproduce on a rebuilt lake (`EVIDENCE#010 → exp 21`). Had the book quoted the pre-reset backtest as fact, it would have shipped a lie. The trail and the traceability loop are what allowed the desk to catch it. Chapter 06 tells that story in full.
+3 -3
View File
@@ -12,7 +12,7 @@ The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. T
## Why traceability is the methodology
TradeAC ran 31+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:
TradeAC ran 40+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:
- **The dirty lake.** Pre-reset exp 8–18 reported strong results that did not survive a clean rebuild `(PROVEN → exp 21)`. Only because the exact YAML, branch, and run were recorded could the desk reproduce — and falsify — the reference. Traceability is what turned a false belief into evidence.
- **Post-hoc cherry-picking.** With 31+ experiments, the best-looking number is expected to be inflated by selection. The counter is pre-registration: hypothesis, change, and acceptance metric are fixed in the run notes *before* the run `(REFERENCED — research practice; see CLAIMS.md multiple-testing note)`. Where the book quotes an experiment whose hypothesis was recorded after the fact, it says so.
@@ -21,11 +21,11 @@ The pre-reset experiments (exp 8–18) are therefore treated as **idea material,
## The isolation discipline
A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio. Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.
A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio; and the Q-campaign (exp 33–43) changed exactly one thing per run against that same reference — features (Q01, Q11), seeds (Q02), topk (Q03), label horizon (Q04/Q05), sizing (Q06), rebalance cadence (Q07), risk limits (Q08), construction (Q09), and an entry gate (Q10). Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31; EVIDENCE#022–032 → exp 33–43)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.
## Falsification is the output
Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.
Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14), and nine of the eleven Q-campaign runs (longer labels, wider books, Kelly sizing, long-short, regime gates, standalone reversal — exp 34–38, 40–43). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.
## From here
+8 -3
View File
@@ -31,16 +31,20 @@ The clean-lake sequence shows the pattern with the same signal, same costs, vary
| exp 23 (general sp only) | 0.0728 / 0.206 | +6.73% | −2.39% | −0.22 |
| exp 24 (compact sp) | 0.0511 / 0.255 | +5.99% | −3.21% | −0.32 |
| exp 26 (compact, n_drop=1) | 0.0511 / 0.255 | +7.02% | +2.13% | +0.21 |
| exp 39 (compact, n_drop=1, weekly) | 0.0511 / 0.255 | +13.59% | **+12.51%** | **+1.24** |
`PROVEN — EVIDENCE#011/012/013/015 → exp 22/23/24/26`. Read the columns, not the rows: even the *best* clean-lake signal, at the default construction, lost roughly **nine to ten percentage points of annualized excess to costs** (exp 24: +5.99% gross → −3.21% net). The signal that produced a high long-short Sharpe (L/S ann Sharpe 4.54) could not survive daily rebalancing at 20 bp round trips.
`PROVEN — EVIDENCE#011/012/013/015/028 → exp 22/23/24/26/39`. Read the columns, not the rows: even the *best* clean-lake signal, at the default daily construction, lost roughly **nine to ten percentage points of annualized excess to costs** (exp 24: +5.99% gross → −3.21% net). The signal that produced a high long-short Sharpe (L/S ann Sharpe 4.54) could not survive daily rebalancing at 20 bp round trips. The weekly-rebalance row (exp 39, Q07) is the contrast that makes the diagnosis airtight: **the same signal, same costs, same topk/n_drop — only the cadence changed — and the cost drag collapsed to ~1.1pp, turning +2.13% into +12.51% net.** `PROVEN — EVIDENCE#028 → exp 39`; see ch. 08/09.
This is the single most important number in the early book: **at this turnover, cost is not a haircut, it is the strategy's budget.** `PROVEN — EVIDENCE#015 → exp 26 (identical IC/RankIC across n_drop 2 and 1; the entire net difference is trading behavior, not signal)`. The pre-reset campaign observed the same shape historically (baseline +6.2% gross → +1.6% net), which is idea material, not evidence: `HYPOTHESIS (idea: pre-clean-lake, EVIDENCE#002 → exp 8)`.
## What fixed it, and what it implies
The only construction change that flipped net from negative to positive was reducing daily forced replacements from `n_drop=2` to `n_drop=1` — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`.
Two construction changes flipped net from negative to positive — and both were cost relief, not signal:
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 07) then confirmed that most candidate additions *reduced* the edge anyway.
1. **n_drop=2 → 1** (exp 26): holding the previously-dropped name instead of trading around it. Signal metrics byte-identical to exp 24; the gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`.
2. **Weekly recompute** (exp 39, Q07): re-selecting the topk once per week instead of every day, same signal, same topk/n_drop. Cost drag fell to ~1.1pp and net reached +12.51% (IR 1.24). `PROVEN — EVIDENCE#028 → exp 39`. The weekly construction is now the campaign's best result and the book's recommended path forward (ch. 08).
Methodological reading: when the gross edge is ~7–14% and the cost drag is measured in percentage points per quarter of turnover, the two levers with the largest expected payoffs are *cost reduction* (turnover, cadence, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 07) then confirmed that most candidate additions *reduced* the edge anyway.
## Desk rules distilled from this chapter
@@ -57,5 +61,6 @@ Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, t
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
| `EVIDENCE#011/012` | exp 22/23, runs `18db5bc1…` / `be5cd314…` |
| `EVIDENCE#015` | exp 26, run `21afc6af…`, branch `exp/26-test-whether-reducing-topkdropout-daily` |
| `EVIDENCE#028` | exp 39 (Q07), run `eb38588c…`, branch `exp/39-q07-weekly-rebalance-recompute-topkdropo` |
| `EVIDENCE#002` | exp 8 (pre-clean-lake, idea only) |
| chat mining | book/data/chat_mining/exp-polluted-lake.txt (null calibration, idea only) |
+36
View File
@@ -0,0 +1,36 @@
# Chapter 04 — Prune, Don't Add: Feature-Family Ablation
Status: drafting. Claim inventory: see `README.md` ch. 04.
The first instinct of a quant desk with a weak signal is to add features. On the TradeAC 50-ETF panel the evidence runs the other way: **the signal that survived the clean-lake reset is a *pruned* set, and the single-feature additions that "should" work mostly did not.** This chapter collects the ablation evidence and its clean-lake re-tests.
## The pattern: generic beats specific
The pre-clean-lake ablation (exp 9) established the direction: dropping model-specific families (ou, hmm) and keeping the general stochastic families improved the rank signal (RankIC 0.030→0.064). Adding moment/volatility families regressed it (exp 11). Those runs are pre-reset idea material, but the *direction* re-proved itself on clean data: OU reversion hurts (exp 25, `EVIDENCE#014`), GARCH adds nothing (exp 31, `EVIDENCE#019`), and the compact general stochastic set is the reference (exp 24, `EVIDENCE#013`). `HYPOTHESIS (pre-clean-lake) → PROVEN (clean-lake direction, exp 24/25/31)`.
## Two clean-lake additions, one prune (Q01, Q11)
The Q-campaign tested the two extremes of the feature axis against the reference:
- **Add (Q01):** `sp_sharpe_22` — the 22-day risk-adjusted Sharpe drift from exp 30's M2 — reproduced exactly on the compact set: IC 0.0464, RankIC 0.0578, net +6.53% (IR 0.62), maxDD −8.0%. This is a **warranted addition**: it carries the M2 edge into the pruned set. `PROVEN — EVIDENCE#022 → exp 33`.
- **Prune to one (Q11):** the single "obvious" mean-reversion feature `sp_trend_slope_5` alone (plus raw OHLCV) — the hypothesis was that a standalone 5d reversal exists. The model trained **positive** IC (+0.0023), so it did not learn reversal at all; gross −10.4%, net −15.2%. `PROVEN — EVIDENCE#032 → exp 43`. The "mean reversion is the stable single-feature edge" claim is refuted on the clean lake; the pooled trend-slope reversal beta does not survive as a standalone.
The lesson is the pair taken together: on a 50-name daily panel, one principled addition (Sharpe drift) helped and reproduced, while the "obvious" reversal feature did not exist standalone. Feature decisions need isolation runs, not intuition (ch. 07).
## What this means for a reader
- Treat "more features" as a hypothesis, tested one at a time against the reference.
- Prefer scale-free general statistics (volatility, jump, trend, signature) over model-specific machinery (HMM/OU states) on a small cross-section.
- A feature that fails as a bundle member is not necessarily dead (OU); a feature that fails standalone is not necessarily live in a bundle — test both directions (Q01 = bundle→isolated-addition PASS, Q11 = standalone FAIL).
- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#022` | exp 33 (Q01), run `c7c12228…`, branch `exp/33-q01-m2-reproduction-add-spsharpe22-to-th` |
| `EVIDENCE#032` | exp 43 (Q11), run `e859adfe…`, branch `exp/43-q11-standalone-5d-reversal-single-featur` |
| `EVIDENCE#014` | exp 25, run `57450d1a…`, branch `exp/25-test-the-clean-data-hypothesis-that-addi` |
| `EVIDENCE#019` | exp 31, run `514cb523…`, branch `exp/31-isolation-run-m3-does-adding-garch11-vol` |
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
| pre-clean ideas | exp 9/11, EVIDENCE#003/004 |
+30
View File
@@ -0,0 +1,30 @@
# Chapter 05 — Ensembles and the Seed-Count Effect
Status: drafting. Claim inventory: see `README.md` ch. 05.
TradeAC's reference signal is a **5-seed RankIC ensemble** of LightGBM rankers. This chapter records what seed count is worth — and, from the Q-campaign, what it is *not* worth.
## What averaging buys (and its limits)
- **2 < 5 seeds (proven):** on the clean lake the 2-seed ensemble lost to the 5-seed on every rank metric and net (RankIC 0.0579 vs 0.0663; net −1.49% IR −0.14 vs +2.13% IR +0.21). `PROVEN — EVIDENCE#016 → exp 28`.
- **5 seeds = the reference** (compact stochastic set): RankIC 0.0663, RankICIR 0.2545. `PROVEN — EVIDENCE#013 → exp 24`.
- **5→10 seeds (new, Q02):** the Q-campaign tested whether more breadth keeps paying. The 10-seed ensemble (parallel 10, seeds `42,7,2026,99,123,17,3,2020,88,55` — the recorded config artifact is authoritative over the trace's prose note) raised the rank metrics: RankIC 0.0671, RankICIR 0.259, L/S Sharpe 4.58 (vs 5-seed 4.54). But the book stayed **negative net of cost**: −0.93%, IR −0.089, maxDD −8.80%. `PROVEN — EVIDENCE#023 → exp 34`.
## The verdict on seed count
More seeds buy a small, real improvement in signal breadth — the RankICIR nudges up and the long-short Sharpe ticks up — but the added breadth **does not cross the cost barrier** (ch. 09). At 5 seeds the ensemble benefit has already done its work; 10 seeds add breadth without changing the construction's economics. Seed count is load-bearing up to ~5 and asymptotically irrelevant beyond, on this panel and cost model. `PROVEN — EVIDENCE#016/023`.
## Desk rules distilled from this chapter
1. Use a small multi-seed ensemble (3–5) as the standard, not a single model — the 2→5 step is the reproducible gain.
2. Do not chase seed count past the point of signal saturation; breadth past ~5 seeds did not pay net of cost.
3. Verify the recorded config artifact for seeds/parallel — the trace prose note disagreed with the YAML for Q02; the config artifact is authoritative.
4. `TODO(evidence-needed: whether the 10-seed breadth improves the *weekly-rebalance* construction (ch. 08), where cost is not the bottleneck)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#023` | exp 34 (Q02), run `ce49e4e0…`, branch `exp/34-q02-seed10-10-seed-rankicensemble-vs-ref` |
| `EVIDENCE#016` | exp 28, run `c4ab1d01…`, branch `exp/28-isolate-the-seed-count-effect-on-the-ndr` |
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
+58
View File
@@ -0,0 +1,58 @@
# Chapter 07 — Isolation Runs: Single-Variable Discipline
Status: drafting. Claim inventory: see `README.md` ch. 07.
The research loop's discipline (ch. 02) is that one variable changes per run. This chapter runs that discipline across the clean-lake campaign and the Q-series (exp 33–43), and shows that most additions fail — the discipline is the value, not the win rate.
## The design
The reference is the compact stochastic set, 5-seed RankIC ensemble, topk=10, n_drop=1, 5-day forward label, 50-ETF panel, train 2016-01-04→2025-09-01 / valid →2026-01-03 / test 2026-01-04→2026-08-10, account $1M, benchmark SPY, cost 5bp open / 15bp close / $5 min. Acceptance bar (from the reference): **net_IR ≥ 0.21, net_ann ≥ +2.13%, maxDD ≤ 7.69%** `(PROVEN — exp 24/26 baseline)`. Every Q-run changed exactly one thing against this reference.
| Q | One variable changed | Verdict |
|---|---------------------|---------|
| Q01 (exp 33) | features +1: `sp_sharpe_22` | PASS — reproduces M2 |
| Q02 (exp 34) | seeds 5→10 (parallel 10) | FAIL — signal up, book negative |
| Q03 (exp 35) | topk 10→20 | FAIL — no edge, lower vol |
| Q04 (exp 36) | label 5d→10d | FAIL — IC up, net collapses |
| Q05 (exp 37) | label 5d→22d | FAIL — best IC, flat gross |
| Q06 (exp 38) | sizing → fractional Kelly (cap 0.5) | FAIL — below bar |
| Q07 (exp 39) | rebalance daily→weekly | PASS — campaign best |
| Q08 (exp 40) | risk-limit gates on | FAIL as alpha (safety net) |
| Q09 (exp 41) | construction → long-short | FAIL — turnover kills |
| Q10 (exp 42) | regime entry gate on | FAIL — churns |
| Q11 (exp 43) | features → single `sp_trend_slope_5` | FAIL — no reversal learned |
`PROVEN — EVIDENCE#022–032 → exp 33–43, all pre-registered in trace start + workflow YAML before each run`.
## What isolation bought
Because each Q-run changed one thing, the verdicts attribute cleanly:
- **Feature axis (Q01, Q11):** adding the risk-adjusted Sharpe-drift feature is reproducible and positive `(EVIDENCE#022 → exp 33, Q01)`; stripping to a single mean-reversion feature is not learnable — the model trained *positive* IC (+0.0023), meaning there is no standalone reversal to find in the pooled cross-section `(EVIDENCE#032 → exp 43, Q11)`. The "mean reversion is the stable single-feature edge" hypothesis is **refuted** on the clean lake.
- **Label axis (Q04, Q05):** longer forward-return labels monotonically *improve* the signal — 10d: IC 0.0925 / RankIC 0.0960; 22d: IC 0.0970 / RankIC 0.1165 — yet net-of-cost performance *worsens* (10d: −9.92% IR −1.15; 22d: −4.60% IR −0.59). `PROVEN — EVIDENCE#025/026 → exp 36/37`. Horizon signal and daily-turnover construction are incompatible.
- **Model axis (Q02):** 10 seeds raise rank breadth (RankIC 0.0671, L/S Sharpe 4.58) but the book stays negative net (−0.93%) — the added breadth never crosses the cost barrier. `PROVEN — EVIDENCE#023 → exp 34`.
- **Construction axis (Q03, Q06, Q07, Q09):** see ch. 08 — weekly recompute (Q07) is the only change that clears the bar by a wide margin.
- **Risk/gate axis (Q08, Q10):** see ch. 10 — both met at most a drawdown leg; neither adds alpha.
## The discipline is the output
Only 2 of 11 Q-runs passed. That is not a failure of the campaign — it is the mechanism doing its job. Each FAIL closed a candidate direction at the cost of one run, and the two PASSes (Q01 reproducing the M2 feature, Q07 the weekly construction) are the campaign's forward path. The campaign's refuted runs were as valuable as its wins: knowing that a 22d label has an IC of 0.097 *and still loses money daily* is exactly the kind of fact a desk must not learn twice. `REFERENCED (falsification) + PROVEN (recorded negatives) — EVIDENCE#022–032`.
## Desk rules distilled from this chapter
1. Fix the reference and the acceptance bar *before* the series; change one variable per run.
2. Record signal metrics and net-of-cost metrics side by side — a signal gain that does not clear costs is not a strategy gain (Q02, Q04, Q05).
3. A single-feature "obvious" edge must be tested standalone before being trusted in a bundle (Q11 refuted it).
4. `TODO(evidence-needed: reproduce Q07 weekly rebalance on a second window, and Q01's sp_sharpe_22 in a live round)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#022` | exp 33 (Q01), run `c7c12228…`, branch `exp/33-q01-m2-reproduction-add-spsharpe22-to-th` |
| `EVIDENCE#023` | exp 34 (Q02), run `ce49e4e0…`, branch `exp/34-q02-seed10-10-seed-rankicensemble-vs-ref` |
| `EVIDENCE#024` | exp 35 (Q03), run `2a844c02…`, branch `exp/35-q03-topk20-widen-topkdropout-portfolio-f` |
| `EVIDENCE#025` | exp 36 (Q04), run `ef211826…`, branch `exp/36-q04-label10d-10d-forward-return-label-vs` |
| `EVIDENCE#026` | exp 37 (Q05), run `daad5042…`, branch `exp/37-q05-label22d-22d-forward-return-label-vs` |
| `EVIDENCE#032` | exp 43 (Q11), run `e859adfe…`, branch `exp/43-q11-standalone-5d-reversal-single-featur` |
| reference | exp 24/26 (compact set / n_drop=1), runs `fe469a19…`/`21afc6af…` |
@@ -0,0 +1,56 @@
# Chapter 08 — Portfolio Construction: Dropout, Sizing, and Cadence
Status: drafting. Claim inventory: see `README.md` ch. 08.
This chapter asks how a given signal should be turned into a book. The answer the TradeAC campaign converged on is that **construction is the performance lever** — more than features, more than seeds — and the best construction found on the clean lake is weekly recompute of a daily signal.
## The construction space tested
All runs share the compact stochastic signal, 5-seed ensemble, 5d label, $1M / SPY benchmark / 5bp·15bp·$5 costs. Only the construction varies:
| Construction | Run | Net ann | Net IR | MaxDD | Gross ann | Note |
|--------------|-----|---------|--------|-------|-----------|------|
| Topk10 n_drop1, daily (reference) | exp 26 | +2.13% | +0.21 | −7.69% | +7.02% | baseline |
| Topk20, daily | Q03 (exp 35) | −1.88% | −0.253 | −8.80% | +0.64% | wider, no edge |
| Fractional-Kelly (cap 0.5) | Q06 (exp 38) | +1.04% | +0.112 | −7.13% | +5.43% | sizing, below bar |
| **Weekly recompute** | **Q07 (exp 39)** | **+12.51%** | **+1.243** | **−4.13%** | **+13.59%** | **wins chapter** |
| Long-short top10/bottom10, daily | Q09 (exp 41) | −8.38% | −0.834 | −11.22% | +6.57% | $96.7k cost |
`PROVEN — EVIDENCE#024/027/028/030 → exp 35/38/39/41`.
## Weekly recompute: the campaign's best result
The reference strategy recomputes the topk book **daily** from the same 5-day-label predictions. Q07 kept the signal, topk, n_drop, and risk_degree identical and changed only the rebalance cadence to **weekly** (ISO-week, recompute topk from the freshest score each week). Result:
- net +12.51% (IR **1.24**) vs +2.13% (IR 0.21) daily;
- maxDD −4.13% vs −7.69%;
- cost drag collapsed to ~1.1pp (gross +13.59% → net +12.51%), versus the ~5–9pp drags that dominated every daily construction;
- signal metrics byte-identical to exp 26 (IC 0.0502, RankIC 0.0660).
`PROVEN — EVIDENCE#028 → exp 39`. The prediction is a 5-day-ahead cross-sectional rank; holding it weekly instead of churning it daily lets the edge survive the 20bp round-trip. This is the strongest single construction result in the book — `TODO(evidence-needed: reproduce on a second window, then take to a live round)`.
## What failed, and why
- **Wider book (Q03):** topk 10→20 halves per-name size and cuts book vol (std 0.0048 vs 0.0065) but adds no edge net of cost (−1.88%). Spreading the same signal thinner does not create value.
- **Fractional Kelly (Q06):** sizing by score magnitude at half-Kelly (cap_frac 0.5) turned the negative daily book mildly positive (+1.04%, IR 0.11) and trimmed maxDD to −7.13% — but it is a weak paste-over of the turnover problem, not a fix, and lands far below the 0.21 acceptance bar.
- **Long-short (Q09):** the top10/bottom10 market-neutral construction has a genuine *pre-cost* edge (gross +6.57%, IR 0.656) — the signal does rank longs over shorts — but daily long-short turnover is prohibitive: **total cost $96,721 ≈ 9.7% of a $1M book**, 2485 trades in ~150 days, fill rate 0.40, net −8.38%. `PROVEN — EVIDENCE#030 → exp 41`. The same weekly cadence that fixed Q07 was deliberately *not* applied here; the pair is a controlled comparison of cadence on the same signal family.
Pre-clean-lake context: stochastic-control OptimalStopControl constructions (exp 13/14) bled ~11pp to cost — the same turnover mechanism, different strategy class. Those are idea material only. `HYPOTHESIS (idea: pre-clean-lake) — EVIDENCE#006/007`.
## Desk rules distilled from this chapter
1. Construction is a first-class lever: identical signal, +10pp of net annual difference between daily and weekly recompute (exp 26 vs 39).
2. Before changing the signal, ask whether turnover is the binding constraint — weekly cadence buys more than most feature additions.
3. Market-neutral structures are only worth the cost if the long-short spread clears two-sided turnover; on this panel it does not.
4. `TODO(evidence-needed: weekly + long-short combination — the pre-cost edge of Q09 may clear costs at weekly cadence)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#028` | exp 39 (Q07), run `eb38588c…`, branch `exp/39-q07-weekly-rebalance-recompute-topkdropo` |
| `EVIDENCE#024` | exp 35 (Q03), run `2a844c02…`, branch `exp/35-q03-topk20-widen-topkdropout-portfolio-f` |
| `EVIDENCE#027` | exp 38 (Q06), run `afca4b80…`, branch `exp/38-q06-kelly-sizing-score-magnitude-fractio` |
| `EVIDENCE#030` | exp 41 (Q09), run `0647eadd…`, branch `exp/41-q09-long-short-market-neutral-long-top-1` |
| reference | exp 26, run `21afc6af…`, branch `exp/26-test-whether-reducing-topkdropout-daily` |
| pre-clean idea | exp 13/14, EVIDENCE#006/007 |
@@ -0,0 +1,49 @@
# Chapter 09 — The Cost/Turnover Frontier
Status: drafting. Claim inventory: see `README.md` ch. 09.
This chapter is the empirical core of the book's cost argument: **turnover, not signal, is the binding constraint.** Ch. 03 established the gross→net collapse on the reference. Ch. 08 showed the fix. This chapter quantifies the frontier — what turnover costs at 20bp round-trips and what the trade-off looks like when you cut it.
## The frontier on the clean lake
The 50-ETF panel, $1M book, 5bp open / 15bp close / $5 minimum. The same underlying signal (compact stochastic set, 5-seed ensemble, 5d label — IC 0.050, RankIC 0.066) expressed at different turnover levels:
| Construction | Turnover character | Cost drag | Net ann | Net IR | Source |
|--------------|--------------------|-----------|---------|--------|--------|
| daily topk10, n_drop 2 | daily forced replacement | ~9–10pp | −3.21% | −0.32 | exp 24 |
| daily topk10, n_drop 1 | daily, hold dropped name | ~5pp | +2.13% | +0.21 | exp 26 |
| **weekly recompute** | **weekly refresh** | **~1.1pp** | **+12.51%** | **+1.24** | **exp 39 (Q07)** |
| daily long-short top10/b10 | two-sided daily | ~9.7% of NAV | −8.38% | −0.83 | exp 41 (Q09) |
`PROVEN — EVIDENCE#015 (exp 26), #028 (exp 39), #030 (exp 41)`.
The n_drop 2→1 step (exp 26) already showed the mechanism with byte-identical signal metrics — the entire net gain was cost relief `(EVIDENCE#015)`. The weekly step (Q07) went further: same signal, same topk/n_drop, cadence only, and cost drag fell to ~1.1pp while net went to +12.51%.
## The long-short lesson
Q09 is the cleanest demonstration that cost, not signal, is the frontier: the long-short construction had a *positive* pre-cost excess (+6.57%, IR 0.656) — the signal genuinely separates longs from shorts — yet cost **$96,721 ≈ 9.7% of NAV** in ~150 days (2485 trades, fill rate 0.40) and net was −8.38%. `PROVEN — EVIDENCE#030 → exp 41`. A construction that spends ~10% of the book annually on two-sided turnover cannot be rescued by signal alone.
## Where the frontier bends
- **Cadence (proven).** Weekly recompute of a 5-day signal is the single biggest lever found: ~1.1pp drag, +12.51% net. `PROVEN — EVIDENCE#028 → exp 39`.
- **Dropped-name policy (proven).** n_drop 2→1 (hold, don't re-trade) bought ~5pp. `PROVEN — EVIDENCE#015 → exp 26`.
- **Sizing (weak).** Kelly-style sizing scaled exposure but did not change the turnover bill (Q06, +1.04% net). `PROVEN — EVIDENCE#027 → exp 38`.
- **Label horizon (counterintuitive).** Longer labels improve the *signal* monotonically (22d IC 0.097, RankIC 0.117) but *worsen net* under daily churn (Q05: −4.60%). The horizon gain is real but unmonetized. `PROVEN — EVIDENCE#026 → exp 37`. `TODO(evidence-needed: long-horizon label at weekly cadence — the combination is untested and is the book's most promising open cell)`.
## Desk rules distilled from this chapter
1. Compute cost drag as a share of NAV before believing any net number; at 20bp round-trips, 1% NAV per quarter is easy to spend.
2. Rank construction changes by cost drag first: cadence > dropped-name policy > sizing > gates.
3. Report gross and net side by side in every experiment; a positive-gross/negative-net run is a turnover problem, not a signal verdict.
4. `TODO(evidence-needed: realized-cost comparison of the weekly construction against the 5bp/15bp/$5 model once it trades live)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#015` | exp 26, run `21afc6af…`, branch `exp/26-test-whether-reducing-topkdropout-daily` |
| `EVIDENCE#028` | exp 39 (Q07), run `eb38588c…`, branch `exp/39-q07-weekly-rebalance-recompute-topkdropo` |
| `EVIDENCE#030` | exp 41 (Q09), run `0647eadd…`, branch `exp/41-q09-long-short-market-neutral-long-top-1` |
| `EVIDENCE#026` | exp 37 (Q05), run `daad5042…`, branch `exp/37-q05-label22d-22d-forward-return-label-vs` |
| `EVIDENCE#027` | exp 38 (Q06), run `afca4b80…`, branch `exp/38-q06-kelly-sizing-score-magnitude-fractio` |
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
+54
View File
@@ -0,0 +1,54 @@
# Chapter 10 — Risk Limits and Gates: Safety Net, Not Alpha
Status: drafting. Claim inventory: see `README.md` ch. 10.
Every desk wants to believe risk controls are a performance lever. On the TradeAC clean lake the evidence says otherwise: **risk limits are a safety net, and overlay gates mostly churn.** This chapter separates the two claims — what a liquidity floor does (defund) and what a regime gate does (churn) — on the post-reset signal.
## The post-reset A/B (Q08)
The reference pred (exp 26, `21afc6af…`) was run through `rd_risk_calibrate` with the live spec `{liquidity_floor_adv: $5M, size_cap_pct: 0.12, concentration_cap_pct: 0.95, drawdown_pause_pct: 0.10}` versus no limits, same window (2026-01-04→2026-08-10), topk10/n_drop1, SPY, $1M:
| Spec | ann return | IR | maxDD |
|------|-----------|-----|-------|
| baseline (no limits) | +27.50% | **1.5804** | −6.91% |
| candidate (5M floor + caps) | +2.20% | **1.5121** | **−0.65%** |
`PROVEN — EVIDENCE#029 → exp 40`. Read the columns carefully:
- **The floor binds.** The $5M ADV floor drops DBA, DBC, ESPO, FDN, REM, TAN, UNG, XAR (8 of 50 names) — it does real work on this panel.
- **No IR edge.** Candidate IR 1.5121 < baseline 1.5804. Gating does not improve the risk-adjusted return; the floor removes small-AVD names but the surviving book has no better rank.
- **The drawdown cut is pure defunding.** size_cap 0.12 × concentration_cap 0.95 folds the effective risk_degree to ≈ 0.0095 — about **$9.5k deployed of a $1M book**. maxDD falls to −0.65% because there is almost nothing at risk, not because risk was managed well.
The pre-clean-lake claim that "$5M liquidity floor improves IR 0.81→0.98" (exp 18) is **not reproduced** on the clean-lake signal. That number stays idea material `(EVIDENCE#008 → exp 18, pre-clean-lake)`. `PROVEN (refutation) — EVIDENCE#029 → exp 40`.
## The regime gate (Q10)
A HMM regime overlay (`sp_hmm_p_regime1 ≥ 0.5` entry gate, `RegimeGateDropoutStrategy`) on the same daily signal:
- net −4.26%, IR −0.382, maxDD −7.38% — meets the drawdown leg (7.38% < 7.69%) but far below the net-IR acceptance;
- the gate churned 276 trades in ~150 days; ~6.3pp of cost erased the +2.02% gross;
- signal metrics byte-identical to the reference (IC 0.0502, RankIC 0.0660).
`PROVEN — EVIDENCE#031 → exp 42`. A regime gate that flips exposure on a regime posterior priced into the features already just adds turnover. This clean-lake re-test refutes the "gates are a free drawdown cut" idea carried from exp 20 (pre-clean-lake, byte-identical no-ops there) `(EVIDENCE#009)`.
## The synthesis
- **Risk limits**: keep them as a live harness (the round-3 live round used the same spec and the funnel held — EVIDENCE#020), but never market them as alpha. On this signal they defund, not improve. `PROVEN — EVIDENCE#029/020`.
- **Gates**: regime/momentum overlays on top of features the model already sees add turnover, not edge. `PROVEN — EVIDENCE#031/009`.
- **Where risk does earn its keep**: as a *cap on damage*, not a return source. The drawdown pause and floor are the reason the live round stays disciplined; their value is the tail, not the mean. `REFERENCED (risk-management practice) + PROVEN (round-3 funnel held under the spec)`.
## Desk rules distilled from this chapter
1. A/B any risk-limit spec against no-limits on the same pred before shipping it; if IR does not improve, it is defunding.
2. Report deployed capital alongside maxDD — a smaller drawdown with 100x less risk is not a risk win.
3. Prefer limits that bind rarely but cap hard (liquidity floor, drawdown pause) over gates that churn every day (regime overlay).
4. `TODO(evidence-needed: a live round under the weekly-rebalance construction with the risk-limit spec, to confirm the safety-net behavior at higher deployed capital)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#029` | exp 40 (Q08), MLflow run `4667984187…` (exp `tac-rd-q08-risklimit`, id 43), branch `exp/40-q08-risk-limit-ab-on-exp-26-reference-si`, `book/data/evidence/q08-risklimit/risk_calibration.json` |
| `EVIDENCE#031` | exp 42 (Q10), run `436acd01…`, branch `exp/42-q10-hmm-regime-overlay-entry-gate-on-sph` |
| `EVIDENCE#020` | round 3, trace 27, branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac` |
| `EVIDENCE#009/008` | exp 20/18 (pre-clean-lake, idea material) |
+50
View File
@@ -0,0 +1,50 @@
# Chapter 12 — Synthesis: How Proved Truth Compounds
Status: drafting. Claim inventory: see `README.md` ch. 12.
This chapter is the scoreboard. It collects everything the campaign proved, in order of what actually moved performance — and why the Q-campaign's 9 refutations were as informative as its 2 passes.
## The scoreboard (clean lake, exp 21–43)
| Lever | Evidence | Net effect |
|-------|----------|-----------|
| **Data quality** | exp 21 (collapse), 22–24 (fix + revalidation) | The single largest swing: pre-reset +7.77% became −20.6% on the same config, then reappeared as a real signal. Everything before the reset is void. |
| **Cost relief / construction** | exp 26 (n_drop 2→1): −3.21%→+2.13%; exp 39 (weekly, Q07): **+12.51%, IR 1.24, maxDD −4.13%** | The dominant *positive* lever. Same signal, cadence changed. |
| **Feature pruning** | exp 9, 24, 25, 31, 33 (Q01), 43 (Q11) | Generic pruned set > model-specific; one warranted addition (sp_sharpe_22, Q01); standalone reversal refuted (Q11). |
| **Ensemble** | exp 28 (2<5 seeds), exp 34 (Q02: 10 seeds) | Real but bounded: breadth saturates ~5 seeds; extra seeds don't clear costs. |
| **Risk limits** | exp 40 (Q08) | Safety net only: floor binds, no IR edge, DD relief is defunding. |
| **Gates** | exp 20, 42 (Q10) | Refuted: regime overlay churns, adds cost, no edge. |
| **Sizing** | exp 38 (Q06) | Weak: half-Kelly mildly positive, below bar. |
| **Long-short** | exp 41 (Q09) | Refuted by turnover: pre-cost edge +6.6% destroyed by $96.7k cost. |
`PROVEN — EVIDENCE#010–032`.
## What the Q-campaign settled
Eleven pre-registered runs, two passes:
1. **Q01 PASS** — M2's risk-adjusted 22d Sharpe-drift feature reproduces on the compact set (net +6.53%, IR 0.62). Promotes exp-30's lone result from HYPOTHESIS to PROVEN. `EVIDENCE#022`.
2. **Q07 PASS** — weekly recompute is the campaign's best construction (net +12.51%, IR 1.24). The forward path. `EVIDENCE#028`.
3. **Nine refutations** — seed breadth (Q02), wider book (Q03), long labels under daily churn (Q04/Q05), Kelly sizing (Q06), risk-limit-as-alpha (Q08), long-short (Q09), regime gate (Q10), standalone reversal (Q11). Each closed a direction the desk had been considering, at one-run cost each. `EVIDENCE#023–027, 029–032`.
The refuted runs were as valuable as the passes: the label-horizon result (22d label, IC 0.097, yet net negative) is exactly the kind of counterintuitive fact a desk must not re-learn. `REFERENCED (falsification) + PROVEN (recorded negatives)`.
## The order of operations a reader should copy
1. **Fix data first** — re-validate the lake before any run (ch. 06).
2. **Attack turnover before signal** — cadence and dropped-name policy are the proven levers (ch. 09).
3. **Test features one at a time** against the reference (ch. 07); prune, don't add (ch. 04).
4. **Use a small ensemble** (5 seeds) and stop there (ch. 05).
5. **A/B risk limits** before shipping; keep them as a safety net (ch. 10).
6. **Reconcile live** — the funnel and slippage are the only claims that count (ch. 00/11).
## Open questions
- `TODO(evidence-needed: reproduce exp 39 weekly rebalance on a second window, then a live round)`
- `TODO(evidence-needed: long-horizon label at weekly cadence — the proven signal edge with the proven low-turnover construction)`
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic set)`
- `TODO(evidence-needed: realized-cost reconciliation of the weekly construction against the 5bp/15bp/$5 model once live)`
## Evidence cited in this chapter
Composite of `EVIDENCE#010–032`; see the per-chapter evidence tables (ch. 04–10) and `EVIDENCE.md` for run/branch level citations.