book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters

- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
This commit is contained in:
zhaoli
2026-08-20 01:23:21 +00:00
parent 436692a620
commit 06fb1e8ee9
15 changed files with 847 additions and 36 deletions
+30
View File
@@ -0,0 +1,30 @@
# Chapter 05 — Ensembles and the Seed-Count Effect
Status: drafting. Claim inventory: see `README.md` ch. 05.
TradeAC's reference signal is a **5-seed RankIC ensemble** of LightGBM rankers. This chapter records what seed count is worth — and, from the Q-campaign, what it is *not* worth.
## What averaging buys (and its limits)
- **2 < 5 seeds (proven):** on the clean lake the 2-seed ensemble lost to the 5-seed on every rank metric and net (RankIC 0.0579 vs 0.0663; net −1.49% IR −0.14 vs +2.13% IR +0.21). `PROVEN — EVIDENCE#016 → exp 28`.
- **5 seeds = the reference** (compact stochastic set): RankIC 0.0663, RankICIR 0.2545. `PROVEN — EVIDENCE#013 → exp 24`.
- **5→10 seeds (new, Q02):** the Q-campaign tested whether more breadth keeps paying. The 10-seed ensemble (parallel 10, seeds `42,7,2026,99,123,17,3,2020,88,55` — the recorded config artifact is authoritative over the trace's prose note) raised the rank metrics: RankIC 0.0671, RankICIR 0.259, L/S Sharpe 4.58 (vs 5-seed 4.54). But the book stayed **negative net of cost**: −0.93%, IR −0.089, maxDD −8.80%. `PROVEN — EVIDENCE#023 → exp 34`.
## The verdict on seed count
More seeds buy a small, real improvement in signal breadth — the RankICIR nudges up and the long-short Sharpe ticks up — but the added breadth **does not cross the cost barrier** (ch. 09). At 5 seeds the ensemble benefit has already done its work; 10 seeds add breadth without changing the construction's economics. Seed count is load-bearing up to ~5 and asymptotically irrelevant beyond, on this panel and cost model. `PROVEN — EVIDENCE#016/023`.
## Desk rules distilled from this chapter
1. Use a small multi-seed ensemble (3–5) as the standard, not a single model — the 2→5 step is the reproducible gain.
2. Do not chase seed count past the point of signal saturation; breadth past ~5 seeds did not pay net of cost.
3. Verify the recorded config artifact for seeds/parallel — the trace prose note disagreed with the YAML for Q02; the config artifact is authoritative.
4. `TODO(evidence-needed: whether the 10-seed breadth improves the *weekly-rebalance* construction (ch. 08), where cost is not the bottleneck)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#023` | exp 34 (Q02), run `ce49e4e0…`, branch `exp/34-q02-seed10-10-seed-rankicensemble-vs-ref` |
| `EVIDENCE#016` | exp 28, run `c4ab1d01…`, branch `exp/28-isolate-the-seed-count-effect-on-the-ndr` |
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |