- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches - CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims - README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign - new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated - Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2.5 KiB
2.5 KiB
Chapter 05 — Ensembles and the Seed-Count Effect
Status: drafting. Claim inventory: see README.md ch. 05.
TradeAC's reference signal is a 5-seed RankIC ensemble of LightGBM rankers. This chapter records what seed count is worth — and, from the Q-campaign, what it is not worth.
What averaging buys (and its limits)
- 2 < 5 seeds (proven): on the clean lake the 2-seed ensemble lost to the 5-seed on every rank metric and net (RankIC 0.0579 vs 0.0663; net −1.49% IR −0.14 vs +2.13% IR +0.21).
PROVEN — EVIDENCE#016 → exp 28. - 5 seeds = the reference (compact stochastic set): RankIC 0.0663, RankICIR 0.2545.
PROVEN — EVIDENCE#013 → exp 24. - 5→10 seeds (new, Q02): the Q-campaign tested whether more breadth keeps paying. The 10-seed ensemble (parallel 10, seeds
42,7,2026,99,123,17,3,2020,88,55— the recorded config artifact is authoritative over the trace's prose note) raised the rank metrics: RankIC 0.0671, RankICIR 0.259, L/S Sharpe 4.58 (vs 5-seed 4.54). But the book stayed negative net of cost: −0.93%, IR −0.089, maxDD −8.80%.PROVEN — EVIDENCE#023 → exp 34.
The verdict on seed count
More seeds buy a small, real improvement in signal breadth — the RankICIR nudges up and the long-short Sharpe ticks up — but the added breadth does not cross the cost barrier (ch. 09). At 5 seeds the ensemble benefit has already done its work; 10 seeds add breadth without changing the construction's economics. Seed count is load-bearing up to ~5 and asymptotically irrelevant beyond, on this panel and cost model. PROVEN — EVIDENCE#016/023.
Desk rules distilled from this chapter
- Use a small multi-seed ensemble (3–5) as the standard, not a single model — the 2→5 step is the reproducible gain.
- Do not chase seed count past the point of signal saturation; breadth past ~5 seeds did not pay net of cost.
- Verify the recorded config artifact for seeds/parallel — the trace prose note disagreed with the YAML for Q02; the config artifact is authoritative.
TODO(evidence-needed: whether the 10-seed breadth improves the *weekly-rebalance* construction (ch. 08), where cost is not the bottleneck)
Evidence cited in this chapter
| Tag | Source |
|---|---|
EVIDENCE#023 |
exp 34 (Q02), run ce49e4e0…, branch exp/34-q02-seed10-10-seed-rankicensemble-vs-ref |
EVIDENCE#016 |
exp 28, run c4ab1d01…, branch exp/28-isolate-the-seed-count-effect-on-the-ndr |
EVIDENCE#013 |
exp 24, run fe469a19…, branch exp/24-run-the-rankic-ensemble-in-mlflow-experi |