diff --git a/book/chapters/05-ensembles-seed-count.md b/book/chapters/05-ensembles-seed-count.md new file mode 100644 index 0000000..ab648c8 --- /dev/null +++ b/book/chapters/05-ensembles-seed-count.md @@ -0,0 +1,31 @@ +# Chapter 05 — Ensembles and the Seed-Count Effect + +Status: drafting. Claim inventory: see `README.md` ch. 05. + +Ensemble averaging is the campaign's one *additive* lever that survived clean-data scrutiny. This chapter separates what the ensemble does (variance reduction on a noisy rank) from what it does not do (add information), and shows that the *count* of seeds is load-bearing. + +## What the ensemble is + +The reference model is a 5-seed LightGBM blend: five models, differing only by random seed, trained on the same features and label, averaged into one prediction. The mechanism is reinforcement against estimation noise — the same metric a noise-dominated signal needs most (ch. 01, decay/reinforcement). + +## The evidence + +- The 5-seed ensemble on the ablated generic features was the pre-reset best result (RankIC 0.0586, RankICIR 0.224, net excess +7.8%, IR 0.79) `(idea → exp 12, pre-clean-lake; EVIDENCE#005)`. Its numbers are inflated by the dirty lake (EVIDENCE#010) but its *design* — ensemble on ablated features, isolated from feature expansion — was re-validated on the clean lake. +- The clean-lake reference is the same design: the 5-seed RankIC ensemble on the compact stochastic set (RankIC 0.0663, RankICIR 0.2545) `(PROVEN → exp 24, EVIDENCE#013)`, reproduced from the same family lineage in exp 22/23 `(PROVEN → EVIDENCE#011/#012)`. +- **Seed count is load-bearing**: the 2-seed blend loses to the 5-seed blend on the identical compact set — RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21) `(PROVEN → exp 28, EVIDENCE#016)`. Fewer seeds is not "cheaper, same signal"; it is a measurably worse signal. + +## Mechanism: variance reduction, not new information + +Two observations pin the mechanism to variance reduction. First, the ensemble's metric benefit shows up most clearly in RankICIR/IR — the noise-adjusted ratios — rather than in raw IC, consistent with error cancellation `(PROVEN → exp 24 vs exp 22/23 schema; interpret as PROVEN direction, magnitude is window-specific)`. Second, the same features and label produce different outcomes by seed count alone, which means the marginal value of the 5th seed is *stability*: the model family is good enough that its remaining error is estimation variance, and averaging it away is the cheapest reliable win available `(HYPOTHESIS → chat-ideas.md: equal-weight seed blend > adaptive blending; the mechanism is not fully isolated)`. + +## The open question + +Does the 5-seed ensemble win by variance reduction or by diversifying model families (e.g. different effective trees/feature interactions per seed)? The two hypotheses make different predictions for a 10-seed run `TODO(evidence-needed: 10-seed vs 5-seed isolation; and whether seed-count benefit survives a wider panel)`. The desk has not yet run either `(open, chat-ideas.md)`. + +## Interaction with pruning and cost + +The ensemble amplifies a pruned signal — it is not a substitute for pruning (ch. 04) and it does not fix cost (ch. 09). The clean-lake sequence is explicit: the ensemble's net result (+2.13% IR 0.21) still barely clears costs; averaging improves the signal-to-noise ratio, and n_drop relief then converts that into net return `(PROVEN → exp 24 + exp 26, EVIDENCE#013/#015)`. + +## Practice note + +Equal-weight seed blending beat a rolling-IC adaptive blend in the campaign's design choices (pre-reset idea, untested head-to-head on clean data): adaptive weights re-fit to noise on a 50-name panel `(HYPOTHESIS → chat-ideas.md)`. The desk's rule of thumb: fix the seed count and the weights; spend experiment budget on pruning and cost relief, where the reproduced wins are. \ No newline at end of file