Files
book-tac/book/chapters/05-ensembles-seed-count.md
T

31 lines
3.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 05 — Ensembles and the Seed-Count Effect
Status: drafting. Claim inventory: see `README.md` ch. 05.
Ensemble averaging is the campaign's one *additive* lever that survived clean-data scrutiny. This chapter separates what the ensemble does (variance reduction on a noisy rank) from what it does not do (add information), and shows that the *count* of seeds is load-bearing.
## What the ensemble is
The reference model is a 5-seed LightGBM blend: five models, differing only by random seed, trained on the same features and label, averaged into one prediction. The mechanism is reinforcement against estimation noise — the same metric a noise-dominated signal needs most (ch. 01, decay/reinforcement).
## The evidence
- The 5-seed ensemble on the ablated generic features was the pre-reset best result (RankIC 0.0586, RankICIR 0.224, net excess +7.8%, IR 0.79) `(idea → exp 12, pre-clean-lake; EVIDENCE#005)`. Its numbers are inflated by the dirty lake (EVIDENCE#010) but its *design* — ensemble on ablated features, isolated from feature expansion — was re-validated on the clean lake.
- The clean-lake reference is the same design: the 5-seed RankIC ensemble on the compact stochastic set (RankIC 0.0663, RankICIR 0.2545) `(PROVEN → exp 24, EVIDENCE#013)`, reproduced from the same family lineage in exp 22/23 `(PROVEN → EVIDENCE#011/#012)`.
- **Seed count is load-bearing**: the 2-seed blend loses to the 5-seed blend on the identical compact set — RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21) `(PROVEN → exp 28, EVIDENCE#016)`. Fewer seeds is not "cheaper, same signal"; it is a measurably worse signal.
## Mechanism: variance reduction, not new information
Two observations pin the mechanism to variance reduction. First, the ensemble's metric benefit shows up most clearly in RankICIR/IR — the noise-adjusted ratios — rather than in raw IC, consistent with error cancellation `(PROVEN → exp 24 vs exp 22/23 schema; interpret as PROVEN direction, magnitude is window-specific)`. Second, the same features and label produce different outcomes by seed count alone, which means the marginal value of the 5th seed is *stability*: the model family is good enough that its remaining error is estimation variance, and averaging it away is the cheapest reliable win available `(HYPOTHESIS → chat-ideas.md: equal-weight seed blend > adaptive blending; the mechanism is not fully isolated)`.
## The open question
Does the 5-seed ensemble win by variance reduction or by diversifying model families (e.g. different effective trees/feature interactions per seed)? The two hypotheses make different predictions for a 10-seed run `TODO(evidence-needed: 10-seed vs 5-seed isolation; and whether seed-count benefit survives a wider panel)`. The desk has not yet run either `(open, chat-ideas.md)`.
## Interaction with pruning and cost
The ensemble amplifies a pruned signal — it is not a substitute for pruning (ch. 04) and it does not fix cost (ch. 09). The clean-lake sequence is explicit: the ensemble's net result (+2.13% IR 0.21) still barely clears costs; averaging improves the signal-to-noise ratio, and n_drop relief then converts that into net return `(PROVEN → exp 24 + exp 26, EVIDENCE#013/#015)`.
## Practice note
Equal-weight seed blending beat a rolling-IC adaptive blend in the campaign's design choices (pre-reset idea, untested head-to-head on clean data): adaptive weights re-fit to noise on a 50-name panel `(HYPOTHESIS → chat-ideas.md)`. The desk's rule of thumb: fix the seed count and the weights; spend experiment budget on pruning and cost relief, where the reproduced wins are.