book: ch04 prune don't add (feature-family ablation) — exp 24/25/29/31, evidence #003-#019
This commit is contained in:
@@ -0,0 +1,41 @@
|
|||||||
|
# Chapter 04 — Prune, Don't Add: Feature-Family Ablation
|
||||||
|
|
||||||
|
Status: drafting. Claim inventory: see `README.md` ch. 04.
|
||||||
|
|
||||||
|
The campaign's strongest and most repeated finding is negative: on a ~50-name daily panel, **adding model-specific feature machinery regresses the signal, while pruning to a compact generic set improves it.** This chapter works out that finding — where it comes from, how it was proved on the clean lake, and the mechanism hypothesis behind it.
|
||||||
|
|
||||||
|
## The pruning thesis, in one sentence
|
||||||
|
|
||||||
|
`CSRankNorm + LightGBM` on a small cross-section rewards a small set of generic, scale-free, well-behaved statistics. Every time the desk added a family built for a specific model (OU mean-reversion, HMM regime, GARCH vol) or a dense family of raw derived fields (moment/volatility moments), the rank signal got worse.
|
||||||
|
|
||||||
|
## The ablation arc (pre-reset, idea material)
|
||||||
|
|
||||||
|
The first clean statement of the thesis came from feature-family ablation on the pre-reset lake: dropping the model-specific families (ou, hmm) and keeping the generic set (jump, har, trend, hurst, signature, ret, max_move) improved RankIC 0.030→0.064 and flipped net excess from −9.4% to +3.1% `(idea → exp 9, pre-clean-lake; EVIDENCE#003)`. The mirror-image run confirmed the failure mode: adding 16 moment/volatility fields regressed every metric (RankIC 0.064→0.047, net excess −16.2%, IR −1.57) `(idea → exp 11, pre-clean-lake; EVIDENCE#004)`.
|
||||||
|
|
||||||
|
These two pre-reset runs are **ideas, not proof** — they ran on the dirty lake. Their thesis survived exactly because the clean-lake runs reproduced the same direction (below).
|
||||||
|
|
||||||
|
## The clean-lake confirmation
|
||||||
|
|
||||||
|
The clean-lake compact set is the proof: the reference signal is built from raw OHLCV plus `sp_ret`, jump, RV1/5/22, vol ratios, trend slopes, logp, Hurst and path-signature L1/L2 — nothing model-specific `(PROVEN → exp 24, EVIDENCE#013)`. Then the isolation runs (ch. 07) proved the negative side on clean data:
|
||||||
|
|
||||||
|
- Adding `sp_ou_zscore` regresses the reference: IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25, EVIDENCE#014)`.
|
||||||
|
- Adding the multi-horizon momentum bundle (M1) regresses it further: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) `(PROVEN → exp 29, EVIDENCE#017)`.
|
||||||
|
- Adding the GARCH(1,1) vol-regime trio adds nothing: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31, EVIDENCE#019)`.
|
||||||
|
|
||||||
|
Three independent clean-lake additions, three failures, all against the same reference and the same metrics. This is the campaign's most reproduced empirical result.
|
||||||
|
|
||||||
|
## Mechanism hypotheses
|
||||||
|
|
||||||
|
Why does pruning win on this panel? Three hypotheses, none yet isolated (all `HYPOTHESIS`):
|
||||||
|
|
||||||
|
1. **Panel width vs feature count.** ~50 names gives ~50 cross-sectional observations per day. Each added feature is another column the model can split on; with so few observations, extra dimensions mostly fit noise that CSRankNorm then amplifies. The minimal generic set "won repeatedly" across three independent expansions (ou/hmm, realized moments, TA) `(HYPOTHESIS → chat-ideas.md)`.
|
||||||
|
2. **Scale-free is necessary but not sufficient.** Features that survive CSRankNorm are scale-free; dense raw fields (moments) are not, and were regressed. But scale-freeness alone doesn't guarantee usefulness — the moments family was still dropped `(HYPOTHESIS → chat-ideas.md)`.
|
||||||
|
3. **Single-feature strength ≠ marginal contribution.** The OU z-score was the strongest standalone time-series predictor (IC −0.15/−0.13) yet degraded the model — the "OU paradox" named in ch. 01 `(HYPOTHESIS → chat-ideas.md; TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`.
|
||||||
|
|
||||||
|
## What practice generally does (and why this differs)
|
||||||
|
|
||||||
|
Industry panels (thousands of names, cross-sectional breadth) routinely feed hundreds of features and let the model prune them. This campaign's panel is 50 ETFs — a breadth-limited, high-correlation cross-section where the dominant signal is common (ch. 01: drift is mostly market-wide). The desk's lesson is not "features are bad"; it is that **feature count must scale with cross-sectional breadth**, and on this breadth the marginal value of the next family is negative. That is a hypothesis about generality `TODO(evidence-needed: out-of-panel universe with wider breadth)`.
|
||||||
|
|
||||||
|
## The operational rule the book carries forward
|
||||||
|
|
||||||
|
Before adding any feature family to the reference, run it as an isolation experiment against the reference metrics (ch. 07). The default assumption is failure; the run must beat IC/RankIC **and** net IR on the clean lake to earn its place. This rule is what made exp 29 and exp 31 cheap negatives instead of silent regressions.
|
||||||
Reference in New Issue
Block a user