Files
book-tac/book/chapters/04-prune-dont-add.md
T

4.6 KiB
Raw Blame History

Chapter 04 — Prune, Don't Add: Feature-Family Ablation

Status: drafting. Claim inventory: see README.md ch. 04.

The campaign's strongest and most repeated finding is negative: on a ~50-name daily panel, adding model-specific feature machinery regresses the signal, while pruning to a compact generic set improves it. This chapter works out that finding — where it comes from, how it was proved on the clean lake, and the mechanism hypothesis behind it.

The pruning thesis, in one sentence

CSRankNorm + LightGBM on a small cross-section rewards a small set of generic, scale-free, well-behaved statistics. Every time the desk added a family built for a specific model (OU mean-reversion, HMM regime, GARCH vol) or a dense family of raw derived fields (moment/volatility moments), the rank signal got worse.

The ablation arc (pre-reset, idea material)

The first clean statement of the thesis came from feature-family ablation on the pre-reset lake: dropping the model-specific families (ou, hmm) and keeping the generic set (jump, har, trend, hurst, signature, ret, max_move) improved RankIC 0.030→0.064 and flipped net excess from −9.4% to +3.1% (idea → exp 9, pre-clean-lake; EVIDENCE#003). The mirror-image run confirmed the failure mode: adding 16 moment/volatility fields regressed every metric (RankIC 0.064→0.047, net excess −16.2%, IR −1.57) (idea → exp 11, pre-clean-lake; EVIDENCE#004).

These two pre-reset runs are ideas, not proof — they ran on the dirty lake. Their thesis survived exactly because the clean-lake runs reproduced the same direction (below).

The clean-lake confirmation

The clean-lake compact set is the proof: the reference signal is built from raw OHLCV plus sp_ret, jump, RV1/5/22, vol ratios, trend slopes, logp, Hurst and path-signature L1/L2 — nothing model-specific (PROVEN → exp 24, EVIDENCE#013). Then the isolation runs (ch. 07) proved the negative side on clean data:

  • Adding sp_ou_zscore regresses the reference: IC 0.0343 vs 0.0511, net −3.76% vs −3.21% (PROVEN → exp 25, EVIDENCE#014).
  • Adding the multi-horizon momentum bundle (M1) regresses it further: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) (PROVEN → exp 29, EVIDENCE#017).
  • Adding the GARCH(1,1) vol-regime trio adds nothing: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) (PROVEN → exp 31, EVIDENCE#019).

Three independent clean-lake additions, three failures, all against the same reference and the same metrics. This is the campaign's most reproduced empirical result.

Mechanism hypotheses

Why does pruning win on this panel? Three hypotheses, none yet isolated (all HYPOTHESIS):

  1. Panel width vs feature count. ~50 names gives ~50 cross-sectional observations per day. Each added feature is another column the model can split on; with so few observations, extra dimensions mostly fit noise that CSRankNorm then amplifies. The minimal generic set "won repeatedly" across three independent expansions (ou/hmm, realized moments, TA) (HYPOTHESIS → chat-ideas.md).
  2. Scale-free is necessary but not sufficient. Features that survive CSRankNorm are scale-free; dense raw fields (moments) are not, and were regressed. But scale-freeness alone doesn't guarantee usefulness — the moments family was still dropped (HYPOTHESIS → chat-ideas.md).
  3. Single-feature strength ≠ marginal contribution. The OU z-score was the strongest standalone time-series predictor (IC −0.15/−0.13) yet degraded the model — the "OU paradox" named in ch. 01 (HYPOTHESIS → chat-ideas.md; TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM)).

What practice generally does (and why this differs)

Industry panels (thousands of names, cross-sectional breadth) routinely feed hundreds of features and let the model prune them. This campaign's panel is 50 ETFs — a breadth-limited, high-correlation cross-section where the dominant signal is common (ch. 01: drift is mostly market-wide). The desk's lesson is not "features are bad"; it is that feature count must scale with cross-sectional breadth, and on this breadth the marginal value of the next family is negative. That is a hypothesis about generality TODO(evidence-needed: out-of-panel universe with wider breadth).

The operational rule the book carries forward

Before adding any feature family to the reference, run it as an isolation experiment against the reference metrics (ch. 07). The default assumption is failure; the run must beat IC/RankIC and net IR on the clean lake to earn its place. This rule is what made exp 29 and exp 31 cheap negatives instead of silent regressions.