Files
tac-exp-dev/book/chapters/04-prune-dont-add.md
T
zhaoli 06fb1e8ee9 book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2026-08-20 01:23:21 +00:00

36 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 04 — Prune, Don't Add: Feature-Family Ablation
Status: drafting. Claim inventory: see `README.md` ch. 04.
The first instinct of a quant desk with a weak signal is to add features. On the TradeAC 50-ETF panel the evidence runs the other way: **the signal that survived the clean-lake reset is a *pruned* set, and the single-feature additions that "should" work mostly did not.** This chapter collects the ablation evidence and its clean-lake re-tests.
## The pattern: generic beats specific
The pre-clean-lake ablation (exp 9) established the direction: dropping model-specific families (ou, hmm) and keeping the general stochastic families improved the rank signal (RankIC 0.030→0.064). Adding moment/volatility families regressed it (exp 11). Those runs are pre-reset idea material, but the *direction* re-proved itself on clean data: OU reversion hurts (exp 25, `EVIDENCE#014`), GARCH adds nothing (exp 31, `EVIDENCE#019`), and the compact general stochastic set is the reference (exp 24, `EVIDENCE#013`). `HYPOTHESIS (pre-clean-lake) → PROVEN (clean-lake direction, exp 24/25/31)`.
## Two clean-lake additions, one prune (Q01, Q11)
The Q-campaign tested the two extremes of the feature axis against the reference:
- **Add (Q01):** `sp_sharpe_22` — the 22-day risk-adjusted Sharpe drift from exp 30's M2 — reproduced exactly on the compact set: IC 0.0464, RankIC 0.0578, net +6.53% (IR 0.62), maxDD −8.0%. This is a **warranted addition**: it carries the M2 edge into the pruned set. `PROVEN — EVIDENCE#022 → exp 33`.
- **Prune to one (Q11):** the single "obvious" mean-reversion feature `sp_trend_slope_5` alone (plus raw OHLCV) — the hypothesis was that a standalone 5d reversal exists. The model trained **positive** IC (+0.0023), so it did not learn reversal at all; gross −10.4%, net −15.2%. `PROVEN — EVIDENCE#032 → exp 43`. The "mean reversion is the stable single-feature edge" claim is refuted on the clean lake; the pooled trend-slope reversal beta does not survive as a standalone.
The lesson is the pair taken together: on a 50-name daily panel, one principled addition (Sharpe drift) helped and reproduced, while the "obvious" reversal feature did not exist standalone. Feature decisions need isolation runs, not intuition (ch. 07).
## What this means for a reader
- Treat "more features" as a hypothesis, tested one at a time against the reference.
- Prefer scale-free general statistics (volatility, jump, trend, signature) over model-specific machinery (HMM/OU states) on a small cross-section.
- A feature that fails as a bundle member is not necessarily dead (OU); a feature that fails standalone is not necessarily live in a bundle — test both directions (Q01 = bundle→isolated-addition PASS, Q11 = standalone FAIL).
- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#022` | exp 33 (Q01), run `c7c12228…`, branch `exp/33-q01-m2-reproduction-add-spsharpe22-to-th` |
| `EVIDENCE#032` | exp 43 (Q11), run `e859adfe…`, branch `exp/43-q11-standalone-5d-reversal-single-featur` |
| `EVIDENCE#014` | exp 25, run `57450d1a…`, branch `exp/25-test-the-clean-data-hypothesis-that-addi` |
| `EVIDENCE#019` | exp 31, run `514cb523…`, branch `exp/31-isolation-run-m3-does-adding-garch11-vol` |
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
| pre-clean ideas | exp 9/11, EVIDENCE#003/004 |