Compare commits

...
4 Commits
4 changed files with 157 additions and 0 deletions
+41
View File
@@ -0,0 +1,41 @@
# Chapter 04 — Prune, Don't Add: Feature-Family Ablation
Status: drafting. Claim inventory: see `README.md` ch. 04.
The campaign's strongest and most repeated finding is negative: on a ~50-name daily panel, **adding model-specific feature machinery regresses the signal, while pruning to a compact generic set improves it.** This chapter works out that finding — where it comes from, how it was proved on the clean lake, and the mechanism hypothesis behind it.
## The pruning thesis, in one sentence
`CSRankNorm + LightGBM` on a small cross-section rewards a small set of generic, scale-free, well-behaved statistics. Every time the desk added a family built for a specific model (OU mean-reversion, HMM regime, GARCH vol) or a dense family of raw derived fields (moment/volatility moments), the rank signal got worse.
## The ablation arc (pre-reset, idea material)
The first clean statement of the thesis came from feature-family ablation on the pre-reset lake: dropping the model-specific families (ou, hmm) and keeping the generic set (jump, har, trend, hurst, signature, ret, max_move) improved RankIC 0.030→0.064 and flipped net excess from −9.4% to +3.1% `(idea → exp 9, pre-clean-lake; EVIDENCE#003)`. The mirror-image run confirmed the failure mode: adding 16 moment/volatility fields regressed every metric (RankIC 0.064→0.047, net excess −16.2%, IR −1.57) `(idea → exp 11, pre-clean-lake; EVIDENCE#004)`.
These two pre-reset runs are **ideas, not proof** — they ran on the dirty lake. Their thesis survived exactly because the clean-lake runs reproduced the same direction (below).
## The clean-lake confirmation
The clean-lake compact set is the proof: the reference signal is built from raw OHLCV plus `sp_ret`, jump, RV1/5/22, vol ratios, trend slopes, logp, Hurst and path-signature L1/L2 — nothing model-specific `(PROVEN → exp 24, EVIDENCE#013)`. Then the isolation runs (ch. 07) proved the negative side on clean data:
- Adding `sp_ou_zscore` regresses the reference: IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25, EVIDENCE#014)`.
- Adding the multi-horizon momentum bundle (M1) regresses it further: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) `(PROVEN → exp 29, EVIDENCE#017)`.
- Adding the GARCH(1,1) vol-regime trio adds nothing: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31, EVIDENCE#019)`.
Three independent clean-lake additions, three failures, all against the same reference and the same metrics. This is the campaign's most reproduced empirical result.
## Mechanism hypotheses
Why does pruning win on this panel? Three hypotheses, none yet isolated (all `HYPOTHESIS`):
1. **Panel width vs feature count.** ~50 names gives ~50 cross-sectional observations per day. Each added feature is another column the model can split on; with so few observations, extra dimensions mostly fit noise that CSRankNorm then amplifies. The minimal generic set "won repeatedly" across three independent expansions (ou/hmm, realized moments, TA) `(HYPOTHESIS → chat-ideas.md)`.
2. **Scale-free is necessary but not sufficient.** Features that survive CSRankNorm are scale-free; dense raw fields (moments) are not, and were regressed. But scale-freeness alone doesn't guarantee usefulness — the moments family was still dropped `(HYPOTHESIS → chat-ideas.md)`.
3. **Single-feature strength ≠ marginal contribution.** The OU z-score was the strongest standalone time-series predictor (IC −0.15/−0.13) yet degraded the model — the "OU paradox" named in ch. 01 `(HYPOTHESIS → chat-ideas.md; TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`.
## What practice generally does (and why this differs)
Industry panels (thousands of names, cross-sectional breadth) routinely feed hundreds of features and let the model prune them. This campaign's panel is 50 ETFs — a breadth-limited, high-correlation cross-section where the dominant signal is common (ch. 01: drift is mostly market-wide). The desk's lesson is not "features are bad"; it is that **feature count must scale with cross-sectional breadth**, and on this breadth the marginal value of the next family is negative. That is a hypothesis about generality `TODO(evidence-needed: out-of-panel universe with wider breadth)`.
## The operational rule the book carries forward
Before adding any feature family to the reference, run it as an isolation experiment against the reference metrics (ch. 07). The default assumption is failure; the run must beat IC/RankIC **and** net IR on the clean lake to earn its place. This rule is what made exp 29 and exp 31 cheap negatives instead of silent regressions.
+31
View File
@@ -0,0 +1,31 @@
# Chapter 05 — Ensembles and the Seed-Count Effect
Status: drafting. Claim inventory: see `README.md` ch. 05.
Ensemble averaging is the campaign's one *additive* lever that survived clean-data scrutiny. This chapter separates what the ensemble does (variance reduction on a noisy rank) from what it does not do (add information), and shows that the *count* of seeds is load-bearing.
## What the ensemble is
The reference model is a 5-seed LightGBM blend: five models, differing only by random seed, trained on the same features and label, averaged into one prediction. The mechanism is reinforcement against estimation noise — the same metric a noise-dominated signal needs most (ch. 01, decay/reinforcement).
## The evidence
- The 5-seed ensemble on the ablated generic features was the pre-reset best result (RankIC 0.0586, RankICIR 0.224, net excess +7.8%, IR 0.79) `(idea → exp 12, pre-clean-lake; EVIDENCE#005)`. Its numbers are inflated by the dirty lake (EVIDENCE#010) but its *design* — ensemble on ablated features, isolated from feature expansion — was re-validated on the clean lake.
- The clean-lake reference is the same design: the 5-seed RankIC ensemble on the compact stochastic set (RankIC 0.0663, RankICIR 0.2545) `(PROVEN → exp 24, EVIDENCE#013)`, reproduced from the same family lineage in exp 22/23 `(PROVEN → EVIDENCE#011/#012)`.
- **Seed count is load-bearing**: the 2-seed blend loses to the 5-seed blend on the identical compact set — RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21) `(PROVEN → exp 28, EVIDENCE#016)`. Fewer seeds is not "cheaper, same signal"; it is a measurably worse signal.
## Mechanism: variance reduction, not new information
Two observations pin the mechanism to variance reduction. First, the ensemble's metric benefit shows up most clearly in RankICIR/IR — the noise-adjusted ratios — rather than in raw IC, consistent with error cancellation `(PROVEN → exp 24 vs exp 22/23 schema; interpret as PROVEN direction, magnitude is window-specific)`. Second, the same features and label produce different outcomes by seed count alone, which means the marginal value of the 5th seed is *stability*: the model family is good enough that its remaining error is estimation variance, and averaging it away is the cheapest reliable win available `(HYPOTHESIS → chat-ideas.md: equal-weight seed blend > adaptive blending; the mechanism is not fully isolated)`.
## The open question
Does the 5-seed ensemble win by variance reduction or by diversifying model families (e.g. different effective trees/feature interactions per seed)? The two hypotheses make different predictions for a 10-seed run `TODO(evidence-needed: 10-seed vs 5-seed isolation; and whether seed-count benefit survives a wider panel)`. The desk has not yet run either `(open, chat-ideas.md)`.
## Interaction with pruning and cost
The ensemble amplifies a pruned signal — it is not a substitute for pruning (ch. 04) and it does not fix cost (ch. 09). The clean-lake sequence is explicit: the ensemble's net result (+2.13% IR 0.21) still barely clears costs; averaging improves the signal-to-noise ratio, and n_drop relief then converts that into net return `(PROVEN → exp 24 + exp 26, EVIDENCE#013/#015)`.
## Practice note
Equal-weight seed blending beat a rolling-IC adaptive blend in the campaign's design choices (pre-reset idea, untested head-to-head on clean data): adaptive weights re-fit to noise on a 50-name panel `(HYPOTHESIS → chat-ideas.md)`. The desk's rule of thumb: fix the seed count and the weights; spend experiment budget on pruning and cost relief, where the reproduced wins are.
+46
View File
@@ -0,0 +1,46 @@
# Chapter 07 — Isolation Runs: Single-Variable Discipline
Status: drafting. Claim inventory: see `README.md` ch. 07.
The isolation run is the campaign's unit of proof: one variable changes against the fixed reference; the same metrics decide the verdict. This chapter walks the clean-lake sequence (exp 28–31) as the worked example, and states the discipline's rules so a reader can run it themselves.
## Why isolation is the discipline
The research loop (ch. 02) is only as honest as its attribution. A run that changes two things cannot say which one moved the result. The clean-lake campaign therefore fixed the reference — exp 26's n_drop=1 configuration on the compact stochastic set `(PROVEN → exp 26, EVIDENCE#015)` — and ran every subsequent experiment as a single-variable change against it:
| Run | One variable changed | Reference | Verdict |
|-----|----------------------|-----------|---------|
| exp 28 | seed count 5 → 2 | same features/book | REFUTED (2 seeds worse) `(PROVEN → EVIDENCE#016)` |
| exp 29 | + multi-horizon momentum bundle (M1) | same book | REFUTED `(PROVEN → EVIDENCE#017)` |
| exp 30 | + risk-adjusted 22d Sharpe drift (M2) | same book | HYPOTHESIS (mixed) `(PROVEN run, EVIDENCE#018)` |
| exp 31 | + GARCH(1,1) vol-regime trio (M3) | same book | REFUTED `(PROVEN → EVIDENCE#019)` |
Three of four refuted; one mixed. The isolation design is what makes "most additions fail" a *finding* rather than an anecdote.
## The acceptance contract
For an addition to earn its place it must beat the reference on **both** layers of the metrics ladder (ch. 01): the rank layer (IC/RankIC/ICIR/RankICIR) *and* the portfolio layer (net IR, MDD). Exp 30 is the canonical trap: M2 looked strong on the portfolio layer (net +6.53%, IR 0.62 vs +2.13%, IR 0.21) while its rank metrics were *lower* than reference (RankIC 0.0576 vs 0.0663) `(PROVEN → exp 30, EVIDENCE#018)`. Because the two layers disagreed and the run was not reproduced, the book labels it HYPOTHESIS rather than PROVEN. The rule: **a single run that improves one layer and degrades the other is a hypothesis, not a win** `TODO(evidence-needed: reproduction of exp 30 M2 on a second window)`.
## What each refutation taught
- **exp 28 (2 seeds):** the ensemble's value is tied to seed count; halving it is not a harmless cost cut (ch. 05).
- **exp 29 (momentum bundle):** drift-as-feature fails even when the underlying structure exists (ch. 01, ch. 04). The mechanism (name-specific scale, collision with trend features) is a hypothesis.
- **exp 31 (GARCH):** parametric vol modeling adds nothing to a model that already has the realized-vol ladder — the generic ladder is the feature; the parametric overlay is not (ch. 01).
- **exp 30 (M2):** the single interesting non-refutation. Risk-adjusting the drift feature changed the portfolio behavior without improving the rank — an unexplained, unreproduced anomaly worth one more run, not a claim.
## How to run an isolation campaign
1. **Freeze a reference** — a configuration, a book spec, and a metric table that everything is judged against (exp 26 n_drop=1 in this campaign).
2. **Change exactly one thing** per run; record the hypothesis and acceptance metric in the run notes *before* running (ch. 02, pre-registration).
3. **Judge on both layers** of the metrics ladder; a single-layer improvement is a hypothesis.
4. **Expect most runs to fail** — that is the point. A refuted run is a recorded negative that protects the next hypothesis from paying for the same mistake twice.
## The discipline as the value
The campaign's net-of-cost performance barely cleared costs at its best `(PROVEN → exp 26, EVIDENCE#015)`. In that regime, undisciplined feature-hunting is not neutral — it is the largest *expected* destroyer of the edge. Isolation runs converted feature-hunting into a bounded cost: a few runs to prove each family dead, instead of silently degrading the live signal. That is why the book counts exp 29 and exp 31 among its wins (ch. 12).
## Open questions
- `TODO(evidence-needed: M2 reproduction — the only surviving clean-lake improvement candidate)`
- `TODO(evidence-needed: 5d-reversal standalone strategy net of costs, the un-isolated idea from ch. 01)`
- `TODO(evidence-needed: HMM regime overlay (long-only gate) on the exp-26 book — an overlay test, not a feature test)`
@@ -0,0 +1,39 @@
# Chapter 11 — Live Execution and Reconciliation
Status: drafting. Claim inventory: see `README.md` ch. 11. HITL review gate applies: live performance numbers, cost/slippage figures.
The book's spine is that claims must be reconcilable (ch. 00). This chapter closes the loop: the live round that executed the proved reference signal, the funnel that held, the realized cost, and what reconciliation says about the backtest's honesty.
## The round
Round 3 (target date 2026-08-17) ran the exp-26 n_drop=1 configuration retrained on a rolling 4-year window, Topk10 with risk limits (liquidity floor $5M, size cap 12%, concentration cap 95%, drawdown pause 10%) `(PROVEN → round 3, EVIDENCE#020; trace 27, run 721ef257…, branch exp/27-scheduled-algo-retrain-on-2026-08-17)`.
## The funnel held
The execution funnel — targets → intents → decisions → placed → filled — closed at **10 → 10 → 10 → 9** `(PROVEN → round 3 funnel)`:
- 10 targets from the strategy's target list,
- 10 decided, 10 placed,
- 9 filled, 1 cancelled, and 1 skipped (SLV, `delta_zero` — the pre-skip gate stopped a zero-delta name).
A 90% fill-to-target ratio with one deliberate skip is a funnel that executed what the research claimed it would — the strategy's intent survived the gates and the broker. The reconciliation (targets vs decisions vs fills, per-symbol residuals) is available from the round's `book_reconcile`.
## Realized cost
Invested notional was **$74,202.85** with **realized slippage ≈ 4.54 bps** and estimated cost ≈ **$45**; turnover ≈ 0.74 `(PROVEN → round 3 metrics, EVIDENCE#020)`. The slippage figure is *realized* — taken from fills versus the expected execution price in the round's order trail — not a backtest assumption. This is the number the backtest cost model must be judged against.
## What reconciliation says about the backtest
The backtest cost model assumes 5bp open / 15bp close / $5 minimum (ch. 03). Realized slippage of 4.54 bps is inside the model's open-side assumption and well under the close-side assumption — the first live round did **not** reveal a cost-model under-estimate. That is a positive but narrow result: one round, ~$74k notional, mostly buys. The honest statement is the one the book keeps making — **live beats backtest, and one round is one round** `(PROVEN → round 3; generality HYPOTHESIS)`.
`TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 model over a full position window — the round's buys are still held at writing)`
`TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
## The trail as ground truth
Every claim in this chapter traces to the tac-rd-book execution trail — round_id, intents (versioned target portfolios), decisions (placed/skipped with reasons), linked Alpaca orders, fills, and the reconcile roll-up `(PROVEN → tac-rd-book schema and round 3 data)`. This is the honest alternative to quoting a backtest as a promise: the round can be re-opened, per-symbol residuals inspected, and the funnel re-counted by anyone with read access.
## Practice note
The funnel and cost figures are the *target* for the next round: the desk expects slippage ≤ ~5 bps and funnel ≥ 9/10 fills under normal conditions; any round that materially breaches either is a reconciliation event, not a rounding error (ch. 10, risk posture).