Compare commits

..
8 Commits
9 changed files with 331 additions and 5 deletions
+7 -5
View File
@@ -105,8 +105,9 @@ Each chapter opens with its claims. The inventory below is the working contract:
### 08 — Portfolio construction
| Claim | Expected status |
|-------|-----------------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `HYPOTHESIS` (idea: pre-clean-lake exp 13/14, not re-tested post-reset) |
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `HYPOTHESIS` (idea: pre-clean-lake exp 13; mechanism consistent with clean exp 26) |
| The reference and live book are TopkDropout n_drop=1, equal weight × risk_degree | `PROVEN` — exp 26, round 3 |
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
### 09 — Cost/turnover frontier
@@ -118,9 +119,10 @@ Each chapter opens with its claims. The inventory below is the working contract:
### 10 — Risk limits
| Claim | Expected status |
|-------|-----------------|
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18 |
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `HYPOTHESIS` (idea: pre-clean-lake exp 18, not comparable post-reset) |
| Size/concentration caps hurt by cutting deployed capital | `HYPOTHESIS` (idea: pre-clean-lake exp 18) |
| Entry/risk gates are no-ops when the signal is the bottleneck | `HYPOTHESIS` (idea: pre-clean-lake exp 20) |
| The risk-limit spec executes and does not break the funnel (liq floor dropped 8, 10→10→10→9) | `PROVEN` — round 3 |
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
### 11 — Live execution and reconciliation
+41
View File
@@ -0,0 +1,41 @@
# Chapter 04 — Prune, Don't Add: Feature-Family Ablation
Status: drafting. Claim inventory: see `README.md` ch. 04.
The campaign's strongest and most repeated finding is negative: on a ~50-name daily panel, **adding model-specific feature machinery regresses the signal, while pruning to a compact generic set improves it.** This chapter works out that finding — where it comes from, how it was proved on the clean lake, and the mechanism hypothesis behind it.
## The pruning thesis, in one sentence
`CSRankNorm + LightGBM` on a small cross-section rewards a small set of generic, scale-free, well-behaved statistics. Every time the desk added a family built for a specific model (OU mean-reversion, HMM regime, GARCH vol) or a dense family of raw derived fields (moment/volatility moments), the rank signal got worse.
## The ablation arc (pre-reset, idea material)
The first clean statement of the thesis came from feature-family ablation on the pre-reset lake: dropping the model-specific families (ou, hmm) and keeping the generic set (jump, har, trend, hurst, signature, ret, max_move) improved RankIC 0.030→0.064 and flipped net excess from −9.4% to +3.1% `(idea → exp 9, pre-clean-lake; EVIDENCE#003)`. The mirror-image run confirmed the failure mode: adding 16 moment/volatility fields regressed every metric (RankIC 0.064→0.047, net excess −16.2%, IR −1.57) `(idea → exp 11, pre-clean-lake; EVIDENCE#004)`.
These two pre-reset runs are **ideas, not proof** — they ran on the dirty lake. Their thesis survived exactly because the clean-lake runs reproduced the same direction (below).
## The clean-lake confirmation
The clean-lake compact set is the proof: the reference signal is built from raw OHLCV plus `sp_ret`, jump, RV1/5/22, vol ratios, trend slopes, logp, Hurst and path-signature L1/L2 — nothing model-specific `(PROVEN → exp 24, EVIDENCE#013)`. Then the isolation runs (ch. 07) proved the negative side on clean data:
- Adding `sp_ou_zscore` regresses the reference: IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25, EVIDENCE#014)`.
- Adding the multi-horizon momentum bundle (M1) regresses it further: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) `(PROVEN → exp 29, EVIDENCE#017)`.
- Adding the GARCH(1,1) vol-regime trio adds nothing: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31, EVIDENCE#019)`.
Three independent clean-lake additions, three failures, all against the same reference and the same metrics. This is the campaign's most reproduced empirical result.
## Mechanism hypotheses
Why does pruning win on this panel? Three hypotheses, none yet isolated (all `HYPOTHESIS`):
1. **Panel width vs feature count.** ~50 names gives ~50 cross-sectional observations per day. Each added feature is another column the model can split on; with so few observations, extra dimensions mostly fit noise that CSRankNorm then amplifies. The minimal generic set "won repeatedly" across three independent expansions (ou/hmm, realized moments, TA) `(HYPOTHESIS → chat-ideas.md)`.
2. **Scale-free is necessary but not sufficient.** Features that survive CSRankNorm are scale-free; dense raw fields (moments) are not, and were regressed. But scale-freeness alone doesn't guarantee usefulness — the moments family was still dropped `(HYPOTHESIS → chat-ideas.md)`.
3. **Single-feature strength ≠ marginal contribution.** The OU z-score was the strongest standalone time-series predictor (IC −0.15/−0.13) yet degraded the model — the "OU paradox" named in ch. 01 `(HYPOTHESIS → chat-ideas.md; TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`.
## What practice generally does (and why this differs)
Industry panels (thousands of names, cross-sectional breadth) routinely feed hundreds of features and let the model prune them. This campaign's panel is 50 ETFs — a breadth-limited, high-correlation cross-section where the dominant signal is common (ch. 01: drift is mostly market-wide). The desk's lesson is not "features are bad"; it is that **feature count must scale with cross-sectional breadth**, and on this breadth the marginal value of the next family is negative. That is a hypothesis about generality `TODO(evidence-needed: out-of-panel universe with wider breadth)`.
## The operational rule the book carries forward
Before adding any feature family to the reference, run it as an isolation experiment against the reference metrics (ch. 07). The default assumption is failure; the run must beat IC/RankIC **and** net IR on the clean lake to earn its place. This rule is what made exp 29 and exp 31 cheap negatives instead of silent regressions.
+31
View File
@@ -0,0 +1,31 @@
# Chapter 05 — Ensembles and the Seed-Count Effect
Status: drafting. Claim inventory: see `README.md` ch. 05.
Ensemble averaging is the campaign's one *additive* lever that survived clean-data scrutiny. This chapter separates what the ensemble does (variance reduction on a noisy rank) from what it does not do (add information), and shows that the *count* of seeds is load-bearing.
## What the ensemble is
The reference model is a 5-seed LightGBM blend: five models, differing only by random seed, trained on the same features and label, averaged into one prediction. The mechanism is reinforcement against estimation noise — the same metric a noise-dominated signal needs most (ch. 01, decay/reinforcement).
## The evidence
- The 5-seed ensemble on the ablated generic features was the pre-reset best result (RankIC 0.0586, RankICIR 0.224, net excess +7.8%, IR 0.79) `(idea → exp 12, pre-clean-lake; EVIDENCE#005)`. Its numbers are inflated by the dirty lake (EVIDENCE#010) but its *design* — ensemble on ablated features, isolated from feature expansion — was re-validated on the clean lake.
- The clean-lake reference is the same design: the 5-seed RankIC ensemble on the compact stochastic set (RankIC 0.0663, RankICIR 0.2545) `(PROVEN → exp 24, EVIDENCE#013)`, reproduced from the same family lineage in exp 22/23 `(PROVEN → EVIDENCE#011/#012)`.
- **Seed count is load-bearing**: the 2-seed blend loses to the 5-seed blend on the identical compact set — RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21) `(PROVEN → exp 28, EVIDENCE#016)`. Fewer seeds is not "cheaper, same signal"; it is a measurably worse signal.
## Mechanism: variance reduction, not new information
Two observations pin the mechanism to variance reduction. First, the ensemble's metric benefit shows up most clearly in RankICIR/IR — the noise-adjusted ratios — rather than in raw IC, consistent with error cancellation `(PROVEN → exp 24 vs exp 22/23 schema; interpret as PROVEN direction, magnitude is window-specific)`. Second, the same features and label produce different outcomes by seed count alone, which means the marginal value of the 5th seed is *stability*: the model family is good enough that its remaining error is estimation variance, and averaging it away is the cheapest reliable win available `(HYPOTHESIS → chat-ideas.md: equal-weight seed blend > adaptive blending; the mechanism is not fully isolated)`.
## The open question
Does the 5-seed ensemble win by variance reduction or by diversifying model families (e.g. different effective trees/feature interactions per seed)? The two hypotheses make different predictions for a 10-seed run `TODO(evidence-needed: 10-seed vs 5-seed isolation; and whether seed-count benefit survives a wider panel)`. The desk has not yet run either `(open, chat-ideas.md)`.
## Interaction with pruning and cost
The ensemble amplifies a pruned signal — it is not a substitute for pruning (ch. 04) and it does not fix cost (ch. 09). The clean-lake sequence is explicit: the ensemble's net result (+2.13% IR 0.21) still barely clears costs; averaging improves the signal-to-noise ratio, and n_drop relief then converts that into net return `(PROVEN → exp 24 + exp 26, EVIDENCE#013/#015)`.
## Practice note
Equal-weight seed blending beat a rolling-IC adaptive blend in the campaign's design choices (pre-reset idea, untested head-to-head on clean data): adaptive weights re-fit to noise on a 50-name panel `(HYPOTHESIS → chat-ideas.md)`. The desk's rule of thumb: fix the seed count and the weights; spend experiment budget on pruning and cost relief, where the reproduced wins are.
+46
View File
@@ -0,0 +1,46 @@
# Chapter 07 — Isolation Runs: Single-Variable Discipline
Status: drafting. Claim inventory: see `README.md` ch. 07.
The isolation run is the campaign's unit of proof: one variable changes against the fixed reference; the same metrics decide the verdict. This chapter walks the clean-lake sequence (exp 28–31) as the worked example, and states the discipline's rules so a reader can run it themselves.
## Why isolation is the discipline
The research loop (ch. 02) is only as honest as its attribution. A run that changes two things cannot say which one moved the result. The clean-lake campaign therefore fixed the reference — exp 26's n_drop=1 configuration on the compact stochastic set `(PROVEN → exp 26, EVIDENCE#015)` — and ran every subsequent experiment as a single-variable change against it:
| Run | One variable changed | Reference | Verdict |
|-----|----------------------|-----------|---------|
| exp 28 | seed count 5 → 2 | same features/book | REFUTED (2 seeds worse) `(PROVEN → EVIDENCE#016)` |
| exp 29 | + multi-horizon momentum bundle (M1) | same book | REFUTED `(PROVEN → EVIDENCE#017)` |
| exp 30 | + risk-adjusted 22d Sharpe drift (M2) | same book | HYPOTHESIS (mixed) `(PROVEN run, EVIDENCE#018)` |
| exp 31 | + GARCH(1,1) vol-regime trio (M3) | same book | REFUTED `(PROVEN → EVIDENCE#019)` |
Three of four refuted; one mixed. The isolation design is what makes "most additions fail" a *finding* rather than an anecdote.
## The acceptance contract
For an addition to earn its place it must beat the reference on **both** layers of the metrics ladder (ch. 01): the rank layer (IC/RankIC/ICIR/RankICIR) *and* the portfolio layer (net IR, MDD). Exp 30 is the canonical trap: M2 looked strong on the portfolio layer (net +6.53%, IR 0.62 vs +2.13%, IR 0.21) while its rank metrics were *lower* than reference (RankIC 0.0576 vs 0.0663) `(PROVEN → exp 30, EVIDENCE#018)`. Because the two layers disagreed and the run was not reproduced, the book labels it HYPOTHESIS rather than PROVEN. The rule: **a single run that improves one layer and degrades the other is a hypothesis, not a win** `TODO(evidence-needed: reproduction of exp 30 M2 on a second window)`.
## What each refutation taught
- **exp 28 (2 seeds):** the ensemble's value is tied to seed count; halving it is not a harmless cost cut (ch. 05).
- **exp 29 (momentum bundle):** drift-as-feature fails even when the underlying structure exists (ch. 01, ch. 04). The mechanism (name-specific scale, collision with trend features) is a hypothesis.
- **exp 31 (GARCH):** parametric vol modeling adds nothing to a model that already has the realized-vol ladder — the generic ladder is the feature; the parametric overlay is not (ch. 01).
- **exp 30 (M2):** the single interesting non-refutation. Risk-adjusting the drift feature changed the portfolio behavior without improving the rank — an unexplained, unreproduced anomaly worth one more run, not a claim.
## How to run an isolation campaign
1. **Freeze a reference** — a configuration, a book spec, and a metric table that everything is judged against (exp 26 n_drop=1 in this campaign).
2. **Change exactly one thing** per run; record the hypothesis and acceptance metric in the run notes *before* running (ch. 02, pre-registration).
3. **Judge on both layers** of the metrics ladder; a single-layer improvement is a hypothesis.
4. **Expect most runs to fail** — that is the point. A refuted run is a recorded negative that protects the next hypothesis from paying for the same mistake twice.
## The discipline as the value
The campaign's net-of-cost performance barely cleared costs at its best `(PROVEN → exp 26, EVIDENCE#015)`. In that regime, undisciplined feature-hunting is not neutral — it is the largest *expected* destroyer of the edge. Isolation runs converted feature-hunting into a bounded cost: a few runs to prove each family dead, instead of silently degrading the live signal. That is why the book counts exp 29 and exp 31 among its wins (ch. 12).
## Open questions
- `TODO(evidence-needed: M2 reproduction — the only surviving clean-lake improvement candidate)`
- `TODO(evidence-needed: 5d-reversal standalone strategy net of costs, the un-isolated idea from ch. 01)`
- `TODO(evidence-needed: HMM regime overlay (long-only gate) on the exp-26 book — an overlay test, not a feature test)`
@@ -0,0 +1,39 @@
# Chapter 08 — Portfolio Construction: Dropout vs Optimal Stop
Status: drafting. Claim inventory: see `README.md` ch. 08.
This chapter compares the two portfolio constructions the campaign actually ran — TopkDropout (the rank-based, turnover-conscious book that became the reference and the live book) and stochastic-control OptimalStopControl (entry/exit/stop parametrized allocation). The honest status is that the comparison is **pre-reset idea material**: both constructions were tested on the dirty lake and the alternates were never re-run on the clean lake. What is PROVEN on the clean lake is that the reference book is TopkDropout and that it executes (ch. 11); what the alternates would do on clean data is unknown.
## The two constructions
- **TopkDropout** (qlib TopkDropoutStrategy): each day, rank the cross-sectional scores, apply the dropout rule, and hold the selected top-k names equally weighted. The campaign's `n_drop` parameter controls which names the strategy refuses to chase, and with it the book's turnover (ch. 09).
- **OptimalStopControl** (stochastic control): allocate toward a target portfolio with entry/exit thresholds, holding-period and stop-loss parameters. The campaign tried the baseline (entry 0.85 / exit 0.7 / hold 10 / stop −0.08) and a V2 with turnover bands, cooldown and a cap.
## The pre-reset comparison (idea material)
On the pre-reset lake, TopkDropout beat both stochastic-control variants: OptimalStopControl net excess −2.7% (IR −0.31) versus TopkDropout +7.8% `(idea → exp 13, pre-clean-lake; EVIDENCE#006)`, and the V2 also refuted (net −6.9%, IR −0.72) `(idea → exp 14, pre-clean-lake; EVIDENCE#007)`. The attributed mechanism was **turnover**: the stop-control constructions churned the book and bled ~11.3pp of cost drag `(idea → exp 13, EVIDENCE#006)`. That mechanism is plausible — it is the same cost drag that proved binding on the clean lake (exp 26, ch. 09) — but the numbers themselves are not usable (dirty lake, EVIDENCE#010).
`TODO(evidence-needed: OptimalStopControl vs TopkDropout A/B on the exp-26 reference and its n_drop=1 book — the clean-lake rerun of this comparison)`
## What is PROVEN on the clean lake
- The reference book is TopkDropout with `n_drop=1`, equal weight × risk_degree, and it is the campaign's best net result `(PROVEN → exp 26, EVIDENCE#015)`.
- The same construction is the live book of round 3: 10 targets, 9 fills, realized slippage 4.54 bps `(PROVEN → round 3, EVIDENCE#020)`.
- Construction is not a substitute for signal or cost work: the n_drop change moved net return by ~5.3pp with *identical* signal metrics `(PROVEN → exp 26, EVIDENCE#015)` — construction is where the cost edge is won or lost, and cost is the binding constraint (ch. 09).
## Sizing
Sizing in the campaign is equal weight × `risk_degree` (0.95) — a fixed fraction of account per name, floored to whole shares at execution `(PROVEN → the sizing used in exp 26 and round 3; tac-rd-book intents)`.
- Fractional-Kelly sizing (exp 15) was never verified — the run never finished `(HYPOTHESIS; TODO(evidence-needed: exp 15 Kelly re-run on the clean lake))`.
- The hypothesis that equal-weight × risk_degree throws away edge-magnitude information (a Kelly-style rule would size by score spread) is untested `(HYPOTHESIS → chat-ideas.md)`.
## Practice note
The book's working rule: prefer the construction that minimizes turnover at a fixed topk (TopkDropout with controlled `n_drop` over parametrized stop-control), because cost is the binding constraint on this signal. That rule is a hypothesis until the clean-lake A/B lands.
## Open questions
- `TODO(evidence-needed: OptimalStopControl vs TopkDropout on clean data)`
- `TODO(evidence-needed: Kelly-style sizing vs equal-weight × risk_degree on the exp-26 book)`
- `TODO(evidence-needed: lower topk vs higher topk on the clean-lake reference — concentration vs diversification)`
@@ -0,0 +1,40 @@
# Chapter 09 — The Cost/Turnover Frontier
Status: drafting. Claim inventory: see `README.md` ch. 09.
This chapter is the cleanest result in the book: on an *identical signal*, the campaign moved net excess return from **−3.21% to +2.13%** purely by relieving turnover/cost pressure — changing the strategy's `n_drop` from 2 to 1. Signal metrics did not move; the outcome did. That is the definition of a binding cost constraint, and it sets the frontier every later improvement must operate on.
## The result
Exp 26 ran the identical compact-stochastic reference signal through two construction variants `(PROVEN → exp 26, EVIDENCE#015)`:
| Configuration | Gross | Net | MaxDD | IR | IC / RankIC |
|---------------|-------|-----|-------|-----|-------------|
| n_drop 2 (chases the drop) | +7.02% | **−3.21%** | — | — | identical |
| n_drop 1 (holds the dropped name) | +7.02% | **+2.13%** | −7.69% | 0.21 | identical |
The gross return and the signal metrics (IC/RankIC) are identical between the two rows — the entire ~5.3pp gap is cost. `n_drop` 2 means the strategy refuses to hold the top-ranked name and buys the next one down, so it chases in and out of the extreme winner every day; `n_drop` 1 holds it. Lower turnover, not a better signal, is what turned the book positive `(PROVEN → exp 26)`.
## Why cost is the binding constraint
Ch. 03 established the noise floor: on this signal, the gross→net collapse is ~9–10pp of cost drag at realistic assumptions (5bp open / 15bp close / $5 minimum) `(PROVEN → exp 22–26 composite, EVIDENCE#011–015)`. Exp 26 then proved the direction of relief: cost is not a fixed tax you subtract, it is a **construction decision**. Turnover is the cost's driver, and turnover is chosen by the strategy — dropout rule, rebalance cadence, and order type.
The campaign's best net IR is 0.21 — a real but thin edge. Any addition that adds turnover faster than it adds gross return loses (the refuted runs of ch. 07 all *added* features that churned the book).
## The frontier
The campaign's measured points on the frontier:
- **n_drop 2 → 1**: the reproduced win; holds the extreme winner, cuts daily churn `(PROVEN → exp 26, EVIDENCE#015)`.
- **Live round 3**: turnover ≈ 0.74 at n_drop 1, invested $74,202.85, realized slippage 4.54 bps `(PROVEN → round 3, EVIDENCE#020)`. The live turnover is the first *measured* number the frontier can be calibrated against.
- **Untested relief levers** (hypotheses from the desk's design notes, not yet isolated on the clean lake): weekly instead of daily rebalance; no-trade buffer bands (skip trades below a return-to-cost threshold); notional instead of qty orders at small sizes `(HYPOTHESIS → chat-ideas.md)`.
`TODO(evidence-needed: weekly-rebalance and no-trade-band isolation runs on the exp-26 book — each would trade ~1pp of cost drag against ~1 day of signal decay)`
## The discipline the frontier imposes
Because the edge is thin and cost is the binding constraint, the acceptance contract for any change tightens: a candidate must beat the n_drop=1 reference on net IR *and* on IC/RankIC (ch. 07), and its turnover must not silently rise. The book treats turnover as a first-class metric to be reported with every run, not a tooling detail `(PROVEN → exp 26 + round 3; the numbers to report are turnover, slippage bps, and cost as % of gross)`.
## Practice note
Hold the winner. Prefer the lowest-turnover construction that preserves the ranking. Measure turnover and realized cost in every round; reconcile them against the backtest's 5bp/15bp/$5 model (ch. 11). The frontier is where this campaign's edge lives, and it is narrower than the backtest suggested.
+42
View File
@@ -0,0 +1,42 @@
# Chapter 10 — Risk Limits That Work
Status: drafting. Claim inventory: see `README.md` ch. 10. HITL review gate applies: risk-limit advice.
Risk limits gate the live book before execution: a liquidity floor, a size cap, a concentration cap, and a drawdown pause. This chapter is deliberately careful about what it claims: the **A/B evidence** that the liquidity floor beats concentration caps is a pre-reset idea (not comparable post-reset); what is PROVEN is that the spec **executed** in round 3 and the funnel held. Any desk acting on the A/B numbers is acting on a hypothesis until the post-reset rerun lands.
## The spec as executed
Round 3 ran with `risk_limits`: liquidity floor **$5M** (min 20-day average dollar volume), size cap **12%** of book per name, concentration cap **95%**, drawdown pause **10%** (pause new buys if equity ≤ 90% of peak) `(PROVEN → round 3, EVIDENCE#020)`. The gates acted: the liquidity floor dropped **8** names from the target list before placement, and one further name was skipped at decision time (SLV, `delta_zero`) — the funnel closed 10 → 10 → 10 → 9 `(PROVEN → round 3 funnel, EVIDENCE#020)`.
Two facts stand out. First, the liquidity floor was the *active* gate — it removed 8 of 10 low-liquidity ETF names, which is exactly the gate's purpose on a panel of thinly-traded funds. Second, the drawdown-pause gate did not trip (equity stayed above the pause threshold), so this round is **not** evidence about the pause's behavior — only about its non-interference.
## What is PROVEN vs what is idea material
- **PROVEN (execution):** the risk-limit spec runs, gates, and does not break the funnel — round 3 `(EVIDENCE#020)`.
- **HYPOTHESIS (idea, pre-clean-lake):** the A/B that the $5M floor *improves* the book — net IR 0.81→0.98 and drawdown 7.9%→5.4% in exp 18 — and that size/concentration caps *hurt* by cutting deployed capital (IR 0.816) `(idea → exp 18, EVIDENCE#008; not comparable post-reset, exp 20 R0 note)`. These are exactly the numbers the book must NOT cite as fact.
- **HYPOTHESIS (idea, pre-clean-lake):** entry/risk gates (momentum, HMM regime) were byte-identical no-ops in exp 20, supporting "the signal is the bottleneck, not the risk layer" `(idea → exp 20, EVIDENCE#009)`.
`TODO(evidence-needed: risk-limit A/B on the post-reset reference — rd_risk_calibrate on the exp-26 lineage, comparing floor-on vs floor-off and the cap grid)`
## Design guidance (derived, hedged)
Reading across the (pre-reset, idea-tier) A/B and the (clean, proven) execution, the book offers hedged guidance — each item marked for what it is:
1. **Liquidity floor first.** It was the only gate that acted in round 3, and it removes names the book cannot actually trade at size. `HYPOTHESIS` that it is the highest-value limit (pre-reset A/B + round-3 execution consistent, not a clean A/B).
2. **Caps that cut deployed capital cost edge.** On a thin-cost book, a size cap that forces smaller positions than the strategy wants spends the exact budget ch. 09 says is binding. `HYPOTHESIS` (idea-tier evidence, mechanism consistent with exp 26).
3. **Gates are no-ops when the signal is weak.** A regime/momentum gate that rarely trips adds complexity, not protection. `HYPOTHESIS` (idea-tier evidence).
4. **Pause gates are for tail events.** The drawdown pause is untested in round 3; it is cheap insurance, and its behavior under stress is unknown. `HYPOTHESIS`.
## Where risk limits sit in the loop
Risk limits are a post-signal gate — they cannot create edge, they can only destroy or preserve it (ch. 02). The campaign's reading is that the signal is the bottleneck (exp 20 idea, consistent with the clean-lake cost finding of ch. 09): limits should remove untradeable names and stop the book from self-destructing in a drawdown, and otherwise stay out of the way. That is a working posture, not a proof.
## Practice note
Run limits as a pre-gate on the same spec that gates backtests (`rd_backtest`/`rd_strategy_targets` share the `risk_limits` spec, so backtest and live are gated identically — the setup the campaign used). Reconcile each round's gate actions (names dropped, pauses tripped) in the trail (ch. 11). Until the post-reset A/B lands, treat the floor's benefit as hypothesis and the spec's execution as fact.
## Open questions
- `TODO(evidence-needed: post-reset risk-limit A/B on the exp-26 lineage)`
- `TODO(evidence-needed: drawdown-pause behavior — it never tripped; no evidence on its trigger/recovery)`
- `TODO(evidence-needed: liquidity floor level sensitivity — is $5M the right cutoff on this panel?)`
@@ -0,0 +1,39 @@
# Chapter 11 — Live Execution and Reconciliation
Status: drafting. Claim inventory: see `README.md` ch. 11. HITL review gate applies: live performance numbers, cost/slippage figures.
The book's spine is that claims must be reconcilable (ch. 00). This chapter closes the loop: the live round that executed the proved reference signal, the funnel that held, the realized cost, and what reconciliation says about the backtest's honesty.
## The round
Round 3 (target date 2026-08-17) ran the exp-26 n_drop=1 configuration retrained on a rolling 4-year window, Topk10 with risk limits (liquidity floor $5M, size cap 12%, concentration cap 95%, drawdown pause 10%) `(PROVEN → round 3, EVIDENCE#020; trace 27, run 721ef257…, branch exp/27-scheduled-algo-retrain-on-2026-08-17)`.
## The funnel held
The execution funnel — targets → intents → decisions → placed → filled — closed at **10 → 10 → 10 → 9** `(PROVEN → round 3 funnel)`:
- 10 targets from the strategy's target list,
- 10 decided, 10 placed,
- 9 filled, 1 cancelled, and 1 skipped (SLV, `delta_zero` — the pre-skip gate stopped a zero-delta name).
A 90% fill-to-target ratio with one deliberate skip is a funnel that executed what the research claimed it would — the strategy's intent survived the gates and the broker. The reconciliation (targets vs decisions vs fills, per-symbol residuals) is available from the round's `book_reconcile`.
## Realized cost
Invested notional was **$74,202.85** with **realized slippage ≈ 4.54 bps** and estimated cost ≈ **$45**; turnover ≈ 0.74 `(PROVEN → round 3 metrics, EVIDENCE#020)`. The slippage figure is *realized* — taken from fills versus the expected execution price in the round's order trail — not a backtest assumption. This is the number the backtest cost model must be judged against.
## What reconciliation says about the backtest
The backtest cost model assumes 5bp open / 15bp close / $5 minimum (ch. 03). Realized slippage of 4.54 bps is inside the model's open-side assumption and well under the close-side assumption — the first live round did **not** reveal a cost-model under-estimate. That is a positive but narrow result: one round, ~$74k notional, mostly buys. The honest statement is the one the book keeps making — **live beats backtest, and one round is one round** `(PROVEN → round 3; generality HYPOTHESIS)`.
`TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 model over a full position window — the round's buys are still held at writing)`
`TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
## The trail as ground truth
Every claim in this chapter traces to the tac-rd-book execution trail — round_id, intents (versioned target portfolios), decisions (placed/skipped with reasons), linked Alpaca orders, fills, and the reconcile roll-up `(PROVEN → tac-rd-book schema and round 3 data)`. This is the honest alternative to quoting a backtest as a promise: the round can be re-opened, per-symbol residuals inspected, and the funnel re-counted by anyone with read access.
## Practice note
The funnel and cost figures are the *target* for the next round: the desk expects slippage ≤ ~5 bps and funnel ≥ 9/10 fills under normal conditions; any round that materially breaches either is a reconciliation event, not a rounding error (ch. 10, risk posture).
+46
View File
@@ -0,0 +1,46 @@
# Chapter 12 — Synthesis: How Proved Truth Compounds
Status: drafting. Claim inventory: see `README.md` ch. 12.
This chapter is the scoreboard: what actually moved performance, what was refuted, and why the book's methodology — not any single experiment — is the durable product of the campaign.
## The scoreboard
Everything PROVEN below is on the clean lake (exp 21+) or a reconciled post-reset round; everything else is labeled what it is.
| Lever | What moved | Status |
|-------|-----------|--------|
| **Data quality** (clean-lake reset) | The single largest event: invalidated all pre-reset results; the reference collapsed and was rebuilt (IC 0.0354→0.0019, then rebuilt to 0.0511) | `PROVEN` — exp 21→22–24 (EVIDENCE#010–013) |
| **Cost/turnover relief** (n_drop 2→1) | The largest *positive* lever: net −3.21%→+2.13% on an identical signal (IR 0.21, MDD −7.69%) | `PROVEN` — exp 26 (EVIDENCE#015) |
| **Feature pruning** (compact generic set) | The biggest series of wins were refutations: OU (exp 25), momentum (exp 29), GARCH (exp 31) all rejected; the compact set stands (RankIC 0.0663) | `PROVEN` — exp 24/25/29/31 (EVIDENCE#013/#014/#017/#019) |
| **Ensemble & seed count** | Variance reduction, not new information; 2 seeds < 5 seeds (net −1.49% vs +2.13%) | `PROVEN` — exp 28 (EVIDENCE#016) |
| **Risk limits** | Executed and non-interfering in round 3 (liquidity floor dropped 8, funnel 10→10→10→9); the floor-beats-caps A/B is pre-reset idea material | `PROVEN` (execution) / `HYPOTHESIS` (A/B) — round 3 + exp 18 |
| **Portfolio construction** | TopkDropout beat stochastic-control on the pre-reset lake; never re-tested post-reset | `HYPOTHESIS` (idea) — exp 13/14 |
| **Live execution** | Funnel held, slippage 4.54 bps, cost ~$45, turnover 0.74 — the first reconciled live number | `PROVEN` — round 3 (EVIDENCE#020) |
## The pattern beneath the scoreboard
Two positive levers (data quality, cost relief), one protective discipline (pruning, whose wins were negatives), one reinforcement (seed count). The pattern: **performance improved by removing lies, removing cost, and removing features — not by adding anything to the signal.** The only surviving clean-lake addition candidate is exp 30's M2 (Sharpe-drift feature), which the book keeps at HYPOTHESIS precisely because it improved one layer and degraded another in a single unreproduced run (ch. 07).
The refuted runs were as valuable as the wins: exp 11, 13, 14, 20, 25, 29, 31 each stopped a wrong direction at the cost of a few runs `(PROVEN — refuted runs recorded in the ledger; REFERENCED — falsification as method)`. A campaign that counts its refutations as output is a campaign that spends its budget learning, not re-learning.
## The methodology that made it compound
None of the scoreboard above is usable without the machinery of ch. 00–02:
1. **The execution trail** (targets→decisions→fills, reconcile) is the spine — it is what let the desk catch the clean-lake collapse and what turns the live round into evidence.
2. **The clean-lake boundary** is the watermark — it is why exp 18's pretty risk numbers are hypotheses and exp 26's thin-but-real numbers are facts.
3. **Isolation and pre-registration** make each verdict attributable (ch. 07).
4. **The two-layer metrics ladder** (rank + portfolio, ch. 01) is why exp 30 is a hypothesis and not a claim.
5. **Live beats backtest** (ch. 11) is the final gate — no metric in this book outranks a reconciled round.
## What the book still does not know
- Whether the 50-ETF panel generalizes — the widest open question `TODO(evidence-needed: out-of-panel universe)`.
- Why the strongest single-feature signal (OU) degrades the model (the OU paradox, ch. 01).
- Whether 5-day reversal trades standalone net of costs (ch. 01, ch. 07).
- Whether M2 reproduces (ch. 07), whether the risk-limit A/B holds on clean data (ch. 10), and whether a second live round confirms the funnel and slippage under a different regime (ch. 11).
## Closing
This book's claims are deliberately thin: a RankIC near 0.066 on 50 names, an IR near 0.2 net, one reconciled live round. That thinness is the point. Every number in it can be re-derived from a recorded run or a re-opened round; every hypothesis is marked as one; every backtest is labeled a backtest. A quant-desk reader can act on the book's method even where its edge is small — and the book expects its own claims to be superseded as the next rounds and experiments land (living document, `AGENTS.md` rule 6).