Compare commits
8
Commits
7d4fd6c12a
..
book
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8389142484 | ||
|
|
0f0ffa7e18 | ||
|
|
2146fdf9e8 | ||
|
|
d43a83b9ae | ||
|
|
40e2ac1a6f | ||
|
|
9b96ee6dbf | ||
|
|
6a1d6f2638 | ||
|
|
dec192b20f |
+7
-5
@@ -105,8 +105,9 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
### 08 — Portfolio construction
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
|
||||
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `HYPOTHESIS` (idea: pre-clean-lake exp 13/14, not re-tested post-reset) |
|
||||
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `HYPOTHESIS` (idea: pre-clean-lake exp 13; mechanism consistent with clean exp 26) |
|
||||
| The reference and live book are TopkDropout n_drop=1, equal weight × risk_degree | `PROVEN` — exp 26, round 3 |
|
||||
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
|
||||
|
||||
### 09 — Cost/turnover frontier
|
||||
@@ -118,9 +119,10 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
### 10 — Risk limits
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
|
||||
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18 |
|
||||
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
|
||||
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `HYPOTHESIS` (idea: pre-clean-lake exp 18, not comparable post-reset) |
|
||||
| Size/concentration caps hurt by cutting deployed capital | `HYPOTHESIS` (idea: pre-clean-lake exp 18) |
|
||||
| Entry/risk gates are no-ops when the signal is the bottleneck | `HYPOTHESIS` (idea: pre-clean-lake exp 20) |
|
||||
| The risk-limit spec executes and does not break the funnel (liq floor dropped 8, 10→10→10→9) | `PROVEN` — round 3 |
|
||||
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
|
||||
|
||||
### 11 — Live execution and reconciliation
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
# Chapter 04 — Prune, Don't Add: Feature-Family Ablation
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 04.
|
||||
|
||||
The campaign's strongest and most repeated finding is negative: on a ~50-name daily panel, **adding model-specific feature machinery regresses the signal, while pruning to a compact generic set improves it.** This chapter works out that finding — where it comes from, how it was proved on the clean lake, and the mechanism hypothesis behind it.
|
||||
|
||||
## The pruning thesis, in one sentence
|
||||
|
||||
`CSRankNorm + LightGBM` on a small cross-section rewards a small set of generic, scale-free, well-behaved statistics. Every time the desk added a family built for a specific model (OU mean-reversion, HMM regime, GARCH vol) or a dense family of raw derived fields (moment/volatility moments), the rank signal got worse.
|
||||
|
||||
## The ablation arc (pre-reset, idea material)
|
||||
|
||||
The first clean statement of the thesis came from feature-family ablation on the pre-reset lake: dropping the model-specific families (ou, hmm) and keeping the generic set (jump, har, trend, hurst, signature, ret, max_move) improved RankIC 0.030→0.064 and flipped net excess from −9.4% to +3.1% `(idea → exp 9, pre-clean-lake; EVIDENCE#003)`. The mirror-image run confirmed the failure mode: adding 16 moment/volatility fields regressed every metric (RankIC 0.064→0.047, net excess −16.2%, IR −1.57) `(idea → exp 11, pre-clean-lake; EVIDENCE#004)`.
|
||||
|
||||
These two pre-reset runs are **ideas, not proof** — they ran on the dirty lake. Their thesis survived exactly because the clean-lake runs reproduced the same direction (below).
|
||||
|
||||
## The clean-lake confirmation
|
||||
|
||||
The clean-lake compact set is the proof: the reference signal is built from raw OHLCV plus `sp_ret`, jump, RV1/5/22, vol ratios, trend slopes, logp, Hurst and path-signature L1/L2 — nothing model-specific `(PROVEN → exp 24, EVIDENCE#013)`. Then the isolation runs (ch. 07) proved the negative side on clean data:
|
||||
|
||||
- Adding `sp_ou_zscore` regresses the reference: IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25, EVIDENCE#014)`.
|
||||
- Adding the multi-horizon momentum bundle (M1) regresses it further: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) `(PROVEN → exp 29, EVIDENCE#017)`.
|
||||
- Adding the GARCH(1,1) vol-regime trio adds nothing: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31, EVIDENCE#019)`.
|
||||
|
||||
Three independent clean-lake additions, three failures, all against the same reference and the same metrics. This is the campaign's most reproduced empirical result.
|
||||
|
||||
## Mechanism hypotheses
|
||||
|
||||
Why does pruning win on this panel? Three hypotheses, none yet isolated (all `HYPOTHESIS`):
|
||||
|
||||
1. **Panel width vs feature count.** ~50 names gives ~50 cross-sectional observations per day. Each added feature is another column the model can split on; with so few observations, extra dimensions mostly fit noise that CSRankNorm then amplifies. The minimal generic set "won repeatedly" across three independent expansions (ou/hmm, realized moments, TA) `(HYPOTHESIS → chat-ideas.md)`.
|
||||
2. **Scale-free is necessary but not sufficient.** Features that survive CSRankNorm are scale-free; dense raw fields (moments) are not, and were regressed. But scale-freeness alone doesn't guarantee usefulness — the moments family was still dropped `(HYPOTHESIS → chat-ideas.md)`.
|
||||
3. **Single-feature strength ≠ marginal contribution.** The OU z-score was the strongest standalone time-series predictor (IC −0.15/−0.13) yet degraded the model — the "OU paradox" named in ch. 01 `(HYPOTHESIS → chat-ideas.md; TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`.
|
||||
|
||||
## What practice generally does (and why this differs)
|
||||
|
||||
Industry panels (thousands of names, cross-sectional breadth) routinely feed hundreds of features and let the model prune them. This campaign's panel is 50 ETFs — a breadth-limited, high-correlation cross-section where the dominant signal is common (ch. 01: drift is mostly market-wide). The desk's lesson is not "features are bad"; it is that **feature count must scale with cross-sectional breadth**, and on this breadth the marginal value of the next family is negative. That is a hypothesis about generality `TODO(evidence-needed: out-of-panel universe with wider breadth)`.
|
||||
|
||||
## The operational rule the book carries forward
|
||||
|
||||
Before adding any feature family to the reference, run it as an isolation experiment against the reference metrics (ch. 07). The default assumption is failure; the run must beat IC/RankIC **and** net IR on the clean lake to earn its place. This rule is what made exp 29 and exp 31 cheap negatives instead of silent regressions.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Chapter 05 — Ensembles and the Seed-Count Effect
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 05.
|
||||
|
||||
Ensemble averaging is the campaign's one *additive* lever that survived clean-data scrutiny. This chapter separates what the ensemble does (variance reduction on a noisy rank) from what it does not do (add information), and shows that the *count* of seeds is load-bearing.
|
||||
|
||||
## What the ensemble is
|
||||
|
||||
The reference model is a 5-seed LightGBM blend: five models, differing only by random seed, trained on the same features and label, averaged into one prediction. The mechanism is reinforcement against estimation noise — the same metric a noise-dominated signal needs most (ch. 01, decay/reinforcement).
|
||||
|
||||
## The evidence
|
||||
|
||||
- The 5-seed ensemble on the ablated generic features was the pre-reset best result (RankIC 0.0586, RankICIR 0.224, net excess +7.8%, IR 0.79) `(idea → exp 12, pre-clean-lake; EVIDENCE#005)`. Its numbers are inflated by the dirty lake (EVIDENCE#010) but its *design* — ensemble on ablated features, isolated from feature expansion — was re-validated on the clean lake.
|
||||
- The clean-lake reference is the same design: the 5-seed RankIC ensemble on the compact stochastic set (RankIC 0.0663, RankICIR 0.2545) `(PROVEN → exp 24, EVIDENCE#013)`, reproduced from the same family lineage in exp 22/23 `(PROVEN → EVIDENCE#011/#012)`.
|
||||
- **Seed count is load-bearing**: the 2-seed blend loses to the 5-seed blend on the identical compact set — RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21) `(PROVEN → exp 28, EVIDENCE#016)`. Fewer seeds is not "cheaper, same signal"; it is a measurably worse signal.
|
||||
|
||||
## Mechanism: variance reduction, not new information
|
||||
|
||||
Two observations pin the mechanism to variance reduction. First, the ensemble's metric benefit shows up most clearly in RankICIR/IR — the noise-adjusted ratios — rather than in raw IC, consistent with error cancellation `(PROVEN → exp 24 vs exp 22/23 schema; interpret as PROVEN direction, magnitude is window-specific)`. Second, the same features and label produce different outcomes by seed count alone, which means the marginal value of the 5th seed is *stability*: the model family is good enough that its remaining error is estimation variance, and averaging it away is the cheapest reliable win available `(HYPOTHESIS → chat-ideas.md: equal-weight seed blend > adaptive blending; the mechanism is not fully isolated)`.
|
||||
|
||||
## The open question
|
||||
|
||||
Does the 5-seed ensemble win by variance reduction or by diversifying model families (e.g. different effective trees/feature interactions per seed)? The two hypotheses make different predictions for a 10-seed run `TODO(evidence-needed: 10-seed vs 5-seed isolation; and whether seed-count benefit survives a wider panel)`. The desk has not yet run either `(open, chat-ideas.md)`.
|
||||
|
||||
## Interaction with pruning and cost
|
||||
|
||||
The ensemble amplifies a pruned signal — it is not a substitute for pruning (ch. 04) and it does not fix cost (ch. 09). The clean-lake sequence is explicit: the ensemble's net result (+2.13% IR 0.21) still barely clears costs; averaging improves the signal-to-noise ratio, and n_drop relief then converts that into net return `(PROVEN → exp 24 + exp 26, EVIDENCE#013/#015)`.
|
||||
|
||||
## Practice note
|
||||
|
||||
Equal-weight seed blending beat a rolling-IC adaptive blend in the campaign's design choices (pre-reset idea, untested head-to-head on clean data): adaptive weights re-fit to noise on a 50-name panel `(HYPOTHESIS → chat-ideas.md)`. The desk's rule of thumb: fix the seed count and the weights; spend experiment budget on pruning and cost relief, where the reproduced wins are.
|
||||
@@ -0,0 +1,46 @@
|
||||
# Chapter 07 — Isolation Runs: Single-Variable Discipline
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 07.
|
||||
|
||||
The isolation run is the campaign's unit of proof: one variable changes against the fixed reference; the same metrics decide the verdict. This chapter walks the clean-lake sequence (exp 28–31) as the worked example, and states the discipline's rules so a reader can run it themselves.
|
||||
|
||||
## Why isolation is the discipline
|
||||
|
||||
The research loop (ch. 02) is only as honest as its attribution. A run that changes two things cannot say which one moved the result. The clean-lake campaign therefore fixed the reference — exp 26's n_drop=1 configuration on the compact stochastic set `(PROVEN → exp 26, EVIDENCE#015)` — and ran every subsequent experiment as a single-variable change against it:
|
||||
|
||||
| Run | One variable changed | Reference | Verdict |
|
||||
|-----|----------------------|-----------|---------|
|
||||
| exp 28 | seed count 5 → 2 | same features/book | REFUTED (2 seeds worse) `(PROVEN → EVIDENCE#016)` |
|
||||
| exp 29 | + multi-horizon momentum bundle (M1) | same book | REFUTED `(PROVEN → EVIDENCE#017)` |
|
||||
| exp 30 | + risk-adjusted 22d Sharpe drift (M2) | same book | HYPOTHESIS (mixed) `(PROVEN run, EVIDENCE#018)` |
|
||||
| exp 31 | + GARCH(1,1) vol-regime trio (M3) | same book | REFUTED `(PROVEN → EVIDENCE#019)` |
|
||||
|
||||
Three of four refuted; one mixed. The isolation design is what makes "most additions fail" a *finding* rather than an anecdote.
|
||||
|
||||
## The acceptance contract
|
||||
|
||||
For an addition to earn its place it must beat the reference on **both** layers of the metrics ladder (ch. 01): the rank layer (IC/RankIC/ICIR/RankICIR) *and* the portfolio layer (net IR, MDD). Exp 30 is the canonical trap: M2 looked strong on the portfolio layer (net +6.53%, IR 0.62 vs +2.13%, IR 0.21) while its rank metrics were *lower* than reference (RankIC 0.0576 vs 0.0663) `(PROVEN → exp 30, EVIDENCE#018)`. Because the two layers disagreed and the run was not reproduced, the book labels it HYPOTHESIS rather than PROVEN. The rule: **a single run that improves one layer and degrades the other is a hypothesis, not a win** `TODO(evidence-needed: reproduction of exp 30 M2 on a second window)`.
|
||||
|
||||
## What each refutation taught
|
||||
|
||||
- **exp 28 (2 seeds):** the ensemble's value is tied to seed count; halving it is not a harmless cost cut (ch. 05).
|
||||
- **exp 29 (momentum bundle):** drift-as-feature fails even when the underlying structure exists (ch. 01, ch. 04). The mechanism (name-specific scale, collision with trend features) is a hypothesis.
|
||||
- **exp 31 (GARCH):** parametric vol modeling adds nothing to a model that already has the realized-vol ladder — the generic ladder is the feature; the parametric overlay is not (ch. 01).
|
||||
- **exp 30 (M2):** the single interesting non-refutation. Risk-adjusting the drift feature changed the portfolio behavior without improving the rank — an unexplained, unreproduced anomaly worth one more run, not a claim.
|
||||
|
||||
## How to run an isolation campaign
|
||||
|
||||
1. **Freeze a reference** — a configuration, a book spec, and a metric table that everything is judged against (exp 26 n_drop=1 in this campaign).
|
||||
2. **Change exactly one thing** per run; record the hypothesis and acceptance metric in the run notes *before* running (ch. 02, pre-registration).
|
||||
3. **Judge on both layers** of the metrics ladder; a single-layer improvement is a hypothesis.
|
||||
4. **Expect most runs to fail** — that is the point. A refuted run is a recorded negative that protects the next hypothesis from paying for the same mistake twice.
|
||||
|
||||
## The discipline as the value
|
||||
|
||||
The campaign's net-of-cost performance barely cleared costs at its best `(PROVEN → exp 26, EVIDENCE#015)`. In that regime, undisciplined feature-hunting is not neutral — it is the largest *expected* destroyer of the edge. Isolation runs converted feature-hunting into a bounded cost: a few runs to prove each family dead, instead of silently degrading the live signal. That is why the book counts exp 29 and exp 31 among its wins (ch. 12).
|
||||
|
||||
## Open questions
|
||||
|
||||
- `TODO(evidence-needed: M2 reproduction — the only surviving clean-lake improvement candidate)`
|
||||
- `TODO(evidence-needed: 5d-reversal standalone strategy net of costs, the un-isolated idea from ch. 01)`
|
||||
- `TODO(evidence-needed: HMM regime overlay (long-only gate) on the exp-26 book — an overlay test, not a feature test)`
|
||||
@@ -0,0 +1,39 @@
|
||||
# Chapter 08 — Portfolio Construction: Dropout vs Optimal Stop
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 08.
|
||||
|
||||
This chapter compares the two portfolio constructions the campaign actually ran — TopkDropout (the rank-based, turnover-conscious book that became the reference and the live book) and stochastic-control OptimalStopControl (entry/exit/stop parametrized allocation). The honest status is that the comparison is **pre-reset idea material**: both constructions were tested on the dirty lake and the alternates were never re-run on the clean lake. What is PROVEN on the clean lake is that the reference book is TopkDropout and that it executes (ch. 11); what the alternates would do on clean data is unknown.
|
||||
|
||||
## The two constructions
|
||||
|
||||
- **TopkDropout** (qlib TopkDropoutStrategy): each day, rank the cross-sectional scores, apply the dropout rule, and hold the selected top-k names equally weighted. The campaign's `n_drop` parameter controls which names the strategy refuses to chase, and with it the book's turnover (ch. 09).
|
||||
- **OptimalStopControl** (stochastic control): allocate toward a target portfolio with entry/exit thresholds, holding-period and stop-loss parameters. The campaign tried the baseline (entry 0.85 / exit 0.7 / hold 10 / stop −0.08) and a V2 with turnover bands, cooldown and a cap.
|
||||
|
||||
## The pre-reset comparison (idea material)
|
||||
|
||||
On the pre-reset lake, TopkDropout beat both stochastic-control variants: OptimalStopControl net excess −2.7% (IR −0.31) versus TopkDropout +7.8% `(idea → exp 13, pre-clean-lake; EVIDENCE#006)`, and the V2 also refuted (net −6.9%, IR −0.72) `(idea → exp 14, pre-clean-lake; EVIDENCE#007)`. The attributed mechanism was **turnover**: the stop-control constructions churned the book and bled ~11.3pp of cost drag `(idea → exp 13, EVIDENCE#006)`. That mechanism is plausible — it is the same cost drag that proved binding on the clean lake (exp 26, ch. 09) — but the numbers themselves are not usable (dirty lake, EVIDENCE#010).
|
||||
|
||||
`TODO(evidence-needed: OptimalStopControl vs TopkDropout A/B on the exp-26 reference and its n_drop=1 book — the clean-lake rerun of this comparison)`
|
||||
|
||||
## What is PROVEN on the clean lake
|
||||
|
||||
- The reference book is TopkDropout with `n_drop=1`, equal weight × risk_degree, and it is the campaign's best net result `(PROVEN → exp 26, EVIDENCE#015)`.
|
||||
- The same construction is the live book of round 3: 10 targets, 9 fills, realized slippage 4.54 bps `(PROVEN → round 3, EVIDENCE#020)`.
|
||||
- Construction is not a substitute for signal or cost work: the n_drop change moved net return by ~5.3pp with *identical* signal metrics `(PROVEN → exp 26, EVIDENCE#015)` — construction is where the cost edge is won or lost, and cost is the binding constraint (ch. 09).
|
||||
|
||||
## Sizing
|
||||
|
||||
Sizing in the campaign is equal weight × `risk_degree` (0.95) — a fixed fraction of account per name, floored to whole shares at execution `(PROVEN → the sizing used in exp 26 and round 3; tac-rd-book intents)`.
|
||||
|
||||
- Fractional-Kelly sizing (exp 15) was never verified — the run never finished `(HYPOTHESIS; TODO(evidence-needed: exp 15 Kelly re-run on the clean lake))`.
|
||||
- The hypothesis that equal-weight × risk_degree throws away edge-magnitude information (a Kelly-style rule would size by score spread) is untested `(HYPOTHESIS → chat-ideas.md)`.
|
||||
|
||||
## Practice note
|
||||
|
||||
The book's working rule: prefer the construction that minimizes turnover at a fixed topk (TopkDropout with controlled `n_drop` over parametrized stop-control), because cost is the binding constraint on this signal. That rule is a hypothesis until the clean-lake A/B lands.
|
||||
|
||||
## Open questions
|
||||
|
||||
- `TODO(evidence-needed: OptimalStopControl vs TopkDropout on clean data)`
|
||||
- `TODO(evidence-needed: Kelly-style sizing vs equal-weight × risk_degree on the exp-26 book)`
|
||||
- `TODO(evidence-needed: lower topk vs higher topk on the clean-lake reference — concentration vs diversification)`
|
||||
@@ -0,0 +1,40 @@
|
||||
# Chapter 09 — The Cost/Turnover Frontier
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 09.
|
||||
|
||||
This chapter is the cleanest result in the book: on an *identical signal*, the campaign moved net excess return from **−3.21% to +2.13%** purely by relieving turnover/cost pressure — changing the strategy's `n_drop` from 2 to 1. Signal metrics did not move; the outcome did. That is the definition of a binding cost constraint, and it sets the frontier every later improvement must operate on.
|
||||
|
||||
## The result
|
||||
|
||||
Exp 26 ran the identical compact-stochastic reference signal through two construction variants `(PROVEN → exp 26, EVIDENCE#015)`:
|
||||
|
||||
| Configuration | Gross | Net | MaxDD | IR | IC / RankIC |
|
||||
|---------------|-------|-----|-------|-----|-------------|
|
||||
| n_drop 2 (chases the drop) | +7.02% | **−3.21%** | — | — | identical |
|
||||
| n_drop 1 (holds the dropped name) | +7.02% | **+2.13%** | −7.69% | 0.21 | identical |
|
||||
|
||||
The gross return and the signal metrics (IC/RankIC) are identical between the two rows — the entire ~5.3pp gap is cost. `n_drop` 2 means the strategy refuses to hold the top-ranked name and buys the next one down, so it chases in and out of the extreme winner every day; `n_drop` 1 holds it. Lower turnover, not a better signal, is what turned the book positive `(PROVEN → exp 26)`.
|
||||
|
||||
## Why cost is the binding constraint
|
||||
|
||||
Ch. 03 established the noise floor: on this signal, the gross→net collapse is ~9–10pp of cost drag at realistic assumptions (5bp open / 15bp close / $5 minimum) `(PROVEN → exp 22–26 composite, EVIDENCE#011–015)`. Exp 26 then proved the direction of relief: cost is not a fixed tax you subtract, it is a **construction decision**. Turnover is the cost's driver, and turnover is chosen by the strategy — dropout rule, rebalance cadence, and order type.
|
||||
|
||||
The campaign's best net IR is 0.21 — a real but thin edge. Any addition that adds turnover faster than it adds gross return loses (the refuted runs of ch. 07 all *added* features that churned the book).
|
||||
|
||||
## The frontier
|
||||
|
||||
The campaign's measured points on the frontier:
|
||||
|
||||
- **n_drop 2 → 1**: the reproduced win; holds the extreme winner, cuts daily churn `(PROVEN → exp 26, EVIDENCE#015)`.
|
||||
- **Live round 3**: turnover ≈ 0.74 at n_drop 1, invested $74,202.85, realized slippage 4.54 bps `(PROVEN → round 3, EVIDENCE#020)`. The live turnover is the first *measured* number the frontier can be calibrated against.
|
||||
- **Untested relief levers** (hypotheses from the desk's design notes, not yet isolated on the clean lake): weekly instead of daily rebalance; no-trade buffer bands (skip trades below a return-to-cost threshold); notional instead of qty orders at small sizes `(HYPOTHESIS → chat-ideas.md)`.
|
||||
|
||||
`TODO(evidence-needed: weekly-rebalance and no-trade-band isolation runs on the exp-26 book — each would trade ~1pp of cost drag against ~1 day of signal decay)`
|
||||
|
||||
## The discipline the frontier imposes
|
||||
|
||||
Because the edge is thin and cost is the binding constraint, the acceptance contract for any change tightens: a candidate must beat the n_drop=1 reference on net IR *and* on IC/RankIC (ch. 07), and its turnover must not silently rise. The book treats turnover as a first-class metric to be reported with every run, not a tooling detail `(PROVEN → exp 26 + round 3; the numbers to report are turnover, slippage bps, and cost as % of gross)`.
|
||||
|
||||
## Practice note
|
||||
|
||||
Hold the winner. Prefer the lowest-turnover construction that preserves the ranking. Measure turnover and realized cost in every round; reconcile them against the backtest's 5bp/15bp/$5 model (ch. 11). The frontier is where this campaign's edge lives, and it is narrower than the backtest suggested.
|
||||
@@ -0,0 +1,42 @@
|
||||
# Chapter 10 — Risk Limits That Work
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 10. HITL review gate applies: risk-limit advice.
|
||||
|
||||
Risk limits gate the live book before execution: a liquidity floor, a size cap, a concentration cap, and a drawdown pause. This chapter is deliberately careful about what it claims: the **A/B evidence** that the liquidity floor beats concentration caps is a pre-reset idea (not comparable post-reset); what is PROVEN is that the spec **executed** in round 3 and the funnel held. Any desk acting on the A/B numbers is acting on a hypothesis until the post-reset rerun lands.
|
||||
|
||||
## The spec as executed
|
||||
|
||||
Round 3 ran with `risk_limits`: liquidity floor **$5M** (min 20-day average dollar volume), size cap **12%** of book per name, concentration cap **95%**, drawdown pause **10%** (pause new buys if equity ≤ 90% of peak) `(PROVEN → round 3, EVIDENCE#020)`. The gates acted: the liquidity floor dropped **8** names from the target list before placement, and one further name was skipped at decision time (SLV, `delta_zero`) — the funnel closed 10 → 10 → 10 → 9 `(PROVEN → round 3 funnel, EVIDENCE#020)`.
|
||||
|
||||
Two facts stand out. First, the liquidity floor was the *active* gate — it removed 8 of 10 low-liquidity ETF names, which is exactly the gate's purpose on a panel of thinly-traded funds. Second, the drawdown-pause gate did not trip (equity stayed above the pause threshold), so this round is **not** evidence about the pause's behavior — only about its non-interference.
|
||||
|
||||
## What is PROVEN vs what is idea material
|
||||
|
||||
- **PROVEN (execution):** the risk-limit spec runs, gates, and does not break the funnel — round 3 `(EVIDENCE#020)`.
|
||||
- **HYPOTHESIS (idea, pre-clean-lake):** the A/B that the $5M floor *improves* the book — net IR 0.81→0.98 and drawdown 7.9%→5.4% in exp 18 — and that size/concentration caps *hurt* by cutting deployed capital (IR 0.816) `(idea → exp 18, EVIDENCE#008; not comparable post-reset, exp 20 R0 note)`. These are exactly the numbers the book must NOT cite as fact.
|
||||
- **HYPOTHESIS (idea, pre-clean-lake):** entry/risk gates (momentum, HMM regime) were byte-identical no-ops in exp 20, supporting "the signal is the bottleneck, not the risk layer" `(idea → exp 20, EVIDENCE#009)`.
|
||||
|
||||
`TODO(evidence-needed: risk-limit A/B on the post-reset reference — rd_risk_calibrate on the exp-26 lineage, comparing floor-on vs floor-off and the cap grid)`
|
||||
|
||||
## Design guidance (derived, hedged)
|
||||
|
||||
Reading across the (pre-reset, idea-tier) A/B and the (clean, proven) execution, the book offers hedged guidance — each item marked for what it is:
|
||||
|
||||
1. **Liquidity floor first.** It was the only gate that acted in round 3, and it removes names the book cannot actually trade at size. `HYPOTHESIS` that it is the highest-value limit (pre-reset A/B + round-3 execution consistent, not a clean A/B).
|
||||
2. **Caps that cut deployed capital cost edge.** On a thin-cost book, a size cap that forces smaller positions than the strategy wants spends the exact budget ch. 09 says is binding. `HYPOTHESIS` (idea-tier evidence, mechanism consistent with exp 26).
|
||||
3. **Gates are no-ops when the signal is weak.** A regime/momentum gate that rarely trips adds complexity, not protection. `HYPOTHESIS` (idea-tier evidence).
|
||||
4. **Pause gates are for tail events.** The drawdown pause is untested in round 3; it is cheap insurance, and its behavior under stress is unknown. `HYPOTHESIS`.
|
||||
|
||||
## Where risk limits sit in the loop
|
||||
|
||||
Risk limits are a post-signal gate — they cannot create edge, they can only destroy or preserve it (ch. 02). The campaign's reading is that the signal is the bottleneck (exp 20 idea, consistent with the clean-lake cost finding of ch. 09): limits should remove untradeable names and stop the book from self-destructing in a drawdown, and otherwise stay out of the way. That is a working posture, not a proof.
|
||||
|
||||
## Practice note
|
||||
|
||||
Run limits as a pre-gate on the same spec that gates backtests (`rd_backtest`/`rd_strategy_targets` share the `risk_limits` spec, so backtest and live are gated identically — the setup the campaign used). Reconcile each round's gate actions (names dropped, pauses tripped) in the trail (ch. 11). Until the post-reset A/B lands, treat the floor's benefit as hypothesis and the spec's execution as fact.
|
||||
|
||||
## Open questions
|
||||
|
||||
- `TODO(evidence-needed: post-reset risk-limit A/B on the exp-26 lineage)`
|
||||
- `TODO(evidence-needed: drawdown-pause behavior — it never tripped; no evidence on its trigger/recovery)`
|
||||
- `TODO(evidence-needed: liquidity floor level sensitivity — is $5M the right cutoff on this panel?)`
|
||||
@@ -0,0 +1,39 @@
|
||||
# Chapter 11 — Live Execution and Reconciliation
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 11. HITL review gate applies: live performance numbers, cost/slippage figures.
|
||||
|
||||
The book's spine is that claims must be reconcilable (ch. 00). This chapter closes the loop: the live round that executed the proved reference signal, the funnel that held, the realized cost, and what reconciliation says about the backtest's honesty.
|
||||
|
||||
## The round
|
||||
|
||||
Round 3 (target date 2026-08-17) ran the exp-26 n_drop=1 configuration retrained on a rolling 4-year window, Topk10 with risk limits (liquidity floor $5M, size cap 12%, concentration cap 95%, drawdown pause 10%) `(PROVEN → round 3, EVIDENCE#020; trace 27, run 721ef257…, branch exp/27-scheduled-algo-retrain-on-2026-08-17)`.
|
||||
|
||||
## The funnel held
|
||||
|
||||
The execution funnel — targets → intents → decisions → placed → filled — closed at **10 → 10 → 10 → 9** `(PROVEN → round 3 funnel)`:
|
||||
|
||||
- 10 targets from the strategy's target list,
|
||||
- 10 decided, 10 placed,
|
||||
- 9 filled, 1 cancelled, and 1 skipped (SLV, `delta_zero` — the pre-skip gate stopped a zero-delta name).
|
||||
|
||||
A 90% fill-to-target ratio with one deliberate skip is a funnel that executed what the research claimed it would — the strategy's intent survived the gates and the broker. The reconciliation (targets vs decisions vs fills, per-symbol residuals) is available from the round's `book_reconcile`.
|
||||
|
||||
## Realized cost
|
||||
|
||||
Invested notional was **$74,202.85** with **realized slippage ≈ 4.54 bps** and estimated cost ≈ **$45**; turnover ≈ 0.74 `(PROVEN → round 3 metrics, EVIDENCE#020)`. The slippage figure is *realized* — taken from fills versus the expected execution price in the round's order trail — not a backtest assumption. This is the number the backtest cost model must be judged against.
|
||||
|
||||
## What reconciliation says about the backtest
|
||||
|
||||
The backtest cost model assumes 5bp open / 15bp close / $5 minimum (ch. 03). Realized slippage of 4.54 bps is inside the model's open-side assumption and well under the close-side assumption — the first live round did **not** reveal a cost-model under-estimate. That is a positive but narrow result: one round, ~$74k notional, mostly buys. The honest statement is the one the book keeps making — **live beats backtest, and one round is one round** `(PROVEN → round 3; generality HYPOTHESIS)`.
|
||||
|
||||
`TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 model over a full position window — the round's buys are still held at writing)`
|
||||
|
||||
`TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
|
||||
|
||||
## The trail as ground truth
|
||||
|
||||
Every claim in this chapter traces to the tac-rd-book execution trail — round_id, intents (versioned target portfolios), decisions (placed/skipped with reasons), linked Alpaca orders, fills, and the reconcile roll-up `(PROVEN → tac-rd-book schema and round 3 data)`. This is the honest alternative to quoting a backtest as a promise: the round can be re-opened, per-symbol residuals inspected, and the funnel re-counted by anyone with read access.
|
||||
|
||||
## Practice note
|
||||
|
||||
The funnel and cost figures are the *target* for the next round: the desk expects slippage ≤ ~5 bps and funnel ≥ 9/10 fills under normal conditions; any round that materially breaches either is a reconciliation event, not a rounding error (ch. 10, risk posture).
|
||||
@@ -0,0 +1,46 @@
|
||||
# Chapter 12 — Synthesis: How Proved Truth Compounds
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 12.
|
||||
|
||||
This chapter is the scoreboard: what actually moved performance, what was refuted, and why the book's methodology — not any single experiment — is the durable product of the campaign.
|
||||
|
||||
## The scoreboard
|
||||
|
||||
Everything PROVEN below is on the clean lake (exp 21+) or a reconciled post-reset round; everything else is labeled what it is.
|
||||
|
||||
| Lever | What moved | Status |
|
||||
|-------|-----------|--------|
|
||||
| **Data quality** (clean-lake reset) | The single largest event: invalidated all pre-reset results; the reference collapsed and was rebuilt (IC 0.0354→0.0019, then rebuilt to 0.0511) | `PROVEN` — exp 21→22–24 (EVIDENCE#010–013) |
|
||||
| **Cost/turnover relief** (n_drop 2→1) | The largest *positive* lever: net −3.21%→+2.13% on an identical signal (IR 0.21, MDD −7.69%) | `PROVEN` — exp 26 (EVIDENCE#015) |
|
||||
| **Feature pruning** (compact generic set) | The biggest series of wins were refutations: OU (exp 25), momentum (exp 29), GARCH (exp 31) all rejected; the compact set stands (RankIC 0.0663) | `PROVEN` — exp 24/25/29/31 (EVIDENCE#013/#014/#017/#019) |
|
||||
| **Ensemble & seed count** | Variance reduction, not new information; 2 seeds < 5 seeds (net −1.49% vs +2.13%) | `PROVEN` — exp 28 (EVIDENCE#016) |
|
||||
| **Risk limits** | Executed and non-interfering in round 3 (liquidity floor dropped 8, funnel 10→10→10→9); the floor-beats-caps A/B is pre-reset idea material | `PROVEN` (execution) / `HYPOTHESIS` (A/B) — round 3 + exp 18 |
|
||||
| **Portfolio construction** | TopkDropout beat stochastic-control on the pre-reset lake; never re-tested post-reset | `HYPOTHESIS` (idea) — exp 13/14 |
|
||||
| **Live execution** | Funnel held, slippage 4.54 bps, cost ~$45, turnover 0.74 — the first reconciled live number | `PROVEN` — round 3 (EVIDENCE#020) |
|
||||
|
||||
## The pattern beneath the scoreboard
|
||||
|
||||
Two positive levers (data quality, cost relief), one protective discipline (pruning, whose wins were negatives), one reinforcement (seed count). The pattern: **performance improved by removing lies, removing cost, and removing features — not by adding anything to the signal.** The only surviving clean-lake addition candidate is exp 30's M2 (Sharpe-drift feature), which the book keeps at HYPOTHESIS precisely because it improved one layer and degraded another in a single unreproduced run (ch. 07).
|
||||
|
||||
The refuted runs were as valuable as the wins: exp 11, 13, 14, 20, 25, 29, 31 each stopped a wrong direction at the cost of a few runs `(PROVEN — refuted runs recorded in the ledger; REFERENCED — falsification as method)`. A campaign that counts its refutations as output is a campaign that spends its budget learning, not re-learning.
|
||||
|
||||
## The methodology that made it compound
|
||||
|
||||
None of the scoreboard above is usable without the machinery of ch. 00–02:
|
||||
|
||||
1. **The execution trail** (targets→decisions→fills, reconcile) is the spine — it is what let the desk catch the clean-lake collapse and what turns the live round into evidence.
|
||||
2. **The clean-lake boundary** is the watermark — it is why exp 18's pretty risk numbers are hypotheses and exp 26's thin-but-real numbers are facts.
|
||||
3. **Isolation and pre-registration** make each verdict attributable (ch. 07).
|
||||
4. **The two-layer metrics ladder** (rank + portfolio, ch. 01) is why exp 30 is a hypothesis and not a claim.
|
||||
5. **Live beats backtest** (ch. 11) is the final gate — no metric in this book outranks a reconciled round.
|
||||
|
||||
## What the book still does not know
|
||||
|
||||
- Whether the 50-ETF panel generalizes — the widest open question `TODO(evidence-needed: out-of-panel universe)`.
|
||||
- Why the strongest single-feature signal (OU) degrades the model (the OU paradox, ch. 01).
|
||||
- Whether 5-day reversal trades standalone net of costs (ch. 01, ch. 07).
|
||||
- Whether M2 reproduces (ch. 07), whether the risk-limit A/B holds on clean data (ch. 10), and whether a second live round confirms the funnel and slippage under a different regime (ch. 11).
|
||||
|
||||
## Closing
|
||||
|
||||
This book's claims are deliberately thin: a RankIC near 0.066 on 50 names, an IR near 0.2 net, one reconciled live round. That thinness is the point. Every number in it can be re-derived from a recorded run or a re-opened round; every hypothesis is marked as one; every backtest is labeled a backtest. A quant-desk reader can act on the book's method even where its edge is small — and the book expects its own claims to be superseded as the next rounds and experiments land (living document, `AGENTS.md` rule 6).
|
||||
Reference in New Issue
Block a user