book: insert ch01 metrics vocabulary + ch02 research loop; renumber 03/06 — evidence exp 21-31, martingale study

This commit is contained in:
TradeAC Book Agent
2026-08-18 23:04:10 +00:00
parent 3562b7f776
commit 436692a620
7 changed files with 186 additions and 34 deletions
+32 -22
View File
@@ -23,17 +23,18 @@ A quant-desk reader should be able to act on this book: replicate a signal pipel
| # | Chapter | Status | Core experiments cited | Core lesson | | # | Chapter | Status | Core experiments cited | Core lesson |
|---|---------|--------|------------------------|-------------| |---|---------|--------|------------------------|-------------|
| 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile | | 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile |
| 01 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology | | 01 | Metrics: the vocabulary of a price series | drafting | exp 21–31 + dataset studies | Every claim reduces to a falsifiable statistic |
| 02 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal | | 02 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
| 03 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific | | 03 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal |
| 04 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing | | 04 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific |
| 05 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise | | 05 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing |
| 06 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value | | 06 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
| 07 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge | | 07 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value |
| 08 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge | | 08 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge |
| 09 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck | | 09 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge |
| 10 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 | | 10 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck |
| 11 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why | | 11 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
| 12 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why |
Status legend: `drafting` → `in-review` → `done`. Status legend: `drafting` → `in-review` → `done`.
@@ -48,21 +49,30 @@ Each chapter opens with its claims. The inventory below is the working contract:
| Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 | | Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 |
| The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema | | The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema |
### 01 — The research loop ### 01 — Metrics: the vocabulary of a price series
| Claim | Expected status |
|-------|-----------------|
| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–12) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) |
| Generic scale-free statistics beat model-specific machinery on a small daily panel | `PROVEN` (exp 23/24/25/29/31) + `HYPOTHESIS` (generality) |
| A statistic is only as good as the falsification it survives (null z-scores, reproduction) | `PROVEN` (exp 21 detection playbook) + `REFERENCED` |
| The strongest single-feature signal (OU z-score) can be worthless inside a rank model — the "OU paradox" | `PROVEN` (exp 25) + open mechanism `TODO(evidence-needed)` |
### 02 — The research loop
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 | | Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 |
| Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) | | Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) |
| The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk | | The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk |
| One variable changes per run (isolation); verdicts attributable | `PROVEN` — exp 26→28/29/30/31 design |
### 02 — Baseline and the cost reality ### 03 — Baseline and the cost reality
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 | | Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 |
| Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 | | Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 |
| A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` | | A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` |
### 03 — Prune, don't add ### 04 — Prune, don't add
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 | | Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
@@ -70,21 +80,21 @@ Each chapter opens with its claims. The inventory below is the working contract:
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 | | Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3 runs, still panel-specific) | | More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3 runs, still panel-specific) |
### 04 — Ensembles ### 05 — Ensembles
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 | | 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 | | Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design | | Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
### 05 — Clean-lake reset ### 06 — Clean-lake reset
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 | | The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 |
| Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction | | Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction |
| Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality | | Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality |
### 06 — Isolation runs ### 07 — Isolation runs
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 design | | Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 design |
@@ -92,20 +102,20 @@ Each chapter opens with its claims. The inventory below is the working contract:
| Risk-adjusted 22d Sharpe drift is promising on portfolio metrics, mixed on rank | `HYPOTHESIS` — exp 30 single run, unreproduced | | Risk-adjusted 22d Sharpe drift is promising on portfolio metrics, mixed on rank | `HYPOTHESIS` — exp 30 single run, unreproduced |
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 | | GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
### 07 — Portfolio construction ### 08 — Portfolio construction
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 | | TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 | | Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished | | Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
### 08 — Cost/turnover frontier ### 09 — Cost/turnover frontier
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 | | n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) | | Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
### 09 — Risk limits ### 10 — Risk limits
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) | | $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
@@ -113,14 +123,14 @@ Each chapter opens with its claims. The inventory below is the working contract:
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) | | Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note | | Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
### 10 — Live execution and reconciliation ### 11 — Live execution and reconciliation
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 | | Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics | | Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology | | Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
### 11 — Synthesis ### 12 — Synthesis
| Claim | Expected status | | Claim | Expected status |
|-------|-----------------| |-------|-----------------|
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26 | | The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26 |
+4 -4
View File
@@ -34,14 +34,14 @@ None of these numbers — slippage in bps, cost as a fraction of gross, the rati
Throughout this book, backtest metrics carry a warning label, not a hiding place: universe, date window, and whether the hypothesis was pre-registered before the run. This matters because TradeAC ran 31+ experiments; with that many draws, some positive results will be luck. The book is explicit about which runs were pre-registered (e.g. isolation runs exp 28–31) and which were exploratory. `REFERENCED` — multiple-testing/cherry-picking risk is standard research practice; see `references/` as it accrues. Throughout this book, backtest metrics carry a warning label, not a hiding place: universe, date window, and whether the hypothesis was pre-registered before the run. This matters because TradeAC ran 31+ experiments; with that many draws, some positive results will be luck. The book is explicit about which runs were pre-registered (e.g. isolation runs exp 28–31) and which were exploratory. `REFERENCED` — multiple-testing/cherry-picking risk is standard research practice; see `references/` as it accrues.
The most important proof of this discipline is the clean-lake reset, which this book treats as a turning point rather than a footnote: the pre-reset reference signal did **not** reproduce on a rebuilt lake (`EVIDENCE#010 → exp 21`). Had the book quoted the pre-reset backtest as fact, it would have shipped a lie. The trail and the traceability loop are what allowed the desk to catch it. Chapter 05 tells that story in full. The most important proof of this discipline is the clean-lake reset, which this book treats as a turning point rather than a footnote: the pre-reset reference signal did **not** reproduce on a rebuilt lake (`EVIDENCE#010 → exp 21`). Had the book quoted the pre-reset backtest as fact, it would have shipped a lie. The trail and the traceability loop are what allowed the desk to catch it. Chapter 06 tells that story in full.
## How to read this book ## How to read this book
- Every claim is tagged `PROVEN` (traced experiment/round), `HYPOTHESIS` (unreproduced), or `REFERENCED` (external source). `EVIDENCE.md` maps each tag to the run, branch, and round behind it. - Every claim is tagged `PROVEN` (traced experiment/round), `HYPOTHESIS` (unreproduced), or `REFERENCED` (external source). `EVIDENCE.md` maps each tag to the run, branch, and round behind it.
- Chapters 02–09 follow the research arc: what was tested, what was proved, what was refuted, and what moved performance. Refuted runs are cited as evidence too — knowing what *doesn't* work is how the desk avoided paying for it twice. - Chapters 03–10 follow the research arc: what was tested, what was proved, what was refuted, and what moved performance. Refuted runs are cited as evidence too — knowing what *doesn't* work is how the desk avoided paying for it twice.
- Chapter 10 is the reality check: live execution against the research claims. - Chapter 11 is the reality check: live execution against the research claims.
- Chapter 11 is the synthesis: the scoreboard of what actually improved performance and why. - Chapter 12 is the synthesis: the scoreboard of what actually improved performance and why.
## Open questions ## Open questions
@@ -0,0 +1,108 @@
# Chapter 01 — Metrics: The Vocabulary of a Price Series
Status: drafting. Claim inventory: see `README.md` ch. 01.
This chapter exists because the rest of the book argues in a vocabulary that must be shared before it can be trusted. Every claim in every later chapter reduces to a statistic computed on the lake — a drift estimate, an IC, a drawdown, a slippage number. If those statistics are ambiguous, the claims built on them are ambiguous. So this chapter defines the metrics, shows where each one lives in the TradeAC feature store and experiment ledger, and states — claim by claim — what the lake study observed versus what the clean-lake experiments actually proved.
The theme that runs through every metric below: **a statistic is only as good as the falsification it survives.** A drift that vanishes when the data is rebuilt is not a drift; an IC that dies after 5bp of cost is not an edge. The book treats each metric as a hypothesis generator, then runs the hypothesis through the research loop (ch. 02) until it is proved, refuted, or parked.
## The metric → lake → experiment map
| Metric | What it measures | Lake feature | Status in this book |
|--------|------------------|--------------|---------------------|
| Drift / trend | conditional mean of returns over horizon | `sp_trend_slope_5/20/60`, `sp_logp`, `sp_ret_*`, `sp_sharpe_22` | HYPOTHESIS (structure); REFUTED as features (exp 29) |
| Jump | discontinuity share of return variation | `sp_jump_ratio/flag/tail`, `sp_max_up/down/move` | HYPOTHESIS (Peso regime) |
| Volatility clustering | volatility-of-volatility, HAR-RV persistence | `sp_rv1/5/22`, `sp_vol_ratio_*`, `sp_garch_*` | PROVEN as reference features; GARCH trio REFUTED (exp 31) |
| Regime | latent state of the return process | `sp_hmm_p_regime1`, `sp_hmm_state` | REFUTED as features; HYPOTHESIS as overlay |
| Mean reversion | short-horizon reversal strength | `sp_ou_zscore`, `sp_ou_half_life`, `sp_hurst_exponent` | REFUTED as features (exp 25); HYPOTHESIS in isolation |
| Memory | long-range dependence | `sp_hurst_exponent`, `sp_sig_level1/2` | PROVEN as features (exp 24); Hurst magnitude HYPOTHESIS |
| Risk | realized vol ratios, drawdown, liquidity | `sp_vol_ratio_5_22`, net_MDD, RankIC noise floor | PROVEN via exp 26/round 3; liquidity-floor spec needs post-reset rerun |
| Error | prediction quality vs a null | IC, RankIC, ICIR, RankICIR, net IR | PROVEN — the book's scorecard |
| Probability | statistical significance | null z-scores, t-stats, sign agreement | PROVEN as method (exp 21 detection playbook) |
| Timeline | horizon at which a signal holds | 1d/5d labels, IC half-life | PROVEN (5d label) + HYPOTHESIS (decay curves) |
| Decay / reinforcement | signal fading and ensemble averaging | momentum half-decay, seed-count blending | PROVEN (exp 28 seed count); decay curves HYPOTHESIS |
Every row is a bridge: the metric is defined here, its evidence is cited here, and its consequences are worked out in the chapters listed.
## Drift / trend
Drift is the conditional mean of returns — the question "on average, does this price series go somewhere?" The lake measures it two ways: directly as multi-horizon log-price slopes (`sp_trend_slope_5/20/60`, `sp_logp`) and momentum totals (`sp_ret_22/63/126/252`, `sp_sharpe_22`), and implicitly as the mean of the label every model is trained on.
What the dataset study observed (pre-clean-lake, hypothesis material): on the 72-asset panel, only 8 names (QQQ, SMH, SPY, VOO, VTI, DIA, GLD, XAR) showed statistically detectable positive drift at t≥2 — a submartingale at long horizons. But drift explains only ~0.5% of daily variance, and at short horizons the panel mean-reverts (VR<1 at 5–20d for ~32/72 assets) `(HYPOTHESIS → book/references/chat-ideas.md: martingale study, pre-clean-lake; idea only)`.
What the clean-lake experiments proved: drift-as-model-feature failed. The multi-horizon momentum bundle (M1) regressed every metric — IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13% `(PROVEN → exp 29)`. A risk-adjusted 22d Sharpe drift (M2) was mixed and unreproduced — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio net +6.53% (IR 0.62) `(PROVEN run, HYPOTHESIS claim → exp 30)`. The open question is whether short-horizon reversal is tradable net of costs on the clean lake, which no isolation run has yet tested `TODO(evidence-needed: standalone 5d-reversal strategy net of costs)`.
Lesson for later chapters: drift is real structure but it is not, by itself, a feature; its observable manifestation in this campaign was reversal, not momentum (ch. 04, ch. 07).
## Jump
A jump is a discontinuity in the price path — a return too large to be explained by the local diffusion. The lake splits variation with bipower variation: `sp_jump_ratio` (jump share of RV), `sp_jump_tail` (z-scored tail move), `sp_max_up/down/move` (signed extremes).
The dataset study flagged a Peso problem in commodities: USO/UNG show apparent drift (+0.94/+0.55 annualized) that is spike-regime compensation, not carry — the drift is earned in rare jumps and given back in between `(HYPOTHESIS → chat-ideas.md: martingale study; idea only)`. The tradable reading: trend-follow the spikes, do not hold the reversion stanza.
On the clean lake, jump features are part of the reference set that survives `(PROVEN → exp 23/24)`, but no isolation run has tested jump *alone*; the jump-share hypothesis (that the tail-to-diffusion ratio, not raw vol, ranks names) remains unisolated `TODO(evidence-needed: jump-only isolation on the clean lake)`.
## Volatility clustering
Volatility clusters: large moves beget large moves. The lake measures the realized-vol ladder (`sp_rv1/5/22`), its ratios (`sp_vol_ratio_1_22`, `sp_vol_ratio_5_22`), HAR-RV ratios and RV lag-1 autocorrelation, and a GARCH(1,1) MLE (`sp_garch_*`: conditional variance, standardized residual, persistence α+β).
Clustering is the most consistently predictive family in this campaign: the clean-lake compact reference set is built around RV/vol-ratio features `(PROVEN → exp 24)`, and the earliest tree splits of the reference model are dominated by realized-vol features `(HYPOTHESIS → chat-ideas.md: single-feature study; feature importance is from the pre-reset tree)`.
But vol clustering is not the same as *vol-regime modeling*. The GARCH(1,1) trio, added to the reference, was refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31)`. The pattern repeated: the generic realized-vol ladder contributes; the parametric vol model does not (ch. 07). This is the book's recurring lesson — **generic, scale-free, well-behaved statistics beat model-specific machinery on a small daily panel.**
## Regime
A regime is a latent state of the return process — a two-state Gaussian HMM is fit on returns, and `sp_hmm_p_regime1` / `sp_hmm_state` carry the posterior.
Regime flags failed as model features twice (exp 9 pre-reset idea, exp 25 clean-lake confirmation that model-specific families regress the signal) `(PROVEN → exp 25; the exp 9 idea is pre-reset idea material)`. The surviving hypothesis is that regime belongs **overlay, not feature**: a long-only/regime-gate that holds names only in the favourable state `(HYPOTHESIS → chat-ideas.md; untested on the clean lake)`. What would settle it: a gated version of the exp-26 n_drop=1 book compared against the ungated book over the same window `TODO(evidence-needed: HMM regime gate as overlay on exp-26 book)`.
## Mean reversion
Mean reversion is the flip side of drift: short-horizon reversal. The lake measures it as `sp_ou_zscore` (distance from a fitted OU/AR(1) mean), `sp_ou_half_life` (mean-reversion speed), and via Hurst < 0.5.
The single-feature study found `sp_ou_zscore` the strongest stable standalone predictor (IC −0.15/−0.13, sign-stable across years) `(HYPOTHESIS → chat-ideas.md: single-feature study; idea only)`. Yet adding it to the reference model regressed every metric — IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25)`. This is the book's named open problem, the "OU paradox": the strongest single-feature signal is worthless — worse, harmful — inside a cross-sectional rank model `(see chat-ideas.md, TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`. The working explanation (hypothesis): the OU z-score carries name-specific scale that survives CSRankNorm poorly and collides with the vol/trend families the model already uses.
## Memory and the path signature
Memory is long-range dependence: a return's persistence beyond the short horizon. The lake measures Hurst exponent (R/S) and the path signature (lead/lag integrals of log-price path, levels 1–2 at lag 1 and 5).
The dataset study found mild persistence across the panel (H ≈ 0.54–0.63), i.e. neither strong trend nor strong mean-reversion at the measured lags `(HYPOTHESIS → chat-ideas.md: martingale study; idea only)`. Signatures are the *generic* memory feature and they are load-bearing on the clean lake: `sp_sig_level1/2` are part of the compact reference set `(PROVEN → exp 24)`. Memory's practical meaning in this book: persistence is weak and horizon-dependent, so the signal must be refreshed on a short label and turned over carefully — which is exactly the cost argument of ch. 03 and ch. 09.
## Risk
Risk here is the denominator of every edge: realized vol ratios (`sp_vol_ratio_5_22`), drawdown, and the noise floor of the rank measurement itself.
Two facts about risk matter throughout the book. First, the measurement floor: on a 50-name cross-section the daily RankIC null std is 1/√(N−1) ≈ 0.143, so a mean RankIC near 0.06 is a small-but-real edge sitting on a wide null — every performance claim in this book is read against that floor `(method, PROVEN via the exp-21 detection playbook; also REFERENCED for the rank-null statistic)`. Second, the binding constraint: on clean data the gross→net collapse is ~9–10pp of cost drag, and IR ≈ 0.21 net is the campaign's best result `(PROVEN → exp 26)`. Risk limits — liquidity floor, size caps, drawdown pause — gate the live book, but the pre-reset evidence that the floor beats caps needs a post-reset rerun before it can be cited as fact `(HYPOTHESIS → exp 18 pre-clean-lake; TODO(evidence-needed: risk-limit A/B on the exp-26 reference))` (ch. 10).
## Error: the scorecard that separates hypothesis from proof
Error is the book's discipline: how wrong was the prediction, in a way that can be measured against a null? The canonical metrics on the clean lake are IC, ICIR, Rank IC, Rank ICIR (the rank-based signal quality) and net IR, net return, L/S Sharpe, max drawdown (the portfolio outcome). Pre-reset runs (exp 8–18) recorded a different schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ir_with_cost`) — the two schemas are never compared directly in this book `(EVIDENCE.md: metric-schema note)`.
Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); live results are reconciled with slippage and cost, not taken from the backtest (ch. 11). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26, round 3)`.
## Probability and significance
Every statistic in this book carries a significance discipline: t-stats on pooled drift regressions (t≥2 for the submartingale reads), null-baseline z-scores for per-day IC (3–4σ single-day ICs were the contamination fingerprint that exposed the dirty lake), and sign-agreement across symbols/years `(method PROVEN via the exp-21 detection playbook; magnitude claims from the martingale study are HYPOTHESIS)`.
The point is procedural: a metric without a null hypothesis is a number, not evidence. The book's probability posture is that ~50-name panels give weak statistical power, so a single improved run is a hypothesis until reproduced — exp 30 (M2) is explicitly labeled HYPOTHESIS for exactly this reason `(PROVEN run, HYPOTHESIS claim → exp 30)`.
## Timeline, decay, reinforcement
The last cluster is about time. The lake's label is 5-day forward return — the 5d horizon is the campaign's best IC lever `(HYPOTHESIS → chat-ideas.md; the 5d label choice predates the clean lake)`, and horizon matters: 5d sees reversal that a 1d label blurs and a 63/126d label can't distinguish from drift. Decay shows up three ways: signal decay (a 5d reversal signal is stale after its horizon — this is why turnover relief, not signal engineering, was the biggest lever), feature decay (momentum/short-slope features weakened from 2025 to 2026 in the single-feature study `(HYPOTHESIS → chat-ideas.md)`), and cost decay (every holding day the cost drag compounds against a thin edge — the n_drop 2→1 result is the book's cleanest example: identical IC/RankIC, net flips from −3.21% to +2.13% purely from holding the dropped name `(PROVEN → exp 26)`).
Reinforcement is the positive half of decay: ensemble averaging. The 5-seed blend raises net performance on clean data, and seed count is load-bearing — 2 seeds lose to 5 (RankIC 0.0579 vs 0.0663, net −1.49% vs +2.13%) `(PROVEN → exp 28)`. Averaging is reinforcement against noise, not against cost; ch. 05 works out the mechanism.
## From statistic to hypothesis to proved practice
The cycle this book runs on: **measure → hypothesize → pre-register → isolate → prove or refute → reconcile live.**
1. **Measure** — a lake study turns a metric into a number (this chapter's vocabulary; dataset studies in `book/data/`).
2. **Hypothesize** — the number becomes a falsifiable claim ("adding risk-adjusted drift helps"), recorded in the run notes before the run.
3. **Pre-register** — the claim and its acceptance metric (IC/RankIC above reference, net IR above reference) are fixed before execution, to block post-hoc cherry-picking across the 31+ experiments.
4. **Isolate** — one variable changes per run; the reference book and its metrics are the control (exp 26 → 29/30/31).
5. **Prove or refute** — on the clean lake only. A reproduced improvement becomes PROVEN; a single un-reproduced run stays HYPOTHESIS (exp 30); a degradation is REFUTED and — critically — is recorded as a win for the discipline (exp 29, exp 31 stopped wrong directions).
6. **Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 11).
Every metric in this chapter sits on this loop. The drift metric produced the momentum hypothesis and the reversal hypothesis; only one survived isolation. The error metrics are the loop's judge. The decay and risk metrics are why ch. 03 and ch. 09 exist at all. The rest of the book is the working-out of this cycle, claim by claim, with each claim traceable to `EVIDENCE.md` and a recorded run.
Open questions for the desk (see `README.md`): why single-feature OU IC does not survive inside the model; whether 5d reversal trades net of cost; whether the HMM regime overlay beats the ungated book; whether M2 reproduction holds; and whether the 50-ETF panel generalizes.
+34
View File
@@ -0,0 +1,34 @@
# Chapter 02 — The Research Loop: Lake → Experiment → Live
Status: drafting. Claim inventory: see `README.md` ch. 02.
The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. This chapter is the machinery that makes them trustworthy: the traced research loop that turns a hypothesis into a proved practice. It is the book's methodology chapter, and it is also the book's proof-of-work — every later chapter is a walkthrough of this loop on a concrete question.
## The loop
1. **Lake** — bars and features live in one hive-partitioned lake (market/timeframe/symbol, `family=ta|sp`), with coverage and calendar metadata. It is the single source of bar/feature truth. The clean-lake rebuild proved the stakes: when the lake was rebuilt, the reference signal collapsed (IC 0.0354 → 0.0019) because the old lake's data quality had silently inflated results `(PROVEN → exp 21)`. A claim built on the lake is only as good as the lake.
2. **Experiment** — every run is a traced experiment: a git branch (`exp/N-…`), an MLflow run with recorded config/params/metrics, and hypothesis/evaluation notes recorded before and after the run. The traceability loop was used on exp 8–31; the branch, run, and notes are the reproducible unit `(PROVEN → the traced experiment store; see `rd_exp_*` tools and `EVIDENCE.md`)`.
3. **Live** — a proved book advances to a round window (targets → intents → decisions → orders → fills), and is reconciled (slippage bps, cost, funnel) `(PROVEN → round 3; tac-rd-book trail)`. Live beats backtest: a claim about trading performance must trace to a round, not to a backtest (ch. 11).
## Why traceability is the methodology
TradeAC ran 31+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:
- **The dirty lake.** Pre-reset exp 8–18 reported strong results that did not survive a clean rebuild `(PROVEN → exp 21)`. Only because the exact YAML, branch, and run were recorded could the desk reproduce — and falsify — the reference. Traceability is what turned a false belief into evidence.
- **Post-hoc cherry-picking.** With 31+ experiments, the best-looking number is expected to be inflated by selection. The counter is pre-registration: hypothesis, change, and acceptance metric are fixed in the run notes *before* the run `(REFERENCED — research practice; see CLAIMS.md multiple-testing note)`. Where the book quotes an experiment whose hypothesis was recorded after the fact, it says so.
The pre-reset experiments (exp 8–18) are therefore treated as **idea material, not fact**: they were demonstrably inflated by lake data quality `(EVIDENCE.md pre-clean-lake section)`. Their ideas feed the hypothesis pipeline (ch. 01); their numbers never stand alone.
## The isolation discipline
A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio. Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.
## Falsification is the output
Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.
## From here
Ch. 03 applies the loop to the book's first worked question (does the signal clear costs?), ch. 06 to the clean-lake reset, and ch. 11 to the live round that closes the loop with reconciliation.
Open questions: purge/walk-forward CV instead of single train/valid split `TODO(evidence-needed: purged CV on the exp-26 reference)`, and a PSI-based drift-aware retraining gate `(HYPOTHESIS → chat-ideas.md)`.
@@ -1,6 +1,6 @@
# Chapter 02 — Baseline and the Cost Reality # Chapter 03 — Baseline and the Cost Reality
Status: drafting. Claim inventory: see `README.md` ch. 02. Status: drafting. Claim inventory: see `README.md` ch. 03.
This chapter answers the question every quant desk must answer before the first dollar is deployed: **what does the raw signal have to be worth, and what survives the cost of trading it?** This chapter answers the question every quant desk must answer before the first dollar is deployed: **what does the raw signal have to be worth, and what survives the cost of trading it?**
@@ -40,7 +40,7 @@ This is the single most important number in the early book: **at this turnover,
The only construction change that flipped net from negative to positive was reducing daily forced replacements from `n_drop=2` to `n_drop=1` — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`. The only construction change that flipped net from negative to positive was reducing daily forced replacements from `n_drop=2` to `n_drop=1` — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`.
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 06) then confirmed that most candidate additions *reduced* the edge anyway. Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 07) then confirmed that most candidate additions *reduced* the edge anyway.
## Desk rules distilled from this chapter ## Desk rules distilled from this chapter
@@ -1,6 +1,6 @@
# Chapter 05 — The Clean-Lake Reset: Data Quality as First-Order Risk # Chapter 06 — The Clean-Lake Reset: Data Quality as First-Order Risk
Status: drafting. Claim inventory: see `README.md` ch. 05. Status: drafting. Claim inventory: see `README.md` ch. 06.
Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on. Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on.
+3 -3
View File
@@ -2,7 +2,7 @@
Source: opencode chat transcripts under `book/data/chat_mining/` (historical context, pre-clean-lake). Per the evidence contract these are **idea material only** — none may be cited as `PROVEN`. Each idea below is a hypothesis to be tested on the clean lake (exp 21+). Source: opencode chat transcripts under `book/data/chat_mining/` (historical context, pre-clean-lake). Per the evidence contract these are **idea material only** — none may be cited as `PROVEN`. Each idea below is a hypothesis to be tested on the clean lake (exp 21+).
## Data-quality failure classes (feed ch. 05) ## Data-quality failure classes (feed ch. 06)
These are the *classes* of failure documented across `exp-polluted-lake.txt`, `exp-dirty-lake.txt`, `cleaned-lake.txt`. Durable lessons even though exact numbers are pre-reset. These are the *classes* of failure documented across `exp-polluted-lake.txt`, `exp-dirty-lake.txt`, `cleaned-lake.txt`. Durable lessons even though exact numbers are pre-reset.
@@ -13,7 +13,7 @@ These are the *classes* of failure documented across `exp-polluted-lake.txt`, `e
5. **Mid-experiment regeneration.** Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files. 5. **Mid-experiment regeneration.** Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files.
6. **Detection playbook** (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines. 6. **Detection playbook** (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines.
## Market-structure hypotheses (feed ch. 03/06; from martingale study + clean-data study) ## Market-structure hypotheses (feed ch. 04/07; from martingale study + clean-data study)
- **Submartingale at long horizons, mean-reverting at short horizons.** Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation. - **Submartingale at long horizons, mean-reverting at short horizons.** Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation.
- **5-day momentum strongly reverses** (pooled regression: `sp_trend_slope_5` β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever. - **5-day momentum strongly reverses** (pooled regression: `sp_trend_slope_5` β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever.
@@ -21,7 +21,7 @@ These are the *classes* of failure documented across `exp-polluted-lake.txt`, `e
- **HMM regime gating as an overlay, not a feature.** Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested. - **HMM regime gating as an overlay, not a feature.** Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested.
- **Edge is long-short, not long-only** (drift is mostly common/market-wide). - **Edge is long-short, not long-only** (drift is mostly common/market-wide).
## Feature methodology hypotheses (feed ch. 03/06) ## Feature methodology hypotheses (feed ch. 04/07)
- **Panel width vs feature count:** three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM. - **Panel width vs feature count:** three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM.
- **Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model.** `sp_ou_zscore` was the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed). - **Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model.** `sp_ou_zscore` was the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed).