@@ -34,14 +34,14 @@ None of these numbers — slippage in bps, cost as a fraction of gross, the rati
Throughout this book, backtest metrics carry a warning label, not a hiding place: universe, date window, and whether the hypothesis was pre-registered before the run. This matters because TradeAC ran 31+ experiments; with that many draws, some positive results will be luck. The book is explicit about which runs were pre-registered (e.g. isolation runs exp 28–31) and which were exploratory. `REFERENCED` — multiple-testing/cherry-picking risk is standard research practice; see `references/` as it accrues.
The most important proof of this discipline is the clean-lake reset, which this book treats as a turning point rather than a footnote: the pre-reset reference signal did **not** reproduce on a rebuilt lake (`EVIDENCE#010 → exp 21`). Had the book quoted the pre-reset backtest as fact, it would have shipped a lie. The trail and the traceability loop are what allowed the desk to catch it. Chapter 05 tells that story in full.
The most important proof of this discipline is the clean-lake reset, which this book treats as a turning point rather than a footnote: the pre-reset reference signal did **not** reproduce on a rebuilt lake (`EVIDENCE#010 → exp 21`). Had the book quoted the pre-reset backtest as fact, it would have shipped a lie. The trail and the traceability loop are what allowed the desk to catch it. Chapter 06 tells that story in full.
## How to read this book
- Every claim is tagged `PROVEN` (traced experiment/round), `HYPOTHESIS` (unreproduced), or `REFERENCED` (external source). `EVIDENCE.md` maps each tag to the run, branch, and round behind it.
- Chapters 02–09 follow the research arc: what was tested, what was proved, what was refuted, and what moved performance. Refuted runs are cited as evidence too — knowing what *doesn't* work is how the desk avoided paying for it twice.
- Chapter 10 is the reality check: live execution against the research claims.
- Chapter 11 is the synthesis: the scoreboard of what actually improved performance and why.
- Chapters 03–10 follow the research arc: what was tested, what was proved, what was refuted, and what moved performance. Refuted runs are cited as evidence too — knowing what *doesn't* work is how the desk avoided paying for it twice.
- Chapter 11 is the reality check: live execution against the research claims.
- Chapter 12 is the synthesis: the scoreboard of what actually improved performance and why.
# Chapter 01 — Metrics: The Vocabulary of a Price Series
Status: drafting. Claim inventory: see `README.md` ch. 01.
This chapter exists because the rest of the book argues in a vocabulary that must be shared before it can be trusted. Every claim in every later chapter reduces to a statistic computed on the lake — a drift estimate, an IC, a drawdown, a slippage number. If those statistics are ambiguous, the claims built on them are ambiguous. So this chapter defines the metrics, shows where each one lives in the TradeAC feature store and experiment ledger, and states — claim by claim — what the lake study observed versus what the clean-lake experiments actually proved.
The theme that runs through every metric below: **a statistic is only as good as the falsification it survives.** A drift that vanishes when the data is rebuilt is not a drift; an IC that dies after 5bp of cost is not an edge. The book treats each metric as a hypothesis generator, then runs the hypothesis through the research loop (ch. 02) until it is proved, refuted, or parked.
## The metric → lake → experiment map
| Metric | What it measures | Lake feature | Status in this book |
| Drift / trend | conditional mean of returns over horizon | `sp_trend_slope_5/20/60`, `sp_logp`, `sp_ret_*`, `sp_sharpe_22` | HYPOTHESIS (structure); REFUTED as features (exp 29) |
Every row is a bridge: the metric is defined here, its evidence is cited here, and its consequences are worked out in the chapters listed.
## Drift / trend
Drift is the conditional mean of returns — the question "on average, does this price series go somewhere?" The lake measures it two ways: directly as multi-horizon log-price slopes (`sp_trend_slope_5/20/60`, `sp_logp`) and momentum totals (`sp_ret_22/63/126/252`, `sp_sharpe_22`), and implicitly as the mean of the label every model is trained on.
What the dataset study observed (pre-clean-lake, hypothesis material): on the 72-asset panel, only 8 names (QQQ, SMH, SPY, VOO, VTI, DIA, GLD, XAR) showed statistically detectable positive drift at t≥2 — a submartingale at long horizons. But drift explains only ~0.5% of daily variance, and at short horizons the panel mean-reverts (VR<1 at 5–20d for ~32/72 assets) `(HYPOTHESIS → book/references/chat-ideas.md: martingale study, pre-clean-lake; idea only)`.
What the clean-lake experiments proved: drift-as-model-feature failed. The multi-horizon momentum bundle (M1) regressed every metric — IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13% `(PROVEN → exp 29)`. A risk-adjusted 22d Sharpe drift (M2) was mixed and unreproduced — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio net +6.53% (IR 0.62) `(PROVEN run, HYPOTHESIS claim → exp 30)`. The open question is whether short-horizon reversal is tradable net of costs on the clean lake, which no isolation run has yet tested `TODO(evidence-needed: standalone 5d-reversal strategy net of costs)`.
Lesson for later chapters: drift is real structure but it is not, by itself, a feature; its observable manifestation in this campaign was reversal, not momentum (ch. 04, ch. 07).
## Jump
A jump is a discontinuity in the price path — a return too large to be explained by the local diffusion. The lake splits variation with bipower variation: `sp_jump_ratio` (jump share of RV), `sp_jump_tail` (z-scored tail move), `sp_max_up/down/move` (signed extremes).
The dataset study flagged a Peso problem in commodities: USO/UNG show apparent drift (+0.94/+0.55 annualized) that is spike-regime compensation, not carry — the drift is earned in rare jumps and given back in between `(HYPOTHESIS → chat-ideas.md: martingale study; idea only)`. The tradable reading: trend-follow the spikes, do not hold the reversion stanza.
On the clean lake, jump features are part of the reference set that survives `(PROVEN → exp 23/24)`, but no isolation run has tested jump *alone*; the jump-share hypothesis (that the tail-to-diffusion ratio, not raw vol, ranks names) remains unisolated `TODO(evidence-needed: jump-only isolation on the clean lake)`.
## Volatility clustering
Volatility clusters: large moves beget large moves. The lake measures the realized-vol ladder (`sp_rv1/5/22`), its ratios (`sp_vol_ratio_1_22`, `sp_vol_ratio_5_22`), HAR-RV ratios and RV lag-1 autocorrelation, and a GARCH(1,1) MLE (`sp_garch_*`: conditional variance, standardized residual, persistence α+β).
Clustering is the most consistently predictive family in this campaign: the clean-lake compact reference set is built around RV/vol-ratio features `(PROVEN → exp 24)`, and the earliest tree splits of the reference model are dominated by realized-vol features `(HYPOTHESIS → chat-ideas.md: single-feature study; feature importance is from the pre-reset tree)`.
But vol clustering is not the same as *vol-regime modeling*. The GARCH(1,1) trio, added to the reference, was refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31)`. The pattern repeated: the generic realized-vol ladder contributes; the parametric vol model does not (ch. 07). This is the book's recurring lesson — **generic, scale-free, well-behaved statistics beat model-specific machinery on a small daily panel.**
## Regime
A regime is a latent state of the return process — a two-state Gaussian HMM is fit on returns, and `sp_hmm_p_regime1` / `sp_hmm_state` carry the posterior.
Regime flags failed as model features twice (exp 9 pre-reset idea, exp 25 clean-lake confirmation that model-specific families regress the signal) `(PROVEN → exp 25; the exp 9 idea is pre-reset idea material)`. The surviving hypothesis is that regime belongs **overlay, not feature**: a long-only/regime-gate that holds names only in the favourable state `(HYPOTHESIS → chat-ideas.md; untested on the clean lake)`. What would settle it: a gated version of the exp-26 n_drop=1 book compared against the ungated book over the same window `TODO(evidence-needed: HMM regime gate as overlay on exp-26 book)`.
## Mean reversion
Mean reversion is the flip side of drift: short-horizon reversal. The lake measures it as `sp_ou_zscore` (distance from a fitted OU/AR(1) mean), `sp_ou_half_life` (mean-reversion speed), and via Hurst < 0.5.
The single-feature study found `sp_ou_zscore` the strongest stable standalone predictor (IC −0.15/−0.13, sign-stable across years) `(HYPOTHESIS → chat-ideas.md: single-feature study; idea only)`. Yet adding it to the reference model regressed every metric — IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25)`. This is the book's named open problem, the "OU paradox": the strongest single-feature signal is worthless — worse, harmful — inside a cross-sectional rank model `(see chat-ideas.md, TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`. The working explanation (hypothesis): the OU z-score carries name-specific scale that survives CSRankNorm poorly and collides with the vol/trend families the model already uses.
## Memory and the path signature
Memory is long-range dependence: a return's persistence beyond the short horizon. The lake measures Hurst exponent (R/S) and the path signature (lead/lag integrals of log-price path, levels 1–2 at lag 1 and 5).
The dataset study found mild persistence across the panel (H ≈ 0.54–0.63), i.e. neither strong trend nor strong mean-reversion at the measured lags `(HYPOTHESIS → chat-ideas.md: martingale study; idea only)`. Signatures are the *generic* memory feature and they are load-bearing on the clean lake: `sp_sig_level1/2` are part of the compact reference set `(PROVEN → exp 24)`. Memory's practical meaning in this book: persistence is weak and horizon-dependent, so the signal must be refreshed on a short label and turned over carefully — which is exactly the cost argument of ch. 03 and ch. 09.
## Risk
Risk here is the denominator of every edge: realized vol ratios (`sp_vol_ratio_5_22`), drawdown, and the noise floor of the rank measurement itself.
Two facts about risk matter throughout the book. First, the measurement floor: on a 50-name cross-section the daily RankIC null std is 1/√(N−1) ≈ 0.143, so a mean RankIC near 0.06 is a small-but-real edge sitting on a wide null — every performance claim in this book is read against that floor `(method, PROVEN via the exp-21 detection playbook; also REFERENCED for the rank-null statistic)`. Second, the binding constraint: on clean data the gross→net collapse is ~9–10pp of cost drag, and IR ≈ 0.21 net is the campaign's best result `(PROVEN → exp 26)`. Risk limits — liquidity floor, size caps, drawdown pause — gate the live book, but the pre-reset evidence that the floor beats caps needs a post-reset rerun before it can be cited as fact `(HYPOTHESIS → exp 18 pre-clean-lake; TODO(evidence-needed: risk-limit A/B on the exp-26 reference))` (ch. 10).
## Error: the scorecard that separates hypothesis from proof
Error is the book's discipline: how wrong was the prediction, in a way that can be measured against a null? The canonical metrics on the clean lake are IC, ICIR, Rank IC, Rank ICIR (the rank-based signal quality) and net IR, net return, L/S Sharpe, max drawdown (the portfolio outcome). Pre-reset runs (exp 8–18) recorded a different schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ir_with_cost`) — the two schemas are never compared directly in this book `(EVIDENCE.md: metric-schema note)`.
Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); live results are reconciled with slippage and cost, not taken from the backtest (ch. 11). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26, round 3)`.
## Probability and significance
Every statistic in this book carries a significance discipline: t-stats on pooled drift regressions (t≥2 for the submartingale reads), null-baseline z-scores for per-day IC (3–4σ single-day ICs were the contamination fingerprint that exposed the dirty lake), and sign-agreement across symbols/years `(method PROVEN via the exp-21 detection playbook; magnitude claims from the martingale study are HYPOTHESIS)`.
The point is procedural: a metric without a null hypothesis is a number, not evidence. The book's probability posture is that ~50-name panels give weak statistical power, so a single improved run is a hypothesis until reproduced — exp 30 (M2) is explicitly labeled HYPOTHESIS for exactly this reason `(PROVEN run, HYPOTHESIS claim → exp 30)`.
## Timeline, decay, reinforcement
The last cluster is about time. The lake's label is 5-day forward return — the 5d horizon is the campaign's best IC lever `(HYPOTHESIS → chat-ideas.md; the 5d label choice predates the clean lake)`, and horizon matters: 5d sees reversal that a 1d label blurs and a 63/126d label can't distinguish from drift. Decay shows up three ways: signal decay (a 5d reversal signal is stale after its horizon — this is why turnover relief, not signal engineering, was the biggest lever), feature decay (momentum/short-slope features weakened from 2025 to 2026 in the single-feature study `(HYPOTHESIS → chat-ideas.md)`), and cost decay (every holding day the cost drag compounds against a thin edge — the n_drop 2→1 result is the book's cleanest example: identical IC/RankIC, net flips from −3.21% to +2.13% purely from holding the dropped name `(PROVEN → exp 26)`).
Reinforcement is the positive half of decay: ensemble averaging. The 5-seed blend raises net performance on clean data, and seed count is load-bearing — 2 seeds lose to 5 (RankIC 0.0579 vs 0.0663, net −1.49% vs +2.13%) `(PROVEN → exp 28)`. Averaging is reinforcement against noise, not against cost; ch. 05 works out the mechanism.
## From statistic to hypothesis to proved practice
The cycle this book runs on: **measure → hypothesize → pre-register → isolate → prove or refute → reconcile live.**
1.**Measure** — a lake study turns a metric into a number (this chapter's vocabulary; dataset studies in `book/data/`).
2.**Hypothesize** — the number becomes a falsifiable claim ("adding risk-adjusted drift helps"), recorded in the run notes before the run.
3.**Pre-register** — the claim and its acceptance metric (IC/RankIC above reference, net IR above reference) are fixed before execution, to block post-hoc cherry-picking across the 31+ experiments.
4.**Isolate** — one variable changes per run; the reference book and its metrics are the control (exp 26 → 29/30/31).
5.**Prove or refute** — on the clean lake only. A reproduced improvement becomes PROVEN; a single un-reproduced run stays HYPOTHESIS (exp 30); a degradation is REFUTED and — critically — is recorded as a win for the discipline (exp 29, exp 31 stopped wrong directions).
6.**Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 11).
Every metric in this chapter sits on this loop. The drift metric produced the momentum hypothesis and the reversal hypothesis; only one survived isolation. The error metrics are the loop's judge. The decay and risk metrics are why ch. 03 and ch. 09 exist at all. The rest of the book is the working-out of this cycle, claim by claim, with each claim traceable to `EVIDENCE.md` and a recorded run.
Open questions for the desk (see `README.md`): why single-feature OU IC does not survive inside the model; whether 5d reversal trades net of cost; whether the HMM regime overlay beats the ungated book; whether M2 reproduction holds; and whether the 50-ETF panel generalizes.
# Chapter 02 — The Research Loop: Lake → Experiment → Live
Status: drafting. Claim inventory: see `README.md` ch. 02.
The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. This chapter is the machinery that makes them trustworthy: the traced research loop that turns a hypothesis into a proved practice. It is the book's methodology chapter, and it is also the book's proof-of-work — every later chapter is a walkthrough of this loop on a concrete question.
## The loop
1.**Lake** — bars and features live in one hive-partitioned lake (market/timeframe/symbol, `family=ta|sp`), with coverage and calendar metadata. It is the single source of bar/feature truth. The clean-lake rebuild proved the stakes: when the lake was rebuilt, the reference signal collapsed (IC 0.0354 → 0.0019) because the old lake's data quality had silently inflated results `(PROVEN → exp 21)`. A claim built on the lake is only as good as the lake.
2.**Experiment** — every run is a traced experiment: a git branch (`exp/N-…`), an MLflow run with recorded config/params/metrics, and hypothesis/evaluation notes recorded before and after the run. The traceability loop was used on exp 8–31; the branch, run, and notes are the reproducible unit `(PROVEN → the traced experiment store; see `rd_exp_*` tools and `EVIDENCE.md`)`.
3.**Live** — a proved book advances to a round window (targets → intents → decisions → orders → fills), and is reconciled (slippage bps, cost, funnel) `(PROVEN → round 3; tac-rd-book trail)`. Live beats backtest: a claim about trading performance must trace to a round, not to a backtest (ch. 11).
## Why traceability is the methodology
TradeAC ran 31+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:
- **The dirty lake.** Pre-reset exp 8–18 reported strong results that did not survive a clean rebuild `(PROVEN → exp 21)`. Only because the exact YAML, branch, and run were recorded could the desk reproduce — and falsify — the reference. Traceability is what turned a false belief into evidence.
- **Post-hoc cherry-picking.** With 31+ experiments, the best-looking number is expected to be inflated by selection. The counter is pre-registration: hypothesis, change, and acceptance metric are fixed in the run notes *before* the run `(REFERENCED — research practice; see CLAIMS.md multiple-testing note)`. Where the book quotes an experiment whose hypothesis was recorded after the fact, it says so.
The pre-reset experiments (exp 8–18) are therefore treated as **idea material, not fact**: they were demonstrably inflated by lake data quality `(EVIDENCE.md pre-clean-lake section)`. Their ideas feed the hypothesis pipeline (ch. 01); their numbers never stand alone.
## The isolation discipline
A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio. Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.
## Falsification is the output
Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.
## From here
Ch. 03 applies the loop to the book's first worked question (does the signal clear costs?), ch. 06 to the clean-lake reset, and ch. 11 to the live round that closes the loop with reconciliation.
Open questions: purge/walk-forward CV instead of single train/valid split `TODO(evidence-needed: purged CV on the exp-26 reference)`, and a PSI-based drift-aware retraining gate `(HYPOTHESIS → chat-ideas.md)`.
Status: drafting. Claim inventory: see `README.md` ch. 02.
Status: drafting. Claim inventory: see `README.md` ch. 03.
This chapter answers the question every quant desk must answer before the first dollar is deployed: **what does the raw signal have to be worth, and what survives the cost of trading it?**
@@ -40,7 +40,7 @@ This is the single most important number in the early book: **at this turnover,
The only construction change that flipped net from negative to positive was reducing daily forced replacements from `n_drop=2` to `n_drop=1` — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`.
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 06) then confirmed that most candidate additions *reduced* the edge anyway.
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 07) then confirmed that most candidate additions *reduced* the edge anyway.
# Chapter 05 — The Clean-Lake Reset: Data Quality as First-Order Risk
# Chapter 06 — The Clean-Lake Reset: Data Quality as First-Order Risk
Status: drafting. Claim inventory: see `README.md` ch. 05.
Status: drafting. Claim inventory: see `README.md` ch. 06.
Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on.
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.