# Chapter 01 — Metrics: The Vocabulary of a Price Series Status: drafting. Claim inventory: see `README.md` ch. 01. This chapter exists because the rest of the book argues in a vocabulary that must be shared before it can be trusted. Every claim in every later chapter reduces to a statistic computed on the lake — a drift estimate, an IC, a drawdown, a slippage number. If those statistics are ambiguous, the claims built on them are ambiguous. So this chapter defines the metrics, shows where each one lives in the TradeAC feature store and experiment ledger, and states — claim by claim — what the lake study observed versus what the clean-lake experiments actually proved. The theme that runs through every metric below: **a statistic is only as good as the falsification it survives.** A drift that vanishes when the data is rebuilt is not a drift; an IC that dies after 5bp of cost is not an edge. The book treats each metric as a hypothesis generator, then runs the hypothesis through the research loop (ch. 02) until it is proved, refuted, or parked. ## The metric → lake → experiment map | Metric | What it measures | Lake feature | Status in this book | |--------|------------------|--------------|---------------------| | Drift / trend | conditional mean of returns over horizon | `sp_trend_slope_5/20/60`, `sp_logp`, `sp_ret_*`, `sp_sharpe_22` | HYPOTHESIS (structure); REFUTED as features (exp 29) | | Jump | discontinuity share of return variation | `sp_jump_ratio/flag/tail`, `sp_max_up/down/move` | HYPOTHESIS (Peso regime) | | Volatility clustering | volatility-of-volatility, HAR-RV persistence | `sp_rv1/5/22`, `sp_vol_ratio_*`, `sp_garch_*` | PROVEN as reference features; GARCH trio REFUTED (exp 31) | | Regime | latent state of the return process | `sp_hmm_p_regime1`, `sp_hmm_state` | REFUTED as features; HYPOTHESIS as overlay | | Mean reversion | short-horizon reversal strength | `sp_ou_zscore`, `sp_ou_half_life`, `sp_hurst_exponent` | REFUTED as features (exp 25); HYPOTHESIS in isolation | | Memory | long-range dependence | `sp_hurst_exponent`, `sp_sig_level1/2` | PROVEN as features (exp 24); Hurst magnitude HYPOTHESIS | | Risk | realized vol ratios, drawdown, liquidity | `sp_vol_ratio_5_22`, net_MDD, RankIC noise floor | PROVEN via exp 26/round 3; liquidity-floor spec needs post-reset rerun | | Error | prediction quality vs a null | IC, RankIC, ICIR, RankICIR, net IR | PROVEN — the book's scorecard | | Probability | statistical significance | null z-scores, t-stats, sign agreement | PROVEN as method (exp 21 detection playbook) | | Timeline | horizon at which a signal holds | 1d/5d labels, IC half-life | PROVEN (5d label) + HYPOTHESIS (decay curves) | | Decay / reinforcement | signal fading and ensemble averaging | momentum half-decay, seed-count blending | PROVEN (exp 28 seed count); decay curves HYPOTHESIS | Every row is a bridge: the metric is defined here, its evidence is cited here, and its consequences are worked out in the chapters listed. ## Drift / trend Drift is the conditional mean of returns — the question "on average, does this price series go somewhere?" The lake measures it two ways: directly as multi-horizon log-price slopes (`sp_trend_slope_5/20/60`, `sp_logp`) and momentum totals (`sp_ret_22/63/126/252`, `sp_sharpe_22`), and implicitly as the mean of the label every model is trained on. What the dataset study observed (pre-clean-lake, hypothesis material): on the 72-asset panel, only 8 names (QQQ, SMH, SPY, VOO, VTI, DIA, GLD, XAR) showed statistically detectable positive drift at t≥2 — a submartingale at long horizons. But drift explains only ~0.5% of daily variance, and at short horizons the panel mean-reverts (VR<1 at 5–20d for ~32/72 assets) `(HYPOTHESIS → book/references/chat-ideas.md: martingale study, pre-clean-lake; idea only)`. What the clean-lake experiments proved: drift-as-model-feature failed. The multi-horizon momentum bundle (M1) regressed every metric — IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13% `(PROVEN → exp 29)`. A risk-adjusted 22d Sharpe drift (M2) was mixed and unreproduced — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio net +6.53% (IR 0.62) `(PROVEN run, HYPOTHESIS claim → exp 30)`. The open question is whether short-horizon reversal is tradable net of costs on the clean lake, which no isolation run has yet tested `TODO(evidence-needed: standalone 5d-reversal strategy net of costs)`. Lesson for later chapters: drift is real structure but it is not, by itself, a feature; its observable manifestation in this campaign was reversal, not momentum (ch. 04, ch. 07). ## Jump A jump is a discontinuity in the price path — a return too large to be explained by the local diffusion. The lake splits variation with bipower variation: `sp_jump_ratio` (jump share of RV), `sp_jump_tail` (z-scored tail move), `sp_max_up/down/move` (signed extremes). The dataset study flagged a Peso problem in commodities: USO/UNG show apparent drift (+0.94/+0.55 annualized) that is spike-regime compensation, not carry — the drift is earned in rare jumps and given back in between `(HYPOTHESIS → chat-ideas.md: martingale study; idea only)`. The tradable reading: trend-follow the spikes, do not hold the reversion stanza. On the clean lake, jump features are part of the reference set that survives `(PROVEN → exp 23/24)`, but no isolation run has tested jump *alone*; the jump-share hypothesis (that the tail-to-diffusion ratio, not raw vol, ranks names) remains unisolated `TODO(evidence-needed: jump-only isolation on the clean lake)`. ## Volatility clustering Volatility clusters: large moves beget large moves. The lake measures the realized-vol ladder (`sp_rv1/5/22`), its ratios (`sp_vol_ratio_1_22`, `sp_vol_ratio_5_22`), HAR-RV ratios and RV lag-1 autocorrelation, and a GARCH(1,1) MLE (`sp_garch_*`: conditional variance, standardized residual, persistence α+β). Clustering is the most consistently predictive family in this campaign: the clean-lake compact reference set is built around RV/vol-ratio features `(PROVEN → exp 24)`, and the earliest tree splits of the reference model are dominated by realized-vol features `(HYPOTHESIS → chat-ideas.md: single-feature study; feature importance is from the pre-reset tree)`. But vol clustering is not the same as *vol-regime modeling*. The GARCH(1,1) trio, added to the reference, was refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13) `(PROVEN → exp 31)`. The pattern repeated: the generic realized-vol ladder contributes; the parametric vol model does not (ch. 07). This is the book's recurring lesson — **generic, scale-free, well-behaved statistics beat model-specific machinery on a small daily panel.** ## Regime A regime is a latent state of the return process — a two-state Gaussian HMM is fit on returns, and `sp_hmm_p_regime1` / `sp_hmm_state` carry the posterior. Regime flags failed as model features twice (exp 9 pre-reset idea, exp 25 clean-lake confirmation that model-specific families regress the signal) `(PROVEN → exp 25; the exp 9 idea is pre-reset idea material)`. The surviving hypothesis is that regime belongs **overlay, not feature**: a long-only/regime-gate that holds names only in the favourable state `(HYPOTHESIS → chat-ideas.md; untested on the clean lake)`. What would settle it: a gated version of the exp-26 n_drop=1 book compared against the ungated book over the same window `TODO(evidence-needed: HMM regime gate as overlay on exp-26 book)`. ## Mean reversion Mean reversion is the flip side of drift: short-horizon reversal. The lake measures it as `sp_ou_zscore` (distance from a fitted OU/AR(1) mean), `sp_ou_half_life` (mean-reversion speed), and via Hurst < 0.5. The single-feature study found `sp_ou_zscore` the strongest stable standalone predictor (IC −0.15/−0.13, sign-stable across years) `(HYPOTHESIS → chat-ideas.md: single-feature study; idea only)`. Yet adding it to the reference model regressed every metric — IC 0.0343 vs 0.0511, net −3.76% vs −3.21% `(PROVEN → exp 25)`. This is the book's named open problem, the "OU paradox": the strongest single-feature signal is worthless — worse, harmful — inside a cross-sectional rank model `(see chat-ideas.md, TODO(evidence-needed: why single-feature IC ≠ marginal contribution in CSRankNorm+LGBM))`. The working explanation (hypothesis): the OU z-score carries name-specific scale that survives CSRankNorm poorly and collides with the vol/trend families the model already uses. ## Memory and the path signature Memory is long-range dependence: a return's persistence beyond the short horizon. The lake measures Hurst exponent (R/S) and the path signature (lead/lag integrals of log-price path, levels 1–2 at lag 1 and 5). The dataset study found mild persistence across the panel (H ≈ 0.54–0.63), i.e. neither strong trend nor strong mean-reversion at the measured lags `(HYPOTHESIS → chat-ideas.md: martingale study; idea only)`. Signatures are the *generic* memory feature and they are load-bearing on the clean lake: `sp_sig_level1/2` are part of the compact reference set `(PROVEN → exp 24)`. Memory's practical meaning in this book: persistence is weak and horizon-dependent, so the signal must be refreshed on a short label and turned over carefully — which is exactly the cost argument of ch. 03 and ch. 09. ## Risk Risk here is the denominator of every edge: realized vol ratios (`sp_vol_ratio_5_22`), drawdown, and the noise floor of the rank measurement itself. Two facts about risk matter throughout the book. First, the measurement floor: on a 50-name cross-section the daily RankIC null std is 1/√(N−1) ≈ 0.143, so a mean RankIC near 0.06 is a small-but-real edge sitting on a wide null — every performance claim in this book is read against that floor `(method, PROVEN via the exp-21 detection playbook; also REFERENCED for the rank-null statistic)`. Second, the binding constraint: on clean data the gross→net collapse is ~9–10pp of cost drag, and IR ≈ 0.21 net is the campaign's best result `(PROVEN → exp 26)`. Risk limits — liquidity floor, size caps, drawdown pause — gate the live book, but the pre-reset evidence that the floor beats caps needs a post-reset rerun before it can be cited as fact `(HYPOTHESIS → exp 18 pre-clean-lake; TODO(evidence-needed: risk-limit A/B on the exp-26 reference))` (ch. 10). ## Error: the scorecard that separates hypothesis from proof Error is the book's discipline: how wrong was the prediction, in a way that can be measured against a null? The canonical metrics on the clean lake are IC, ICIR, Rank IC, Rank ICIR (the rank-based signal quality) and net IR, net return, L/S Sharpe, max drawdown (the portfolio outcome). Pre-reset runs (exp 8–18) recorded a different schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ir_with_cost`) — the two schemas are never compared directly in this book `(EVIDENCE.md: metric-schema note)`. Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); live results are reconciled with slippage and cost, not taken from the backtest (ch. 11). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26, round 3)`. ## Probability and significance Every statistic in this book carries a significance discipline: t-stats on pooled drift regressions (t≥2 for the submartingale reads), null-baseline z-scores for per-day IC (3–4σ single-day ICs were the contamination fingerprint that exposed the dirty lake), and sign-agreement across symbols/years `(method PROVEN via the exp-21 detection playbook; magnitude claims from the martingale study are HYPOTHESIS)`. The point is procedural: a metric without a null hypothesis is a number, not evidence. The book's probability posture is that ~50-name panels give weak statistical power, so a single improved run is a hypothesis until reproduced — exp 30 (M2) is explicitly labeled HYPOTHESIS for exactly this reason `(PROVEN run, HYPOTHESIS claim → exp 30)`. ## Timeline, decay, reinforcement The last cluster is about time. The lake's label is 5-day forward return — the 5d horizon is the campaign's best IC lever `(HYPOTHESIS → chat-ideas.md; the 5d label choice predates the clean lake)`, and horizon matters: 5d sees reversal that a 1d label blurs and a 63/126d label can't distinguish from drift. Decay shows up three ways: signal decay (a 5d reversal signal is stale after its horizon — this is why turnover relief, not signal engineering, was the biggest lever), feature decay (momentum/short-slope features weakened from 2025 to 2026 in the single-feature study `(HYPOTHESIS → chat-ideas.md)`), and cost decay (every holding day the cost drag compounds against a thin edge — the n_drop 2→1 result is the book's cleanest example: identical IC/RankIC, net flips from −3.21% to +2.13% purely from holding the dropped name `(PROVEN → exp 26)`). Reinforcement is the positive half of decay: ensemble averaging. The 5-seed blend raises net performance on clean data, and seed count is load-bearing — 2 seeds lose to 5 (RankIC 0.0579 vs 0.0663, net −1.49% vs +2.13%) `(PROVEN → exp 28)`. Averaging is reinforcement against noise, not against cost; ch. 05 works out the mechanism. ## From statistic to hypothesis to proved practice The cycle this book runs on: **measure → hypothesize → pre-register → isolate → prove or refute → reconcile live.** 1. **Measure** — a lake study turns a metric into a number (this chapter's vocabulary; dataset studies in `book/data/`). 2. **Hypothesize** — the number becomes a falsifiable claim ("adding risk-adjusted drift helps"), recorded in the run notes before the run. 3. **Pre-register** — the claim and its acceptance metric (IC/RankIC above reference, net IR above reference) are fixed before execution, to block post-hoc cherry-picking across the 31+ experiments. 4. **Isolate** — one variable changes per run; the reference book and its metrics are the control (exp 26 → 29/30/31). 5. **Prove or refute** — on the clean lake only. A reproduced improvement becomes PROVEN; a single un-reproduced run stays HYPOTHESIS (exp 30); a degradation is REFUTED and — critically — is recorded as a win for the discipline (exp 29, exp 31 stopped wrong directions). 6. **Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 11). Every metric in this chapter sits on this loop. The drift metric produced the momentum hypothesis and the reversal hypothesis; only one survived isolation. The error metrics are the loop's judge. The decay and risk metrics are why ch. 03 and ch. 09 exist at all. The rest of the book is the working-out of this cycle, claim by claim, with each claim traceable to `EVIDENCE.md` and a recorded run. Open questions for the desk (see `README.md`): why single-feature OU IC does not survive inside the model; whether 5d reversal trades net of cost; whether the HMM regime overlay beats the ungated book; whether M2 reproduction holds; and whether the 50-ETF panel generalizes.