book: evidence boundary (clean-lake watermark) + ch02 cost reality + ch05 clean-lake reset — exp 21-31, chat mining
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
# Chapter 02 — Baseline and the Cost Reality
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 02.
|
||||
|
||||
This chapter answers the question every quant desk must answer before the first dollar is deployed: **what does the raw signal have to be worth, and what survives the cost of trading it?**
|
||||
|
||||
The honest answer on the TradeAC stack, measured on the clean lake, is that the signal itself was modest — and the cost of expressing it was nearly its entire gross value. The order of magnitude is the lesson.
|
||||
|
||||
## The noise floor first
|
||||
|
||||
Before quoting a single IC, establish what noise looks like. For a cross-section of `N` independent names, daily RankIC under the null has standard deviation roughly `1/√(N−1)`. On the 50-ETF panel that is ≈ 0.143 per day. A signal whose daily RankIC mean is a small fraction of that standard deviation is statistically indistinguishable from noise day-to-day, however it may look averaged.
|
||||
|
||||
The post-reset clean-lake reference signal (exp 24, the compact stochastic set) reports mean RankIC ≈ 0.066 and RankICIR ≈ 0.255 — positive and above the informal 0.2 RankICIR "noise threshold" used on this desk, but far from overwhelming: `PROVEN — EVIDENCE#013 → exp 24`. `HYPOTHESIS (chat-derived null calibration: clean-lake mean RankIC of earlier runs sat ≈ 0.18–0.4σ of the null per-day distribution — book/data/chat_mining/exp-polluted-lake.txt)`. Treat the statistical significance of a 7-month, 50-name cross-section as fragile, not robust.
|
||||
|
||||
## The cost model that decides everything
|
||||
|
||||
The backtest and live sizing on this stack use a fixed cost model:
|
||||
|
||||
- open cost 0.0005 (5 bp), close cost 0.0015 (15 bp), minimum $5 per side;
|
||||
- fills assumed at the close (`deal_price = $close`), benchmark SPY, $1M starting account.
|
||||
|
||||
`PROVEN — strategy config of exp 21–31`. These are round-trip costs of ~20 bp, which is ordinary for liquid US ETFs at retail/PT sizes but not free. At ~20% of the book traded daily (topk=10, n_drop=2), the annualized cost drag is enormous relative to a signal worth single-digit annual excess.
|
||||
|
||||
## The gross → net collapse on clean data
|
||||
|
||||
The clean-lake sequence shows the pattern with the same signal, same costs, varying only the feature set and turnover:
|
||||
|
||||
| Run | Signal (IC / RankICIR) | Gross excess vs SPY | Net excess vs SPY | Net IR |
|
||||
|-----|------------------------|---------------------|-------------------|--------|
|
||||
| exp 22 (full TA+SP) | 0.0486 / 0.243 | +0.12% | −9.09% | −0.80 |
|
||||
| exp 23 (general sp only) | 0.0728 / 0.206 | +6.73% | −2.39% | −0.22 |
|
||||
| exp 24 (compact sp) | 0.0511 / 0.255 | +5.99% | −3.21% | −0.32 |
|
||||
| exp 26 (compact, n_drop=1) | 0.0511 / 0.255 | +7.02% | +2.13% | +0.21 |
|
||||
|
||||
`PROVEN — EVIDENCE#011/012/013/015 → exp 22/23/24/26`. Read the columns, not the rows: even the *best* clean-lake signal, at the default construction, lost roughly **nine to ten percentage points of annualized excess to costs** (exp 24: +5.99% gross → −3.21% net). The signal that produced a high long-short Sharpe (L/S ann Sharpe 4.54) could not survive daily rebalancing at 20 bp round trips.
|
||||
|
||||
This is the single most important number in the early book: **at this turnover, cost is not a haircut, it is the strategy's budget.** `PROVEN — EVIDENCE#015 → exp 26 (identical IC/RankIC across n_drop 2 and 1; the entire net difference is trading behavior, not signal)`. The pre-reset campaign observed the same shape historically (baseline +6.2% gross → +1.6% net), which is idea material, not evidence: `HYPOTHESIS (idea: pre-clean-lake, EVIDENCE#002 → exp 8)`.
|
||||
|
||||
## What fixed it, and what it implies
|
||||
|
||||
The only construction change that flipped net from negative to positive was reducing daily forced replacements from `n_drop=2` to `n_drop=1` — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`.
|
||||
|
||||
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 06) then confirmed that most candidate additions *reduced* the edge anyway.
|
||||
|
||||
## Desk rules distilled from this chapter
|
||||
|
||||
1. Establish the null noise floor before believing any IC/RankIC mean on a small cross-section.
|
||||
2. Report gross and net excess side by side, always with universe + window + cost model.
|
||||
3. Treat net-IR-of-signal as the bar for any construction change; signal metrics alone are not a strategy claim.
|
||||
4. When net is negative and gross is positive by ~10pp, attack turnover before features.
|
||||
5. `TODO(evidence-needed: realized-cost comparison of round 3 vs the 5bp/15bp/$5 model once the position window closes)`.
|
||||
|
||||
## Evidence cited in this chapter
|
||||
|
||||
| Tag | Source |
|
||||
|-----|--------|
|
||||
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
|
||||
| `EVIDENCE#011/012` | exp 22/23, runs `18db5bc1…` / `be5cd314…` |
|
||||
| `EVIDENCE#015` | exp 26, run `21afc6af…`, branch `exp/26-test-whether-reducing-topkdropout-daily` |
|
||||
| `EVIDENCE#002` | exp 8 (pre-clean-lake, idea only) |
|
||||
| chat mining | book/data/chat_mining/exp-polluted-lake.txt (null calibration, idea only) |
|
||||
@@ -0,0 +1,62 @@
|
||||
# Chapter 05 — The Clean-Lake Reset: Data Quality as First-Order Risk
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 05.
|
||||
|
||||
Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on.
|
||||
|
||||
## The result
|
||||
|
||||
The reference was the 5-seed RankIC ensemble on the 50-ETF panel, trained on the old lake. The clean-lake re-execution ran the exact same YAML — same universe, features, model, windows, strategy, costs.
|
||||
|
||||
| Metric | Pre-reset reference | Clean-lake re-execution |
|
||||
|--------|---------------------|-------------------------|
|
||||
| IC | 0.0354 | 0.0019 |
|
||||
| ICIR | 0.150 | 0.0115 |
|
||||
| Rank IC | 0.0586 | 0.0259 |
|
||||
| Rank ICIR | 0.224 | 0.143 |
|
||||
| Net-of-cost excess vs SPY (ann) | +7.77% | −20.6% |
|
||||
| Net IR | +0.79 | −2.70 |
|
||||
| Max drawdown | −7.9% | −15.2% |
|
||||
|
||||
`PROVEN — EVIDENCE#010 → exp 21, run f1bd3c28…, branch exp/21-clean-lake-re-execution-of-the-tac-rd-ra`. Same config, opposite sign. There is no softer way to say it: the pre-reset campaign's headline result was inflated by the lake's data-quality problems and may not be cited as fact anywhere in this book.
|
||||
|
||||
## Why the signal moved so much
|
||||
|
||||
The failures were in the feature layer, not the bars and not the labels. In the pre-reset investigation the team documented, and the clean-lake rebuild confirmed, a family of silent failure modes:
|
||||
|
||||
1. **Provider-path mismatch.** The feature reader pointed at a path that did not exist under the lake's `family=ta|sp` partitioning; features loaded as NaN and `DropAllNaN` silently removed them, so workflows trained on OHLCV only — without knowing it.
|
||||
2. **Silent column-dropping in feature regeneration.** A regeneration omitted the `har` family, dropping `sp_rv1/5/22` and `sp_vol_ratio_1_22/5_22` from 71 of 72 parquet files; a 25-feature model silently became a 20-feature model.
|
||||
3. **Schema fragmentation.** The 72 feature files carried 4 different column schemas (24/53/58/66 columns), so "the same feature set" was not actually the same feature set across the lake.
|
||||
4. **Stale coverage / truncated feature range.** Feature files covered only a trailing ~30-day window while bars spanned 2016–2026 (SPY: 2669 bar rows, 20 feature rows).
|
||||
5. **Mid-experiment regeneration.** Feature files were rewritten between the reference run and a later run, so two runs nominally sharing a config trained on different feature files.
|
||||
|
||||
`HYPOTHESIS (chat-documented failure classes; book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt — idea material, not evidence)`. The post-reset reproduction of the *detection* is what is `PROVEN`: exp 22 re-ran after the feature-routing fix and the signal reappeared (IC 0.0486), establishing that the routing bug — not the model, not the data-generating process — had been suppressing features (`EVIDENCE#011 → exp 22`).
|
||||
|
||||
## The detection playbook
|
||||
|
||||
What allowed the team to catch this, in order of power:
|
||||
|
||||
1. **Byte-identical reproduction.** Keep configs frozen; a same-config collapse isolates data as the cause.
|
||||
2. **Prediction-distribution comparison.** Compare pred scale, rank correlation, and top-k overlap across runs of the same config.
|
||||
3. **Null-baseline calibration.** Compare mean daily RankIC to the null std of `1/√(N−1)`; a signal only a fraction of a sigma above null is not evidence of edge.
|
||||
4. **Per-day IC outlier fingerprint.** Single-day ICs of 3–4σ on a 50-name correlated panel are the signature of contamination, not insight.
|
||||
5. **Feature-vs-bar alignment and coverage checks.** Bars, labels and features must cover the same window and rows; columns must not silently vanish.
|
||||
6. **Same-environment baselines.** An environment reset or code change corrupts cross-run comparison; establish a fresh same-env baseline before judging any overlay.
|
||||
|
||||
`HYPOTHESIS (detection methods, chat-documented and later institutionalized as the lake validation gate: validate_lake_dataset — see book/references/chat-ideas.md)`. The one piece that is directly `PROVEN` from clean data: after the rebuild and routing fix, signal and backtest both reappeared at economically meaningful magnitudes (exp 22–24, `EVIDENCE#011/012/013`), which is the positive control that the reset worked.
|
||||
|
||||
## What this means for the rest of the book
|
||||
|
||||
- **Only exp 21+ is evidence.** All chapters in this book cite the clean-lake lineage (exp 21–31) and post-reset live rounds. Pre-reset runs and chat transcripts are hypotheses and ideas, clearly labeled.
|
||||
- **Reproducibility is a research activity, not a chore.** The traceability loop — per-experiment git branch, MLflow run, pre-registered hypothesis, recorded evaluation — is what made the collapse *detectable* rather than embarrassing.
|
||||
- **A "fix" is not proven by one run.** The route from exp 21 (collapse) to exp 22 (fix) to exp 23/24 (independent re-validations) is the pattern: reproduce, isolate, reproduce again.
|
||||
- **`TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`** — the hollow-coverage and schema-drift classes recurred; the desk's validation gate exists but is not yet a mandatory pre-run step.
|
||||
|
||||
## Evidence cited in this chapter
|
||||
|
||||
| Tag | Source |
|
||||
|-----|--------|
|
||||
| `EVIDENCE#010` | exp 21, run `f1bd3c28…`, branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` |
|
||||
| `EVIDENCE#011` | exp 22, run `18db5bc1…`, branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` |
|
||||
| `EVIDENCE#012/013` | exp 23/24, runs `be5cd314…` / `fe469a19…` |
|
||||
| chat mining | book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt (idea only) |
|
||||
Reference in New Issue
Block a user