# Chapter 05 — The Clean-Lake Reset: Data Quality as First-Order Risk Status: drafting. Claim inventory: see `README.md` ch. 05. Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on. ## The result The reference was the 5-seed RankIC ensemble on the 50-ETF panel, trained on the old lake. The clean-lake re-execution ran the exact same YAML — same universe, features, model, windows, strategy, costs. | Metric | Pre-reset reference | Clean-lake re-execution | |--------|---------------------|-------------------------| | IC | 0.0354 | 0.0019 | | ICIR | 0.150 | 0.0115 | | Rank IC | 0.0586 | 0.0259 | | Rank ICIR | 0.224 | 0.143 | | Net-of-cost excess vs SPY (ann) | +7.77% | −20.6% | | Net IR | +0.79 | −2.70 | | Max drawdown | −7.9% | −15.2% | `PROVEN — EVIDENCE#010 → exp 21, run f1bd3c28…, branch exp/21-clean-lake-re-execution-of-the-tac-rd-ra`. Same config, opposite sign. There is no softer way to say it: the pre-reset campaign's headline result was inflated by the lake's data-quality problems and may not be cited as fact anywhere in this book. ## Why the signal moved so much The failures were in the feature layer, not the bars and not the labels. In the pre-reset investigation the team documented, and the clean-lake rebuild confirmed, a family of silent failure modes: 1. **Provider-path mismatch.** The feature reader pointed at a path that did not exist under the lake's `family=ta|sp` partitioning; features loaded as NaN and `DropAllNaN` silently removed them, so workflows trained on OHLCV only — without knowing it. 2. **Silent column-dropping in feature regeneration.** A regeneration omitted the `har` family, dropping `sp_rv1/5/22` and `sp_vol_ratio_1_22/5_22` from 71 of 72 parquet files; a 25-feature model silently became a 20-feature model. 3. **Schema fragmentation.** The 72 feature files carried 4 different column schemas (24/53/58/66 columns), so "the same feature set" was not actually the same feature set across the lake. 4. **Stale coverage / truncated feature range.** Feature files covered only a trailing ~30-day window while bars spanned 2016–2026 (SPY: 2669 bar rows, 20 feature rows). 5. **Mid-experiment regeneration.** Feature files were rewritten between the reference run and a later run, so two runs nominally sharing a config trained on different feature files. `HYPOTHESIS (chat-documented failure classes; book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt — idea material, not evidence)`. The post-reset reproduction of the *detection* is what is `PROVEN`: exp 22 re-ran after the feature-routing fix and the signal reappeared (IC 0.0486), establishing that the routing bug — not the model, not the data-generating process — had been suppressing features (`EVIDENCE#011 → exp 22`). ## The detection playbook What allowed the team to catch this, in order of power: 1. **Byte-identical reproduction.** Keep configs frozen; a same-config collapse isolates data as the cause. 2. **Prediction-distribution comparison.** Compare pred scale, rank correlation, and top-k overlap across runs of the same config. 3. **Null-baseline calibration.** Compare mean daily RankIC to the null std of `1/√(N−1)`; a signal only a fraction of a sigma above null is not evidence of edge. 4. **Per-day IC outlier fingerprint.** Single-day ICs of 3–4σ on a 50-name correlated panel are the signature of contamination, not insight. 5. **Feature-vs-bar alignment and coverage checks.** Bars, labels and features must cover the same window and rows; columns must not silently vanish. 6. **Same-environment baselines.** An environment reset or code change corrupts cross-run comparison; establish a fresh same-env baseline before judging any overlay. `HYPOTHESIS (detection methods, chat-documented and later institutionalized as the lake validation gate: validate_lake_dataset — see book/references/chat-ideas.md)`. The one piece that is directly `PROVEN` from clean data: after the rebuild and routing fix, signal and backtest both reappeared at economically meaningful magnitudes (exp 22–24, `EVIDENCE#011/012/013`), which is the positive control that the reset worked. ## What this means for the rest of the book - **Only exp 21+ is evidence.** All chapters in this book cite the clean-lake lineage (exp 21–31) and post-reset live rounds. Pre-reset runs and chat transcripts are hypotheses and ideas, clearly labeled. - **Reproducibility is a research activity, not a chore.** The traceability loop — per-experiment git branch, MLflow run, pre-registered hypothesis, recorded evaluation — is what made the collapse *detectable* rather than embarrassing. - **A "fix" is not proven by one run.** The route from exp 21 (collapse) to exp 22 (fix) to exp 23/24 (independent re-validations) is the pattern: reproduce, isolate, reproduce again. - **`TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`** — the hollow-coverage and schema-drift classes recurred; the desk's validation gate exists but is not yet a mandatory pre-run step. ## Evidence cited in this chapter | Tag | Source | |-----|--------| | `EVIDENCE#010` | exp 21, run `f1bd3c28…`, branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` | | `EVIDENCE#011` | exp 22, run `18db5bc1…`, branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` | | `EVIDENCE#012/013` | exp 23/24, runs `be5cd314…` / `fe469a19…` | | chat mining | book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt (idea only) |