Files
book-tac/book/chapters/06-clean-lake-reset.md
T

62 lines
5.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 06 — The Clean-Lake Reset: Data Quality as First-Order Risk
Status: drafting. Claim inventory: see `README.md` ch. 06.
Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on.
## The result
The reference was the 5-seed RankIC ensemble on the 50-ETF panel, trained on the old lake. The clean-lake re-execution ran the exact same YAML — same universe, features, model, windows, strategy, costs.
| Metric | Pre-reset reference | Clean-lake re-execution |
|--------|---------------------|-------------------------|
| IC | 0.0354 | 0.0019 |
| ICIR | 0.150 | 0.0115 |
| Rank IC | 0.0586 | 0.0259 |
| Rank ICIR | 0.224 | 0.143 |
| Net-of-cost excess vs SPY (ann) | +7.77% | −20.6% |
| Net IR | +0.79 | −2.70 |
| Max drawdown | −7.9% | −15.2% |
`PROVEN — EVIDENCE#010 → exp 21, run f1bd3c28…, branch exp/21-clean-lake-re-execution-of-the-tac-rd-ra`. Same config, opposite sign. There is no softer way to say it: the pre-reset campaign's headline result was inflated by the lake's data-quality problems and may not be cited as fact anywhere in this book.
## Why the signal moved so much
The failures were in the feature layer, not the bars and not the labels. In the pre-reset investigation the team documented, and the clean-lake rebuild confirmed, a family of silent failure modes:
1. **Provider-path mismatch.** The feature reader pointed at a path that did not exist under the lake's `family=ta|sp` partitioning; features loaded as NaN and `DropAllNaN` silently removed them, so workflows trained on OHLCV only — without knowing it.
2. **Silent column-dropping in feature regeneration.** A regeneration omitted the `har` family, dropping `sp_rv1/5/22` and `sp_vol_ratio_1_22/5_22` from 71 of 72 parquet files; a 25-feature model silently became a 20-feature model.
3. **Schema fragmentation.** The 72 feature files carried 4 different column schemas (24/53/58/66 columns), so "the same feature set" was not actually the same feature set across the lake.
4. **Stale coverage / truncated feature range.** Feature files covered only a trailing ~30-day window while bars spanned 2016–2026 (SPY: 2669 bar rows, 20 feature rows).
5. **Mid-experiment regeneration.** Feature files were rewritten between the reference run and a later run, so two runs nominally sharing a config trained on different feature files.
`HYPOTHESIS (chat-documented failure classes; book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt — idea material, not evidence)`. The post-reset reproduction of the *detection* is what is `PROVEN`: exp 22 re-ran after the feature-routing fix and the signal reappeared (IC 0.0486), establishing that the routing bug — not the model, not the data-generating process — had been suppressing features (`EVIDENCE#011 → exp 22`).
## The detection playbook
What allowed the team to catch this, in order of power:
1. **Byte-identical reproduction.** Keep configs frozen; a same-config collapse isolates data as the cause.
2. **Prediction-distribution comparison.** Compare pred scale, rank correlation, and top-k overlap across runs of the same config.
3. **Null-baseline calibration.** Compare mean daily RankIC to the null std of `1/√(N−1)`; a signal only a fraction of a sigma above null is not evidence of edge.
4. **Per-day IC outlier fingerprint.** Single-day ICs of 3–4σ on a 50-name correlated panel are the signature of contamination, not insight.
5. **Feature-vs-bar alignment and coverage checks.** Bars, labels and features must cover the same window and rows; columns must not silently vanish.
6. **Same-environment baselines.** An environment reset or code change corrupts cross-run comparison; establish a fresh same-env baseline before judging any overlay.
`HYPOTHESIS (detection methods, chat-documented and later institutionalized as the lake validation gate: validate_lake_dataset — see book/references/chat-ideas.md)`. The one piece that is directly `PROVEN` from clean data: after the rebuild and routing fix, signal and backtest both reappeared at economically meaningful magnitudes (exp 22–24, `EVIDENCE#011/012/013`), which is the positive control that the reset worked.
## What this means for the rest of the book
- **Only exp 21+ is evidence.** All chapters in this book cite the clean-lake lineage (exp 21–31) and post-reset live rounds. Pre-reset runs and chat transcripts are hypotheses and ideas, clearly labeled.
- **Reproducibility is a research activity, not a chore.** The traceability loop — per-experiment git branch, MLflow run, pre-registered hypothesis, recorded evaluation — is what made the collapse *detectable* rather than embarrassing.
- **A "fix" is not proven by one run.** The route from exp 21 (collapse) to exp 22 (fix) to exp 23/24 (independent re-validations) is the pattern: reproduce, isolate, reproduce again.
- **`TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`** — the hollow-coverage and schema-drift classes recurred; the desk's validation gate exists but is not yet a mandatory pre-run step.
## Evidence cited in this chapter
| Tag | Source |
|-----|--------|
| `EVIDENCE#010` | exp 21, run `f1bd3c28…`, branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` |
| `EVIDENCE#011` | exp 22, run `18db5bc1…`, branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` |
| `EVIDENCE#012/013` | exp 23/24, runs `be5cd314…` / `fe469a19…` |
| chat mining | book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt (idea only) |