62 lines
5.9 KiB
Markdown
62 lines
5.9 KiB
Markdown
# Chapter 05 — The Clean-Lake Reset: Data Quality as First-Order Risk
|
||
|
||
Status: drafting. Claim inventory: see `README.md` ch. 05.
|
||
|
||
Every number quoted before this chapter was a warning shot. This chapter is the impact. On 2026-08-18 the TradeAC team rebuilt the data lake and re-executed its best reference experiment with byte-identical configuration. The signal collapsed. This is the most important methodological result in the book: **a positive backtest that does not reproduce on clean data was not a strategy, it was a data-quality artifact** — and the tools that caught it were the same traceability tools the book is built on.
|
||
|
||
## The result
|
||
|
||
The reference was the 5-seed RankIC ensemble on the 50-ETF panel, trained on the old lake. The clean-lake re-execution ran the exact same YAML — same universe, features, model, windows, strategy, costs.
|
||
|
||
| Metric | Pre-reset reference | Clean-lake re-execution |
|
||
|--------|---------------------|-------------------------|
|
||
| IC | 0.0354 | 0.0019 |
|
||
| ICIR | 0.150 | 0.0115 |
|
||
| Rank IC | 0.0586 | 0.0259 |
|
||
| Rank ICIR | 0.224 | 0.143 |
|
||
| Net-of-cost excess vs SPY (ann) | +7.77% | −20.6% |
|
||
| Net IR | +0.79 | −2.70 |
|
||
| Max drawdown | −7.9% | −15.2% |
|
||
|
||
`PROVEN — EVIDENCE#010 → exp 21, run f1bd3c28…, branch exp/21-clean-lake-re-execution-of-the-tac-rd-ra`. Same config, opposite sign. There is no softer way to say it: the pre-reset campaign's headline result was inflated by the lake's data-quality problems and may not be cited as fact anywhere in this book.
|
||
|
||
## Why the signal moved so much
|
||
|
||
The failures were in the feature layer, not the bars and not the labels. In the pre-reset investigation the team documented, and the clean-lake rebuild confirmed, a family of silent failure modes:
|
||
|
||
1. **Provider-path mismatch.** The feature reader pointed at a path that did not exist under the lake's `family=ta|sp` partitioning; features loaded as NaN and `DropAllNaN` silently removed them, so workflows trained on OHLCV only — without knowing it.
|
||
2. **Silent column-dropping in feature regeneration.** A regeneration omitted the `har` family, dropping `sp_rv1/5/22` and `sp_vol_ratio_1_22/5_22` from 71 of 72 parquet files; a 25-feature model silently became a 20-feature model.
|
||
3. **Schema fragmentation.** The 72 feature files carried 4 different column schemas (24/53/58/66 columns), so "the same feature set" was not actually the same feature set across the lake.
|
||
4. **Stale coverage / truncated feature range.** Feature files covered only a trailing ~30-day window while bars spanned 2016–2026 (SPY: 2669 bar rows, 20 feature rows).
|
||
5. **Mid-experiment regeneration.** Feature files were rewritten between the reference run and a later run, so two runs nominally sharing a config trained on different feature files.
|
||
|
||
`HYPOTHESIS (chat-documented failure classes; book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt — idea material, not evidence)`. The post-reset reproduction of the *detection* is what is `PROVEN`: exp 22 re-ran after the feature-routing fix and the signal reappeared (IC 0.0486), establishing that the routing bug — not the model, not the data-generating process — had been suppressing features (`EVIDENCE#011 → exp 22`).
|
||
|
||
## The detection playbook
|
||
|
||
What allowed the team to catch this, in order of power:
|
||
|
||
1. **Byte-identical reproduction.** Keep configs frozen; a same-config collapse isolates data as the cause.
|
||
2. **Prediction-distribution comparison.** Compare pred scale, rank correlation, and top-k overlap across runs of the same config.
|
||
3. **Null-baseline calibration.** Compare mean daily RankIC to the null std of `1/√(N−1)`; a signal only a fraction of a sigma above null is not evidence of edge.
|
||
4. **Per-day IC outlier fingerprint.** Single-day ICs of 3–4σ on a 50-name correlated panel are the signature of contamination, not insight.
|
||
5. **Feature-vs-bar alignment and coverage checks.** Bars, labels and features must cover the same window and rows; columns must not silently vanish.
|
||
6. **Same-environment baselines.** An environment reset or code change corrupts cross-run comparison; establish a fresh same-env baseline before judging any overlay.
|
||
|
||
`HYPOTHESIS (detection methods, chat-documented and later institutionalized as the lake validation gate: validate_lake_dataset — see book/references/chat-ideas.md)`. The one piece that is directly `PROVEN` from clean data: after the rebuild and routing fix, signal and backtest both reappeared at economically meaningful magnitudes (exp 22–24, `EVIDENCE#011/012/013`), which is the positive control that the reset worked.
|
||
|
||
## What this means for the rest of the book
|
||
|
||
- **Only exp 21+ is evidence.** All chapters in this book cite the clean-lake lineage (exp 21–31) and post-reset live rounds. Pre-reset runs and chat transcripts are hypotheses and ideas, clearly labeled.
|
||
- **Reproducibility is a research activity, not a chore.** The traceability loop — per-experiment git branch, MLflow run, pre-registered hypothesis, recorded evaluation — is what made the collapse *detectable* rather than embarrassing.
|
||
- **A "fix" is not proven by one run.** The route from exp 21 (collapse) to exp 22 (fix) to exp 23/24 (independent re-validations) is the pattern: reproduce, isolate, reproduce again.
|
||
- **`TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`** — the hollow-coverage and schema-drift classes recurred; the desk's validation gate exists but is not yet a mandatory pre-run step.
|
||
|
||
## Evidence cited in this chapter
|
||
|
||
| Tag | Source |
|
||
|-----|--------|
|
||
| `EVIDENCE#010` | exp 21, run `f1bd3c28…`, branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` |
|
||
| `EVIDENCE#011` | exp 22, run `18db5bc1…`, branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` |
|
||
| `EVIDENCE#012/013` | exp 23/24, runs `be5cd314…` / `fe469a19…` |
|
||
| chat mining | book/data/chat_mining/exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt (idea only) | |