# Chapter 02 — The Research Loop: Lake → Experiment → Live Status: drafting. Claim inventory: see `README.md` ch. 02. The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. This chapter is the machinery that makes them trustworthy: the traced research loop that turns a hypothesis into a proved practice. It is the book's methodology chapter, and it is also the book's proof-of-work — every later chapter is a walkthrough of this loop on a concrete question. ## The loop 1. **Lake** — bars and features live in one hive-partitioned lake (market/timeframe/symbol, `family=ta|sp`), with coverage and calendar metadata. It is the single source of bar/feature truth. The clean-lake rebuild proved the stakes: when the lake was rebuilt, the reference signal collapsed (IC 0.0354 → 0.0019) because the old lake's data quality had silently inflated results `(PROVEN → exp 21)`. A claim built on the lake is only as good as the lake. 2. **Experiment** — every run is a traced experiment: a git branch (`exp/N-…`), an MLflow run with recorded config/params/metrics, and hypothesis/evaluation notes recorded before and after the run. The traceability loop was used on exp 8–31; the branch, run, and notes are the reproducible unit `(PROVEN → the traced experiment store; see `rd_exp_*` tools and `EVIDENCE.md`)`. 3. **Live** — a proved book advances to a round window (targets → intents → decisions → orders → fills), and is reconciled (slippage bps, cost, funnel) `(PROVEN → round 3; tac-rd-book trail)`. Live beats backtest: a claim about trading performance must trace to a round, not to a backtest (ch. 12). ## Why traceability is the methodology TradeAC ran 40+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case: - **The dirty lake.** Pre-reset exp 8–18 reported strong results that did not survive a clean rebuild `(PROVEN → exp 21)`. Only because the exact YAML, branch, and run were recorded could the desk reproduce — and falsify — the reference. Traceability is what turned a false belief into evidence. - **Post-hoc cherry-picking.** With 31+ experiments, the best-looking number is expected to be inflated by selection. The counter is pre-registration: hypothesis, change, and acceptance metric are fixed in the run notes *before* the run `(REFERENCED — research practice; see CLAIMS.md multiple-testing note)`. Where the book quotes an experiment whose hypothesis was recorded after the fact, it says so. The pre-reset experiments (exp 8–18) are therefore treated as **idea material, not fact**: they were demonstrably inflated by lake data quality `(EVIDENCE.md pre-clean-lake section)`. Their ideas feed the hypothesis pipeline (ch. 01); their numbers never stand alone. ## The isolation discipline A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio; and the Q-campaign (exp 33–43) changed exactly one thing per run against that same reference — features (Q01, Q11), seeds (Q02), topk (Q03), label horizon (Q04/Q05), sizing (Q06), rebalance cadence (Q07), risk limits (Q08), construction (Q09), and an entry gate (Q10). Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31; EVIDENCE#022–032 → exp 33–43)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS. ## Falsification is the output Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14), and nine of the eleven Q-campaign runs (longer labels, wider books, Kelly sizing, long-short, regime gates, standalone reversal — exp 34–38, 40–43). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat. ## From here Ch. 03 applies the loop to the book's first worked question (does the signal clear costs?), ch. 06 to the clean-lake reset, ch. 11 to walk-forward re-validation of the campaign's headline results, and ch. 12 to the live round that closes the loop with reconciliation. Open questions: purge/walk-forward CV instead of single train/valid split `TODO(evidence-needed: purged CV on the exp-26 reference)`, and a PSI-based drift-aware retraining gate `(HYPOTHESIS → chat-ideas.md)`.