Files
book-tac/book/chapters/02-research-loop.md
T

4.3 KiB
Raw Blame History

Chapter 02 — The Research Loop: Lake → Experiment → Live

Status: drafting. Claim inventory: see README.md ch. 02.

The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. This chapter is the machinery that makes them trustworthy: the traced research loop that turns a hypothesis into a proved practice. It is the book's methodology chapter, and it is also the book's proof-of-work — every later chapter is a walkthrough of this loop on a concrete question.

The loop

  1. Lake — bars and features live in one hive-partitioned lake (market/timeframe/symbol, family=ta|sp), with coverage and calendar metadata. It is the single source of bar/feature truth. The clean-lake rebuild proved the stakes: when the lake was rebuilt, the reference signal collapsed (IC 0.0354 → 0.0019) because the old lake's data quality had silently inflated results (PROVEN → exp 21). A claim built on the lake is only as good as the lake.
  2. Experiment — every run is a traced experiment: a git branch (exp/N-…), an MLflow run with recorded config/params/metrics, and hypothesis/evaluation notes recorded before and after the run. The traceability loop was used on exp 8–31; the branch, run, and notes are the reproducible unit (PROVEN → the traced experiment store; see rd_exp_*tools andEVIDENCE.md).
  3. Live — a proved book advances to a round window (targets → intents → decisions → orders → fills), and is reconciled (slippage bps, cost, funnel) (PROVEN → round 3; tac-rd-book trail). Live beats backtest: a claim about trading performance must trace to a round, not to a backtest (ch. 11).

Why traceability is the methodology

TradeAC ran 31+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:

  • The dirty lake. Pre-reset exp 8–18 reported strong results that did not survive a clean rebuild (PROVEN → exp 21). Only because the exact YAML, branch, and run were recorded could the desk reproduce — and falsify — the reference. Traceability is what turned a false belief into evidence.
  • Post-hoc cherry-picking. With 31+ experiments, the best-looking number is expected to be inflated by selection. The counter is pre-registration: hypothesis, change, and acceptance metric are fixed in the run notes before the run (REFERENCED — research practice; see CLAIMS.md multiple-testing note). Where the book quotes an experiment whose hypothesis was recorded after the fact, it says so.

The pre-reset experiments (exp 8–18) are therefore treated as idea material, not fact: they were demonstrably inflated by lake data quality (EVIDENCE.md pre-clean-lake section). Their ideas feed the hypothesis pipeline (ch. 01); their numbers never stand alone.

The isolation discipline

A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio. Because each changed one thing against the same reference, each verdict is attributable (PROVEN → exp 28–31). Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.

Falsification is the output

Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14). The campaign's refuted runs were as valuable as its wins (PROVEN → refuted runs recorded; REFERENCED for the falsification principle). This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.

From here

Ch. 03 applies the loop to the book's first worked question (does the signal clear costs?), ch. 06 to the clean-lake reset, and ch. 11 to the live round that closes the loop with reconciliation.

Open questions: purge/walk-forward CV instead of single train/valid split TODO(evidence-needed: purged CV on the exp-26 reference), and a PSI-based drift-aware retraining gate (HYPOTHESIS → chat-ideas.md).