Files
book-tac/book/README.md
T

142 lines
9.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TradeAC Quant Trading Guide — Table of Contents & Status
A practitioner's guide to quantitative trading written the only way it is worth reading: grounded in a real research loop and a real execution trail. Every number in this book was either reproduced from a recorded TradeAC experiment (MLflow run + traced git branch) or a reconciled live round, or it is explicitly labeled a hypothesis. See `AGENTS.md` (repo root) for the truth contract; `EVIDENCE.md` for the ledger; `CLAIMS.md` for the proven-vs-hypothesis matrix.
## What this book is for
A quant-desk reader should be able to act on this book: replicate a signal pipeline, size a book, gate it with risk limits, execute it, and reconcile what actually happened. The book's spine is **how performance improved with research-proved truth** — the actual arc of TradeAC's campaign from a baseline that barely cleared costs to a live, reconciled round.
## How to read evidence tags
- `PROVEN` — reproduced from a recorded run or reconciled round. Citation is an `experiment_id`/`run_id` or `round_id`.
- `HYPOTHESIS` — plausible but not yet reproduced; never stated as fact.
- `REFERENCED` — industry/academic practice; citation is an external source.
- `TODO(evidence-needed: …)` — an open question the desk should settle.
## Table of contents
| # | Chapter | Status | Core experiments cited | Core lesson |
|---|---------|--------|------------------------|-------------|
| 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile |
| 01 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
| 02 | Baseline and the cost reality | drafting | exp 8 | A signal that dies after 5bp/15bp is not a signal |
| 03 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific |
| 04 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing |
| 05 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
| 06 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value |
| 07 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge |
| 08 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge |
| 09 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck |
| 10 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
| 11 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why |
Status legend: `drafting` → `in-review` → `done`.
## Per-chapter claim inventory (expected truth status)
Each chapter opens with its claims. The inventory below is the working contract: what the chapter asserts, and what evidence tier it must land in. It is updated as chapters pass their HITL review gate.
### 00 — Why a real execution trail matters
| Claim | Expected status |
|-------|-----------------|
| A book's claims must be reconcilable to a real trail (targets→decisions→fills) | `PROVEN` — round 3 funnel |
| Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 |
| The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema |
### 01 — The research loop
| Claim | Expected status |
|-------|-----------------|
| Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 |
| Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) |
| The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk |
### 02 — Baseline and the cost reality
| Claim | Expected status |
|-------|-----------------|
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 |
| Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 |
| A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` |
### 03 — Prune, don't add
| Claim | Expected status |
|-------|-----------------|
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
| Adding moment/volatility families regresses the signal (exp 11), same failure mode as ou/hmm | `PROVEN` — exp 11 |
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3 runs, still panel-specific) |
### 04 — Ensembles
| Claim | Expected status |
|-------|-----------------|
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
### 05 — Clean-lake reset
| Claim | Expected status |
|-------|-----------------|
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 |
| Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction |
| Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality |
### 06 — Isolation runs
| Claim | Expected status |
|-------|-----------------|
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 design |
| Multi-horizon momentum degrades the reference (net IR 0.21→-1.12) | `PROVEN` — exp 29 |
| Risk-adjusted 22d Sharpe drift is promising on portfolio metrics, mixed on rank | `HYPOTHESIS` — exp 30 single run, unreproduced |
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
### 07 — Portfolio construction
| Claim | Expected status |
|-------|-----------------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
### 08 — Cost/turnover frontier
| Claim | Expected status |
|-------|-----------------|
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
### 09 — Risk limits
| Claim | Expected status |
|-------|-----------------|
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18 |
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
### 10 — Live execution and reconciliation
| Claim | Expected status |
|-------|-----------------|
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
### 11 — Synthesis
| Claim | Expected status |
|-------|-----------------|
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26 |
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
| Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) |
## Repository layout
```
book/
README.md # this file
EVIDENCE.md # ledger: id → claim → source → verified?
CLAIMS.md # proven-vs-hypothesis matrix, updated every chapter
chapters/00-intro.md ... # one file per chapter
data/ # ad-hoc validation scripts + outputs
references/ # external citations
```
## Open questions for the desk
- `TODO(evidence-needed: reproduction of exp 30 M2 Sharpe-drift run on a second window)`
- `TODO(evidence-needed: exp 15 Kelly sizing — run never finished; re-run on the clean lake)`
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic feature set)`
- `TODO(evidence-needed: reconciliation of exp 18 risk-limit spec on the post-reset reference signal)`