Files
tac-exp-dev/book/README.md
T
zhaoli 06fb1e8ee9 book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2026-08-20 01:23:21 +00:00

170 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TradeAC Quant Trading Guide — Table of Contents & Status
A practitioner's guide to quantitative trading written the only way it is worth reading: grounded in a real research loop and a real execution trail. Every number in this book was either reproduced from a recorded TradeAC experiment (MLflow run + traced git branch) **on the clean lake (exp 21+)** or a reconciled post-reset live round, or it is explicitly labeled a hypothesis. See `AGENTS.md` (repo root) for the truth contract; `EVIDENCE.md` for the ledger; `CLAIMS.md` for the proven-vs-hypothesis matrix.
## Evidence boundary and living-document status
- **The clean-lake boundary (2026-08-18, exp 21) is the evidence watermark.** Anything before it — exp 8–18 and their backtests, pre-reset live rounds — is historical context and idea material only, never cited as fact (they were demonstrably inflated by lake data-quality problems, `EVIDENCE#010 → exp 21`). Pre-reset experiments and all opencode chat transcripts (see `data/chat_mining/` and `references/chat-ideas.md`) feed the book's hypothesis pipeline.
- **Every section is living.** As new experimental results land on the clean lake, chapters are updated; a chapter marked `done` is done for its window, not forever.
## What this book is for
A quant-desk reader should be able to act on this book: replicate a signal pipeline, size a book, gate it with risk limits, execute it, and reconcile what actually happened. The book's spine is **how performance improved with research-proved truth** — the actual arc of TradeAC's campaign from a baseline that barely cleared costs to a live, reconciled round.
## How to read evidence tags
- `PROVEN` — reproduced from a recorded run or reconciled round. Citation is an `experiment_id`/`run_id` or `round_id`.
- `HYPOTHESIS` — plausible but not yet reproduced; never stated as fact.
- `REFERENCED` — industry/academic practice; citation is an external source.
- `TODO(evidence-needed: …)` — an open question the desk should settle.
## Table of contents
| # | Chapter | Status | Core experiments cited | Core lesson |
|---|---------|--------|------------------------|-------------|
| 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile |
| 01 | Metrics: the vocabulary of a price series | drafting | exp 21–31 + dataset studies | Every claim reduces to a falsifiable statistic |
| 02 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
| 03 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal |
| 04 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25, 43 | On a 50-name panel, generic beats model-specific; standalone reversal never existed |
| 05 | Ensembles and the seed-count effect | drafting | exp 12, 28, 34 | Averaging raises ICIR; more seeds raise breadth but not net — cost is the ceiling |
| 06 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
| 07 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31, 33–37, 43 | Most additions fail; the discipline is the value |
| 08 | Portfolio construction: dropout vs the rest | drafting | exp 13, 14, 15, 35, 38, 39, 41 | Turnover-sensitive construction bleeds the edge; weekly recompute wins |
| 09 | The cost/turnover frontier | drafting | exp 26, 39, 41 | Cut turnover before adding signal; weekly rebalance is the proven lever |
| 10 | Risk limits and gates that work | drafting | exp 18, 20, 40, 42 | Limits are a safety net, not alpha; gates churn without signal |
| 11 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
| 12 | Synthesis: how proved truth compounds | drafting | all, exp 33–43 | Cost relief > signal; the Q-campaign scoreboard |
Status legend: `drafting` → `in-review` → `done`.
## Per-chapter claim inventory (expected truth status)
Each chapter opens with its claims. The inventory below is the working contract: what the chapter asserts, and what evidence tier it must land in. It is updated as chapters pass their HITL review gate.
### 00 — Why a real execution trail matters
| Claim | Expected status |
|-------|-----------------|
| A book's claims must be reconcilable to a real trail (targets→decisions→fills) | `PROVEN` — round 3 funnel |
| Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 |
| The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema |
### 01 — Metrics: the vocabulary of a price series
| Claim | Expected status |
|-------|-----------------|
| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–12) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) |
| Generic scale-free statistics beat model-specific machinery on a small daily panel | `PROVEN` (exp 23/24/25/29/31) + `HYPOTHESIS` (generality) |
| A statistic is only as good as the falsification it survives (null z-scores, reproduction) | `PROVEN` (exp 21 detection playbook) + `REFERENCED` |
| The strongest single-feature signal (OU z-score) can be worthless inside a rank model — the "OU paradox" | `PROVEN` (exp 25) + open mechanism `TODO(evidence-needed)` |
### 02 — The research loop
| Claim | Expected status |
|-------|-----------------|
| Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 |
| Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) |
| The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk |
| One variable changes per run (isolation); verdicts attributable | `PROVEN` — exp 26→28/29/30/31 design |
### 03 — Baseline and the cost reality
| Claim | Expected status |
|-------|-----------------|
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 |
| Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 |
| A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` |
### 04 — Prune, don't add
| Claim | Expected status |
|-------|-----------------|
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
| Adding moment/volatility families regresses the signal (exp 11), same failure mode as ou/hmm | `PROVEN` — exp 11 |
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
| Standalone 5d reversal (single feature sp_trend_slope_5) is not learnable — model trains positive IC | `PROVEN` — exp 43 (Q11) |
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3+ runs, still panel-specific) |
### 05 — Ensembles
| Claim | Expected status |
|-------|-----------------|
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
| 10 seeds raise rank breadth (RankIC 0.0671, L/S Sharpe 4.58) but the book stays negative net | `PROVEN` — exp 34 (Q02) |
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
### 06 — Clean-lake reset
| Claim | Expected status |
|-------|-----------------|
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 |
| Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction |
| Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality |
### 07 — Isolation runs
| Claim | Expected status |
|-------|-----------------|
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 + exp 33–43 Q-runs design |
| Multi-horizon momentum degrades the reference (net IR 0.21→-1.12) | `PROVEN` — exp 29 |
| Risk-adjusted 22d Sharpe drift (M2): reproduced on the compact set by Q01 | `PROVEN` — exp 30 + exp 33 (Q01) |
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
| Longer labels raise IC monotonically but net worsens under daily turnover (10d/22d) | `PROVEN` — exp 36/37 (Q04/Q05) |
| Standalone reversal feature does not reproduce | `PROVEN` — exp 43 (Q11) |
### 08 — Portfolio construction
| Claim | Expected status |
|-------|-----------------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
| Weekly rebalance recompute of the daily signal is the campaign's best construction (net +12.51%, IR 1.24) | `PROVEN` — exp 39 (Q07) |
| Fractional-Kelly sizing (exp 15) is refuted on the clean lake (net +1.04%, IR 0.11) | `PROVEN` — exp 38 (Q06) |
| Widening the book (topk 20) adds no edge; long-short top/bottom is destroyed by turnover | `PROVEN` — exp 35/41 (Q03/Q09) |
### 09 — Cost/turnover frontier
| Claim | Expected status |
|-------|-----------------|
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
| Weekly recompute cuts cost drag to ~1.1pp and unlocks +12.51% net | `PROVEN` — exp 39 (Q07) |
| Long-short daily turnover costs 9.7% of NAV ($96.7k); fill rate 0.40 | `PROVEN` — exp 41 (Q09) |
### 10 — Risk limits
| Claim | Expected status |
|-------|-----------------|
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18; re-confirmed clean-lake exp 40 (Q08) |
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
| Post-reset A/B: the floor binds but adds no IR edge; DD relief is pure defunding | `PROVEN` — exp 40 (Q08) |
| HMM regime gate meets only the drawdown leg and churns | `PROVEN` — exp 42 (Q10) |
### 11 — Live execution and reconciliation
| Claim | Expected status |
|-------|-----------------|
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
### 12 — Synthesis
| Claim | Expected status |
|-------|-----------------|
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26, 39 |
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31, Q02–Q06, Q09–Q11) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
| Turnover reduction is the dominant net-performance lever (weekly rebalance +12.51% vs daily −3.21%–+2.13%) | `PROVEN` — exp 26 vs 39 |
| Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) |
## Repository layout
```
book/
README.md # this file
EVIDENCE.md # ledger: id → claim → source → verified?
CLAIMS.md # proven-vs-hypothesis matrix, updated every chapter
chapters/00-intro.md ... # one file per chapter
data/ # ad-hoc validation scripts + outputs
data/chat_mining/ # raw opencode chat transcripts (idea sources)
references/chat-ideas.md # distilled ideas/hypotheses from chats + pre-reset experiments
references/ # external citations
```
## Open questions for the desk
- `TODO(evidence-needed: weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round)`
- `TODO(evidence-needed: long-horizon label (10d/22d) paired with a low-turnover construction — the signal edge is proven, the cost kills daily churn)`
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic feature set)`
- `TODO(evidence-needed: exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha)`