Files
tac-exp-dev/book/README.md
T

192 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TradeAC Quant Trading Guide — Table of Contents & Status
A practitioner's guide to quantitative trading written the only way it is worth reading: grounded in a real research loop and a real execution trail. Every number in this book was either reproduced from a recorded TradeAC experiment (MLflow run + traced git branch) **on the clean lake (exp 21+)** or a reconciled post-reset live round, or it is explicitly labeled a hypothesis. See `AGENTS.md` (repo root) for the truth contract; `EVIDENCE.md` for the ledger; `CLAIMS.md` for the proven-vs-hypothesis matrix.
## Evidence boundary and living-document status
- **The clean-lake boundary (2026-08-18, exp 21) is the evidence watermark.** Anything before it — exp 8–18 and their backtests, pre-reset live rounds — is historical context and idea material only, never cited as fact (they were demonstrably inflated by lake data-quality problems, `EVIDENCE#010 → exp 21`). Pre-reset experiments and all opencode chat transcripts (see `data/chat_mining/` and `references/chat-ideas.md`) feed the book's hypothesis pipeline.
- **Every section is living.** As new experimental results land on the clean lake, chapters are updated; a chapter marked `done` is done for its window, not forever.
## What this book is for
A quant-desk reader should be able to act on this book: replicate a signal pipeline, size a book, gate it with risk limits, execute it, and reconcile what actually happened. The book's spine is **how performance improved with research-proved truth** — the actual arc of TradeAC's campaign from a baseline that barely cleared costs to a live, reconciled round.
## How to read evidence tags
- `PROVEN` — reproduced from a recorded run or reconciled round. Citation is an `experiment_id`/`run_id` or `round_id`.
- `HYPOTHESIS` — plausible but not yet reproduced; never stated as fact.
- `REFERENCED` — industry/academic practice; citation is an external source.
- `TODO(evidence-needed: …)` — an open question the desk should settle.
## Table of contents
| # | Chapter | Status | Core experiments cited | Core lesson |
|---|---------|--------|------------------------|-------------|
| 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile |
| 01 | Metrics: the vocabulary of a price series | drafting | exp 21–31 + dataset studies | Every claim reduces to a falsifiable statistic |
| 02 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
| 03 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal |
| 04 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25, 43 | On a 50-name panel, generic beats model-specific; standalone reversal never existed |
| 05 | Ensembles and the seed-count effect | drafting | exp 12, 28, 34 | Averaging raises ICIR; more seeds raise breadth but not net — cost is the ceiling |
| 06 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
| 07 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31, 33–37, 43 | Most additions fail; the discipline is the value |
| 08 | Portfolio construction: dropout vs the rest | drafting | exp 13, 14, 15, 35, 38, 39, 41 | Turnover-sensitive construction bleeds the edge; weekly recompute wins |
| 09 | The cost/turnover frontier | drafting | exp 26, 39, 41 | Cut turnover before adding signal; weekly rebalance is the proven lever |
| 10 | Risk limits and gates that work | drafting | exp 18, 20, 40, 42 | Limits are a safety net, not alpha; gates churn without signal |
| 11 | Walk-forward re-validation and guard candidates | drafting | exp 52–56 | The edge is a 2025–2026 regime artifact; all 5 guards refuted |
| 12 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
| 13 | Synthesis: how proved truth compounds | drafting | all, exp 33–43, 52–56 | Cost relief > signal; the Q-campaign scoreboard |
Status legend: `drafting` → `in-review` → `done`.
## Per-chapter claim inventory (expected truth status)
Each chapter opens with its claims. The inventory below is the working contract: what the chapter asserts, and what evidence tier it must land in. It is updated as chapters pass their HITL review gate.
### 00 — Why a real execution trail matters
| Claim | Expected status |
|-------|-----------------|
| A book's claims must be reconcilable to a real trail (targets→decisions→fills) | `PROVEN` — round 3 funnel |
| Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 |
| The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema |
### 01 — Metrics: the vocabulary of a price series
| Claim | Expected status |
|-------|-----------------|
| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–13) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) |
| Generic scale-free statistics beat model-specific machinery on a small daily panel | `PROVEN` (exp 23/24/25/29/31) + `HYPOTHESIS` (generality) |
| A statistic is only as good as the falsification it survives (null z-scores, reproduction) | `PROVEN` (exp 21 detection playbook) + `REFERENCED` |
| The strongest single-feature signal (OU z-score) can be worthless inside a rank model — the "OU paradox" | `PROVEN` (exp 25) + open mechanism `TODO(evidence-needed)` |
### 02 — The research loop
| Claim | Expected status |
|-------|-----------------|
| Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 |
| Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) |
| The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk |
| One variable changes per run (isolation); verdicts attributable | `PROVEN` — exp 26→28/29/30/31 design |
### 03 — Baseline and the cost reality
| Claim | Expected status |
|-------|-----------------|
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 |
| Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 |
| A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` |
### 04 — Prune, don't add
| Claim | Expected status |
|-------|-----------------|
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
| Adding moment/volatility families regresses the signal (exp 11), same failure mode as ou/hmm | `PROVEN` — exp 11 |
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
| Standalone 5d reversal (single feature sp_trend_slope_5) is not learnable — model trains positive IC | `PROVEN` — exp 43 (Q11) |
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3+ runs, still panel-specific) |
### 05 — Ensembles
| Claim | Expected status |
|-------|-----------------|
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
| 10 seeds raise rank breadth (RankIC 0.0671, L/S Sharpe 4.58) but the book stays negative net | `PROVEN` — exp 34 (Q02) |
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
### 06 — Clean-lake reset
| Claim | Expected status |
|-------|-----------------|
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 |
| Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction |
| Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality |
### 07 — Isolation runs
| Claim | Expected status |
|-------|-----------------|
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 + exp 33–43 Q-runs design |
| Multi-horizon momentum degrades the reference (net IR 0.21→-1.12) | `PROVEN` — exp 29 |
| Risk-adjusted 22d Sharpe drift (M2): reproduced on the compact set by Q01 | `PROVEN` — exp 30 + exp 33 (Q01) |
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
| Longer labels raise IC monotonically but net worsens under daily turnover (10d/22d) | `PROVEN` — exp 36/37 (Q04/Q05) |
| Standalone reversal feature does not reproduce | `PROVEN` — exp 43 (Q11) |
### 08 — Portfolio construction
| Claim | Expected status |
|-------|-----------------|
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
| Weekly rebalance recompute of the daily signal is the campaign's best construction (net +12.51%, IR 1.24) | `PROVEN` — exp 39 (Q07) |
| Fractional-Kelly sizing (exp 15) is refuted on the clean lake (net +1.04%, IR 0.11) | `PROVEN` — exp 38 (Q06) |
| Widening the book (topk 20) adds no edge; long-short top/bottom is destroyed by turnover | `PROVEN` — exp 35/41 (Q03/Q09) |
### 09 — Cost/turnover frontier
| Claim | Expected status |
|-------|-----------------|
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
| Weekly recompute cuts cost drag to ~1.1pp and unlocks +12.51% net | `PROVEN` — exp 39 (Q07) |
| Long-short daily turnover costs 9.7% of NAV ($96.7k); fill rate 0.40 | `PROVEN` — exp 41 (Q09) |
### 10 — Risk limits
| Claim | Expected status |
|-------|-----------------|
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18; re-confirmed clean-lake exp 40 (Q08) |
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
| Post-reset A/B: the floor binds but adds no IR edge; DD relief is pure defunding | `PROVEN` — exp 40 (Q08) |
| HMM regime gate meets only the drawdown leg and churns | `PROVEN` — exp 42 (Q10) |
### 11 — Walk-forward re-validation and guard candidates
| Claim | Expected status |
|-------|-----------------|
| The headline results (weekly +12.51%, m2-sharpe22 +6.5%) are 2026-window-specific; walk-forward re-training across 2024/2025 is negative or flat | `PROVEN` — exp 52/53 |
| A and C share identical predictions; the strategy layer alone decides the outcome | `PROVEN` — exp 52 |
| No pre-deployment measurable gate (feature-PSI, label-regime PSI, streaming IC, window length, staleness) selects a profitable year | `PROVEN` — exp 52–56, all 5 guards refuted |
| The edge is a 2025–2026 regime artifact; live capital must be cut until the regime returns | `PROVEN` (walk-forward) + `HYPOTHESIS` (forward-looking) |
### 12 — Live execution and reconciliation
| Claim | Expected status |
|-------|-----------------|
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
### 13 — Synthesis
| Claim | Expected status |
|-------|-----------------|
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26, 39 |
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31, Q02–Q06, Q09–Q11) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
| Turnover reduction is the dominant net-performance lever (weekly rebalance +12.51% vs daily −3.21%–+2.13%) | `PROVEN` — exp 26 vs 39 |
| The campaign's headline edges were a 2025–2026 regime artifact, not robust OOS | `PROVEN` — exp 52–56 |
| Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) |
## Repository layout
```
book/
README.md # this file
EVIDENCE.md # ledger: id → claim → source → verified?
CLAIMS.md # proven-vs-hypothesis matrix, updated every chapter
chapters/00-intro.md ... # one file per chapter
data/ # ad-hoc validation scripts + outputs
data/chat_mining/ # raw opencode chat transcripts (idea sources)
references/chat-ideas.md # distilled ideas/hypotheses from chats + pre-reset experiments
references/ # external citations
```
## Open questions for the desk
- `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
- `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)`
- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
- `TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`
- `TODO(evidence-needed: live round under weekly-rebalance construction with risk-limit spec, to confirm safety-net behavior at higher deployed capital)`
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
- `TODO(evidence-needed: a causal (no-lookahead) regime-change detector that selects the 2026 window before the fact — none of the five guards did)`
### Settled open questions (no longer active)
- ~~`weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`.
- ~~`long-horizon label (10d/22d) paired with a low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`.
- ~~`out-of-universe (non-ETF) validation of the compact stochastic feature set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`.
- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`.
- ~~`do the headline results survive walk-forward re-training?`~~ — **ANSWERED (negatively):** exp 52/53/54 re-ran weekly, moments, ndrop2, and m2-sharpe22 across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable; all prior years negative or flat. Edge = 2025–2026 regime artifact. `EVIDENCE#043–045`.
- ~~`is there a pre-deployment guard that isolates the profitable regime?`~~ — **ANSWERED (negatively):** feature-PSI, label-regime PSI, streaming IC (`ic_min_rankic`), adaptive short-window, and staleness guards all refuted. `EVIDENCE#043–047`.