book: scaffold + ch00 (execution trail as spine) — evidence exp 8-31, round 3
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
# CLAIMS.md — Proven vs Hypothesis Matrix
|
||||
|
||||
The running scoreboard of every quantitative claim in the book. Updated per chapter after HITL review. Status codes: `PROVEN` (reproduced from recorded run / reconciled round), `HYPOTHESIS` (plausible, tested once or never), `REFUTED` (tested and contradicted), `REFERENCED` (external citation).
|
||||
|
||||
## Signal & features
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC 0.040, ICIR 0.062) | PROVEN | EVIDENCE#001 → exp 8 |
|
||||
| Costs erase most of the baseline edge (+6.2% gross → +1.6% net) | PROVEN | EVIDENCE#002 → exp 8 |
|
||||
| Dropping model-specific feature families (ou, hmm) improves rank signal (RankIC 0.030→0.064) | PROVEN | EVIDENCE#003 → exp 9 |
|
||||
| Adding moment/volatility families regresses the signal | PROVEN (refuted direction) | EVIDENCE#004 → exp 11 |
|
||||
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | PROVEN (refuted direction) | EVIDENCE#014 → exp 25 |
|
||||
| Multi-horizon momentum (M1) degrades the reference | PROVEN (refuted direction) | EVIDENCE#017 → exp 29 |
|
||||
| GARCH(1,1) vol-regime features add no signal | PROVEN (refuted direction) | EVIDENCE#019 → exp 31 |
|
||||
| Risk-adjusted 22d Sharpe drift (M2) improves portfolio metrics | HYPOTHESIS (one run, unreproduced) | EVIDENCE#018 → exp 30 |
|
||||
| More features ≠ better signal on a small (50-name) cross-section | HYPOTHESIS (3 supporting runs, panel-specific) | EVIDENCE#003/004/014/017/019 |
|
||||
| General stochastic features (no TA/HMM/OU) have highest ICIR 0.340 | PROVEN | EVIDENCE#012 → exp 23 |
|
||||
|
||||
## Model
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| 5-seed RankIC ensemble raises performance vs single model on ablated set | PROVEN (pre-reset); re-validated post-reset exp 22–24 | EVIDENCE#005/011/013 |
|
||||
| Seed count is load-bearing: 2 seeds < 5 seeds on clean data | PROVEN | EVIDENCE#016 → exp 28 |
|
||||
| n_drop 2→1 flips net excess (−3.21% → +2.13%) with identical signal metrics | PROVEN | EVIDENCE#015 → exp 26 |
|
||||
| Cost drag is the binding constraint, not signal quality | PROVEN | EVIDENCE#015 → exp 26 (IC/RankIC identical across n_drop) |
|
||||
| Fractional-Kelly sizing beats equal-weight top-k net of costs | HYPOTHESIS (exp 15 never finished) | run never completed |
|
||||
|
||||
## Portfolio construction & risk
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | PROVEN | EVIDENCE#006/007 → exp 13/14 |
|
||||
| Stop-control churns and bleeds costs (−11.3pp cost drag) | PROVEN | EVIDENCE#006 → exp 13 |
|
||||
| $5M liquidity floor improves net IR (0.81→0.98) and cuts drawdown (7.9%→5.4%) | PROVEN (pre-clean-lake; not comparable post-reset) | EVIDENCE#008 → exp 18 |
|
||||
| Size/concentration caps hurt by cutting deployed capital | PROVEN (pre-clean-lake) | EVIDENCE#008 → exp 18 |
|
||||
| Entry/risk gates (momentum, HMM) are byte-identical no-ops on the reference signal | PROVEN | EVIDENCE#009 → exp 20 |
|
||||
| Signal quality is the bottleneck, not the execution/risk layer | PROVEN (on the exp-20 reference) | EVIDENCE#009 → exp 20 |
|
||||
|
||||
## Data & reproducibility
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | PROVEN | EVIDENCE#010 → exp 21 |
|
||||
| Old-lake data quality inflated the signal and backtest | PROVEN | EVIDENCE#010 → exp 21 |
|
||||
| Signal work must be re-validated after any data rebuild | PROVEN (exp 21) / HYPOTHESIS (generality) | EVIDENCE#010 |
|
||||
| Pre-reset experiment baselines are not comparable to post-reset runs | PROVEN | EVIDENCE#009/010 (exp 20 R0 note, exp 21) |
|
||||
|
||||
## Live execution
|
||||
|
||||
| Claim | Status | Evidence |
|
||||
|-------|--------|----------|
|
||||
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled | PROVEN | EVIDENCE#020 → round 3 |
|
||||
| Realized slippage ≈ 4.54 bps, est. cost ≈ $45, turnover 0.74 | PROVEN | EVIDENCE#020 → round 3 metrics |
|
||||
| Execution claims trace to round_id + reconcile, not backtest | PROVEN (methodology, round 3 settled) | EVIDENCE#020 |
|
||||
| 50-ETF panel results generalize to other universes | HYPOTHESIS — TODO(evidence-needed) | — |
|
||||
|
||||
## Open questions (settled by further experiments)
|
||||
|
||||
- exp 30 M2 Sharpe-drift: reproduce on a second window before promoting past HYPOTHESIS.
|
||||
- exp 15 Kelly sizing: re-run on the clean lake.
|
||||
- exp 18 risk-limit spec: re-validate $5M liquidity floor on the post-reset reference signal (exp 26 lineage).
|
||||
- Out-of-universe validation: non-ETF universe for the compact stochastic feature set.
|
||||
@@ -0,0 +1,55 @@
|
||||
# Evidence Ledger
|
||||
|
||||
Every quantitative claim in the book lands here: id → claim → source (experiment/run/branch, round_id, script, citation) → verified?.
|
||||
|
||||
## Key metric-schema note
|
||||
|
||||
Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ann_with_cost`, `excess_ir_with_cost`, `ls_ann_return`). Experiments 21+ use the canonical `IC / ICIR / Rank IC / Rank ICIR / net_IR / net_ann_return / gross_* / Long-Short_Ann_Sharpe / net_max_drawdown`. Do not compare schemas directly; chapter text states which schema a number comes from. Additionally, exp 20's R0 note states the exp-18 baseline is not comparable to post-reset runs due to environment non-determinism, and exp 21 invalidated all pre-clean-lake positive results.
|
||||
|
||||
## Pre-clean-lake period (exp 8–18) — historical, superseded
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| EVIDENCE#001 | Baseline 1-day LGB signal weak on 2026 OOS: IC 0.017, ICIR 0.062, RankIC 0.040, RankICIR 0.161 (below 0.2 noise threshold). L/S ann +4.9%. | exp 8, run `e65cf1ec…` (mlflow exp 10), branch `exp/8-baseline-lightgbm-on-the-full-60etf-univ` | yes |
|
||||
| EVIDENCE#002 | Costs erase most of the raw edge on baseline: excess +6.2% ann w/o cost (IR 0.31, MaxDD −20.4%) vs +1.6% ann after costs (IR 0.08). | exp 8 (same run) | yes |
|
||||
| EVIDENCE#003 | Feature-family ablation: generic-only (jump,har,trend,hurst,signature,ret,max_move) beats all-24: RankIC 0.030→0.064, RankICIR 0.146→0.276, L/S Sharpe −0.83→+2.55, net excess −9.4%→+3.1%. | exp 9, run `7b1e7972…` (mlflow exp 11), branch `exp/9-sp5d-feature-family-ablation` | yes |
|
||||
| EVIDENCE#004 | Adding 16 moment/volatility fields regresses every metric (RankIC 0.064→0.047, net excess −16.2% IR −1.57) — same failure mode as ou/hmm. | exp 11, run `a3f7d1d4…` (mlflow exp 12), branch `exp/11-sp5d-momentfeature-extension-after-exten` | yes |
|
||||
| EVIDENCE#005 | 5-seed RankIC ensemble on ablated generic features: RankIC 0.0586, RankICIR 0.224, net excess +7.8% (IR 0.79), L/S Sharpe 3.71, MDD −7.9%. Best pre-clean-lake net result. | exp 12, run `0cea66d9…` (mlflow exp 16), branch `exp/12-isolate-the-multiseed-rankic-ensemble-ef` | yes (superseded by EVIDENCE#011 on clean data) |
|
||||
| EVIDENCE#006 | OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) worse than TopkDropout: net excess −2.7% (IR −0.31) vs +7.8%; cost drag −11.3pp. | exp 13, run `4e1f77b4…` (mlflow exp 17), branch `exp/13-portfolioconstruction-variant-of-the-iso` | yes |
|
||||
| EVIDENCE#007 | OptimalStopControlV2 (turnover band/cooldown/cap) also refuted: net −6.9% (IR −0.72) vs TopkDropout +7.8% (IR 0.79). | exp 14, run `83d7e27e…` (mlflow exp 18), branch `exp/14-enhanced-stochasticcontrol-allocation-fo` | yes |
|
||||
| EVIDENCE#008 | Risk-limit A/B: $5M liquidity floor → net IR 0.81→0.98, cumDD 7.93%→5.44%; size cap 15% + conc 60% hurts (IR 0.816, ann 6.11%). | exp 18, run `28c7fa08…` (mlflow exp 21), branch `exp/18-risk-limit-control-on-the-reference-ense` | yes (pre-clean-lake, see note) |
|
||||
| EVIDENCE#009 | Improvement sweep (R1-R5): 4/5 refuted; R2 momentum gate and R3 HMM gate are byte-identical no-ops; R5 MA3/EWMA marginal (IR 0.049). Conclusion: signal quality is the bottleneck, not the execution/risk layer. | exp 20, run `958198a8…` (mlflow exp 21), branch `exp/20-improve-the-risk-limit-reference-signal` | yes |
|
||||
|
||||
## Post-reset period (exp 21–31) — canonical, current
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| EVIDENCE#010 | Clean-lake re-execution of the reference collapsed: IC 0.0019 (vs ref 0.0354), RankIC 0.0259, net −20.6% (IR −2.70). Old lake data quality had inflated the signal. | exp 21, run `f1bd3c28…` (mlflow exp 23), branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` | yes |
|
||||
| EVIDENCE#011 | Re-run after fixing feature routing: IC 0.0486, RankIC 0.0617, ICIR 0.235, RankICIR 0.243, L/S Sharpe 3.23. | exp 22, run `18db5bc1…` (mlflow exp 24), branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` | yes |
|
||||
| EVIDENCE#012 | General stochastic features only (no TA/HMM/OU): IC 0.0728, ICIR 0.340, L/S Sharpe 4.56. | exp 23, run `be5cd314…` (mlflow exp 25), branch `exp/23-test-whether-the-5-day-rankic-ensemble-i` | yes |
|
||||
| EVIDENCE#013 | Compact stochastic set (raw OHLCV + sp_ret, jump, RV1/5/22, vol ratios, trend slopes, logp, hurst, signature L1/L2): IC 0.0511, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54. | exp 24, run `fe469a19…` (mlflow exp 25), branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` | yes |
|
||||
| EVIDENCE#014 | Adding sp_ou_zscore hurts on clean data: IC 0.0343 vs 0.0511, net −3.76% vs −3.21%. | exp 25, run `57450d1a…` (mlflow exp 25), branch `exp/25-test-the-clean-data-hypothesis-that-addi` | yes |
|
||||
| EVIDENCE#015 | n_drop 2→1 on identical compact stochastic signal: gross +7.02%, net +2.13% (vs −3.21%), MDD −7.69%, IR 0.21. IC/RankIC identical to n_drop 2 — the gain is turnover/cost relief. | exp 26, run `21afc6af…` (mlflow exp 25), branch `exp/26-test-whether-reducing-topkdropout-daily` | yes — best result of the campaign |
|
||||
| EVIDENCE#016 | 2-seed ensemble loses to 5-seed on clean data: RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21). Seed count is load-bearing. | exp 28, run `c4ab1d01…` (mlflow exp 27), branch `exp/28-isolate-the-seed-count-effect-on-the-ndr` | yes |
|
||||
| EVIDENCE#017 | Multi-horizon momentum bundle refuted: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13%. | exp 29, run `b4586675…` (mlflow exp 28), branch `exp/29-isolation-run-m1-does-adding-multi-horiz` | yes |
|
||||
| EVIDENCE#018 | Risk-adjusted 22d Sharpe drift: mixed — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio strong (net +6.53% IR 0.62 vs +2.13% IR 0.21). Single run, unreproduced. | exp 30, run `d5d775f9…` (mlflow exp 29), branch `exp/30-isolation-run-m2-does-adding-risk-adjust` | yes — mark HYPOTHESIS in text |
|
||||
| EVIDENCE#019 | GARCH(1,1) vol-regime trio refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13). | exp 31, run `514cb523…` (mlflow exp 30), branch `exp/31-isolation-run-m3-does-adding-garch11-vol` | yes |
|
||||
|
||||
## Live execution trail
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| EVIDENCE#020 | Live round 3 (target 2026-08-17): retrained exp-26 n_drop=1 config on rolling 4y window; Topk10/n_drop1 with risk limits (liq floor $5M dropped 8, size cap 12%, conc 95%, drawdown pause 10%); funnel 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped (SLV delta_zero); invested $74,202.85, slippage 4.54 bps, est. cost ~$45. | round 3 (`tac-rd-book`), trace 27, run `721ef257…` (mlflow exp 26), branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac` | yes — settled, reconcile available |
|
||||
| EVIDENCE#021 | Scheduled retrain on 2026-08-14 (pre-reset reference): 10 buys + 6 sells placed, 0 cancelled by sentiment gate; sized on live equity $99,999.93. | trace 16, run `3b858b2b…` (mlflow exp 13), branch `exp/16-scheduled-algo-retrain-on-20260814-tacrd` | yes — historical, pre-reset signal |
|
||||
|
||||
## Ad-hoc scripts (book/data/)
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| (none yet) | — | — | — |
|
||||
|
||||
## External references (book/references/)
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| (none yet) | — | — | — |
|
||||
+142
@@ -0,0 +1,142 @@
|
||||
# TradeAC Quant Trading Guide — Table of Contents & Status
|
||||
|
||||
A practitioner's guide to quantitative trading written the only way it is worth reading: grounded in a real research loop and a real execution trail. Every number in this book was either reproduced from a recorded TradeAC experiment (MLflow run + traced git branch) or a reconciled live round, or it is explicitly labeled a hypothesis. See `AGENTS.md` (repo root) for the truth contract; `EVIDENCE.md` for the ledger; `CLAIMS.md` for the proven-vs-hypothesis matrix.
|
||||
|
||||
## What this book is for
|
||||
|
||||
A quant-desk reader should be able to act on this book: replicate a signal pipeline, size a book, gate it with risk limits, execute it, and reconcile what actually happened. The book's spine is **how performance improved with research-proved truth** — the actual arc of TradeAC's campaign from a baseline that barely cleared costs to a live, reconciled round.
|
||||
|
||||
## How to read evidence tags
|
||||
|
||||
- `PROVEN` — reproduced from a recorded run or reconciled round. Citation is an `experiment_id`/`run_id` or `round_id`.
|
||||
- `HYPOTHESIS` — plausible but not yet reproduced; never stated as fact.
|
||||
- `REFERENCED` — industry/academic practice; citation is an external source.
|
||||
- `TODO(evidence-needed: …)` — an open question the desk should settle.
|
||||
|
||||
## Table of contents
|
||||
|
||||
| # | Chapter | Status | Core experiments cited | Core lesson |
|
||||
|---|---------|--------|------------------------|-------------|
|
||||
| 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile |
|
||||
| 01 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
|
||||
| 02 | Baseline and the cost reality | drafting | exp 8 | A signal that dies after 5bp/15bp is not a signal |
|
||||
| 03 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific |
|
||||
| 04 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing |
|
||||
| 05 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
|
||||
| 06 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value |
|
||||
| 07 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge |
|
||||
| 08 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge |
|
||||
| 09 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck |
|
||||
| 10 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
|
||||
| 11 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why |
|
||||
|
||||
Status legend: `drafting` → `in-review` → `done`.
|
||||
|
||||
## Per-chapter claim inventory (expected truth status)
|
||||
|
||||
Each chapter opens with its claims. The inventory below is the working contract: what the chapter asserts, and what evidence tier it must land in. It is updated as chapters pass their HITL review gate.
|
||||
|
||||
### 00 — Why a real execution trail matters
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| A book's claims must be reconcilable to a real trail (targets→decisions→fills) | `PROVEN` — round 3 funnel |
|
||||
| Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 |
|
||||
| The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema |
|
||||
|
||||
### 01 — The research loop
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 |
|
||||
| Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) |
|
||||
| The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk |
|
||||
|
||||
### 02 — Baseline and the cost reality
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 |
|
||||
| Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 |
|
||||
| A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` |
|
||||
|
||||
### 03 — Prune, don't add
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
|
||||
| Adding moment/volatility families regresses the signal (exp 11), same failure mode as ou/hmm | `PROVEN` — exp 11 |
|
||||
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
|
||||
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3 runs, still panel-specific) |
|
||||
|
||||
### 04 — Ensembles
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
|
||||
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
|
||||
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
|
||||
|
||||
### 05 — Clean-lake reset
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 |
|
||||
| Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction |
|
||||
| Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality |
|
||||
|
||||
### 06 — Isolation runs
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 design |
|
||||
| Multi-horizon momentum degrades the reference (net IR 0.21→-1.12) | `PROVEN` — exp 29 |
|
||||
| Risk-adjusted 22d Sharpe drift is promising on portfolio metrics, mixed on rank | `HYPOTHESIS` — exp 30 single run, unreproduced |
|
||||
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
|
||||
|
||||
### 07 — Portfolio construction
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
|
||||
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
|
||||
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
|
||||
|
||||
### 08 — Cost/turnover frontier
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
|
||||
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
|
||||
|
||||
### 09 — Risk limits
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
|
||||
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18 |
|
||||
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
|
||||
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
|
||||
|
||||
### 10 — Live execution and reconciliation
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
|
||||
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
|
||||
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
|
||||
|
||||
### 11 — Synthesis
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26 |
|
||||
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
|
||||
| Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) |
|
||||
|
||||
## Repository layout
|
||||
|
||||
```
|
||||
book/
|
||||
README.md # this file
|
||||
EVIDENCE.md # ledger: id → claim → source → verified?
|
||||
CLAIMS.md # proven-vs-hypothesis matrix, updated every chapter
|
||||
chapters/00-intro.md ... # one file per chapter
|
||||
data/ # ad-hoc validation scripts + outputs
|
||||
references/ # external citations
|
||||
```
|
||||
|
||||
## Open questions for the desk
|
||||
|
||||
- `TODO(evidence-needed: reproduction of exp 30 M2 Sharpe-drift run on a second window)`
|
||||
- `TODO(evidence-needed: exp 15 Kelly sizing — run never finished; re-run on the clean lake)`
|
||||
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic feature set)`
|
||||
- `TODO(evidence-needed: reconciliation of exp 18 risk-limit spec on the post-reset reference signal)`
|
||||
@@ -0,0 +1,56 @@
|
||||
# Chapter 00 — Why a Real Execution Trail Matters
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 00.
|
||||
|
||||
Most quant books are written backwards: the author knows the answer, then builds a narrative to fit it. Backtests are quoted as if they were the outcome, the fill price is assumed to be the signal price, and cost is a footnote. This book is written the other way: every claim that could survive contact with a trading desk must survive contact with a *trail* — a record of what was intended, what was decided, what was placed, and what actually filled, at what price.
|
||||
|
||||
This chapter sets the spine: the `tac-rd-book` execution trail, which records every live round end-to-end.
|
||||
|
||||
## The funnel is the minimal honesty structure
|
||||
|
||||
A live trading round on the TradeAC stack is a chain of five checkpoints:
|
||||
|
||||
```
|
||||
targets (intent) → decided → placed (order) → filled → reconciled
|
||||
```
|
||||
|
||||
The trail records each step as first-class evidence. A round is only settled when the funnel has been reconciled — targets versus decisions versus fills, with per-symbol residuals and roll-ups for cash/buying-power impact, slippage in basis points, and cost as a fraction of gross traded notional. `PROVEN` — this is the schema of `tac-rd-book` (`trail_query`, `book_reconcile`), the same tooling used for the live rounds this book cites.
|
||||
|
||||
Why this structure and not a spreadsheet of P&L? Because P&L is the *last* place problems show up. By the time net return is wrong, you no longer know whether the intent was wrong (bad signal), the decision was wrong (bad gating), the fill was wrong (bad execution), or the book was wrong (bad risk). A funnel isolates the four.
|
||||
|
||||
## What the trail proved that a backtest could not
|
||||
|
||||
The book's live ground truth is round 3 (target date 2026-08-17), the first fully reconciled round of the post-reset signal: `EVIDENCE#020 → round 3`.
|
||||
|
||||
- The funnel held under live conditions: **10 targets → 10 decided → 10 placed → 9 filled**. One order was cancelled, and one target (SLV) was skipped because its requested delta was zero.
|
||||
- Realized slippage was **4.54 bps**; estimated cost **≈ $45** on $74,202.85 invested; turnover **0.74**.
|
||||
- The intent included risk limits from the research campaign: a $5M liquidity floor that dropped 8 of the 50 names, a 12% size cap, 95% concentration cap, and a 10% drawdown pause. `EVIDENCE#020`.
|
||||
|
||||
None of these numbers — slippage in bps, cost as a fraction of gross, the ratio of filled to placed — exists in a backtest. A backtest assumes a cost model (on this stack, 5 bp open / 15 bp close / $5 minimum) and a fill at the close price. The trail records what the market actually charged. That is the difference between a research claim and a trading claim.
|
||||
|
||||
`TODO(evidence-needed: the round-3 reconcile's realized-cost-vs-model comparison once the position window closes)`
|
||||
|
||||
## Backtests are historical, not promises
|
||||
|
||||
Throughout this book, backtest metrics carry a warning label, not a hiding place: universe, date window, and whether the hypothesis was pre-registered before the run. This matters because TradeAC ran 31+ experiments; with that many draws, some positive results will be luck. The book is explicit about which runs were pre-registered (e.g. isolation runs exp 28–31) and which were exploratory. `REFERENCED` — multiple-testing/cherry-picking risk is standard research practice; see `references/` as it accrues.
|
||||
|
||||
The most important proof of this discipline is the clean-lake reset, which this book treats as a turning point rather than a footnote: the pre-reset reference signal did **not** reproduce on a rebuilt lake (`EVIDENCE#010 → exp 21`). Had the book quoted the pre-reset backtest as fact, it would have shipped a lie. The trail and the traceability loop are what allowed the desk to catch it. Chapter 05 tells that story in full.
|
||||
|
||||
## How to read this book
|
||||
|
||||
- Every claim is tagged `PROVEN` (traced experiment/round), `HYPOTHESIS` (unreproduced), or `REFERENCED` (external source). `EVIDENCE.md` maps each tag to the run, branch, and round behind it.
|
||||
- Chapters 02–09 follow the research arc: what was tested, what was proved, what was refuted, and what moved performance. Refuted runs are cited as evidence too — knowing what *doesn't* work is how the desk avoided paying for it twice.
|
||||
- Chapter 10 is the reality check: live execution against the research claims.
|
||||
- Chapter 11 is the synthesis: the scoreboard of what actually improved performance and why.
|
||||
|
||||
## Open questions
|
||||
|
||||
- `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
|
||||
- `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)`
|
||||
|
||||
## Evidence cited in this chapter
|
||||
|
||||
| Tag | Source |
|
||||
|-----|--------|
|
||||
| `EVIDENCE#020` | round 3, `tac-rd-book`, trace 27 (branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac`) |
|
||||
| `EVIDENCE#010` | exp 21, run `f1bd3c28…`, branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` |
|
||||
Reference in New Issue
Block a user