book: insert ch01 metrics vocabulary + ch02 research loop; renumber 03/06 — evidence exp 21-31, martingale study
This commit is contained in:
+32
-22
@@ -23,17 +23,18 @@ A quant-desk reader should be able to act on this book: replicate a signal pipel
|
||||
| # | Chapter | Status | Core experiments cited | Core lesson |
|
||||
|---|---------|--------|------------------------|-------------|
|
||||
| 00 | Why a real execution trail matters | drafting | round 3 | A book claims nothing it cannot reconcile |
|
||||
| 01 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
|
||||
| 02 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal |
|
||||
| 03 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific |
|
||||
| 04 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing |
|
||||
| 05 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
|
||||
| 06 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value |
|
||||
| 07 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge |
|
||||
| 08 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge |
|
||||
| 09 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck |
|
||||
| 10 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
|
||||
| 11 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why |
|
||||
| 01 | Metrics: the vocabulary of a price series | drafting | exp 21–31 + dataset studies | Every claim reduces to a falsifiable statistic |
|
||||
| 02 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
|
||||
| 03 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal |
|
||||
| 04 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific |
|
||||
| 05 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing |
|
||||
| 06 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
|
||||
| 07 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value |
|
||||
| 08 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge |
|
||||
| 09 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge |
|
||||
| 10 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck |
|
||||
| 11 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
|
||||
| 12 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why |
|
||||
|
||||
Status legend: `drafting` → `in-review` → `done`.
|
||||
|
||||
@@ -48,21 +49,30 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
| Backtest claims without live reconciliation are hypotheses about execution | `HYPOTHESIS` → settled by round 3 |
|
||||
| The funnel (targets→decided→placed→filled) is the minimal honesty structure | `REFERENCED` (industry ops practice) + `PROVEN` via tac-rd-book schema |
|
||||
|
||||
### 01 — The research loop
|
||||
### 01 — Metrics: the vocabulary of a price series
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–12) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) |
|
||||
| Generic scale-free statistics beat model-specific machinery on a small daily panel | `PROVEN` (exp 23/24/25/29/31) + `HYPOTHESIS` (generality) |
|
||||
| A statistic is only as good as the falsification it survives (null z-scores, reproduction) | `PROVEN` (exp 21 detection playbook) + `REFERENCED` |
|
||||
| The strongest single-feature signal (OU z-score) can be worthless inside a rank model — the "OU paradox" | `PROVEN` (exp 25) + open mechanism `TODO(evidence-needed)` |
|
||||
|
||||
### 02 — The research loop
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Experiments must be traced: branch + MLflow run + notes (hypothesis before run) | `PROVEN` — traceability loop used on exp 8–31 |
|
||||
| Pre-registration protects against post-hoc cherry-picking | `REFERENCED` (research practice; see CLAIMS for multiple-testing note) |
|
||||
| The lake is the single source of bar/feature truth | `PROVEN` — exp 21 showed dirty-lake risk |
|
||||
| One variable changes per run (isolation); verdicts attributable | `PROVEN` — exp 26→28/29/30/31 design |
|
||||
|
||||
### 02 — Baseline and the cost reality
|
||||
### 03 — Baseline and the cost reality
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Baseline 1-day LGB signal is weak on 2026 OOS (RankIC ≈ 0.04, below the 0.2 ICIR noise threshold) | `PROVEN` — exp 8 |
|
||||
| Costs erase most of the raw edge: +6.2% ann gross → +1.6% net | `PROVEN` — exp 8 |
|
||||
| A viable signal must clear realistic execution costs | `PROVEN` (exp 8, exp 26) + `REFERENCED` |
|
||||
|
||||
### 03 — Prune, don't add
|
||||
### 04 — Prune, don't add
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
|
||||
@@ -70,21 +80,21 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
|
||||
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3 runs, still panel-specific) |
|
||||
|
||||
### 04 — Ensembles
|
||||
### 05 — Ensembles
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
|
||||
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
|
||||
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
|
||||
|
||||
### 05 — Clean-lake reset
|
||||
### 06 — Clean-lake reset
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The reference signal did not reproduce on a rebuilt lake (IC 0.035→0.002) | `PROVEN` — exp 21 |
|
||||
| Data-quality problems had inflated earlier results; post-reset signal is the only valid one | `PROVEN` — exp 21 + exp 22–24 reproduction |
|
||||
| Signal work must be re-validated after any data rebuild | `PROVEN` (exp 21) + `HYPOTHESIS` for generality |
|
||||
|
||||
### 06 — Isolation runs
|
||||
### 07 — Isolation runs
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 design |
|
||||
@@ -92,20 +102,20 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
| Risk-adjusted 22d Sharpe drift is promising on portfolio metrics, mixed on rank | `HYPOTHESIS` — exp 30 single run, unreproduced |
|
||||
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
|
||||
|
||||
### 07 — Portfolio construction
|
||||
### 08 — Portfolio construction
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
|
||||
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
|
||||
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
|
||||
|
||||
### 08 — Cost/turnover frontier
|
||||
### 09 — Cost/turnover frontier
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
|
||||
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
|
||||
|
||||
### 09 — Risk limits
|
||||
### 10 — Risk limits
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
|
||||
@@ -113,14 +123,14 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
|
||||
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
|
||||
|
||||
### 10 — Live execution and reconciliation
|
||||
### 11 — Live execution and reconciliation
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
|
||||
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
|
||||
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
|
||||
|
||||
### 11 — Synthesis
|
||||
### 12 — Synthesis
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26 |
|
||||
|
||||
Reference in New Issue
Block a user