book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches - CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims - README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign - new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated - Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
This commit is contained in:
+28
-17
@@ -26,15 +26,15 @@ A quant-desk reader should be able to act on this book: replicate a signal pipel
|
||||
| 01 | Metrics: the vocabulary of a price series | drafting | exp 21–31 + dataset studies | Every claim reduces to a falsifiable statistic |
|
||||
| 02 | The research loop: lake → experiment → live | drafting | exp 8–31 | Traceability is the methodology |
|
||||
| 03 | Baseline and the cost reality | drafting | exp 22–26, 28–31 | A signal that dies after 5bp/15bp is not a signal |
|
||||
| 04 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25 | On a 50-name panel, generic beats model-specific |
|
||||
| 05 | Ensembles and the seed-count effect | drafting | exp 12, 28 | Averaging raises ICIR; seed count is load-bearing |
|
||||
| 04 | Prune, don't add: feature-family ablation | drafting | exp 9, 10, 11, 25, 43 | On a 50-name panel, generic beats model-specific; standalone reversal never existed |
|
||||
| 05 | Ensembles and the seed-count effect | drafting | exp 12, 28, 34 | Averaging raises ICIR; more seeds raise breadth but not net — cost is the ceiling |
|
||||
| 06 | The clean-lake reset: data quality as first-order risk | drafting | exp 21–24 | If it doesn't reproduce on clean data, it was noise |
|
||||
| 07 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31 | Most additions fail; the discipline is the value |
|
||||
| 08 | Portfolio construction: dropout vs optimal stop | drafting | exp 13, 14, 15 | Turnover-sensitive construction bleeds the edge |
|
||||
| 09 | The cost/turnover frontier | drafting | exp 26 | n_drop 2→1: hold the dropped name, keep the edge |
|
||||
| 10 | Risk limits that work | drafting | exp 18, 20 | Liquidity floor > concentration caps; gates are no-ops when signal is the bottleneck |
|
||||
| 07 | Isolation runs: single-variable discipline | drafting | exp 26, 29–31, 33–37, 43 | Most additions fail; the discipline is the value |
|
||||
| 08 | Portfolio construction: dropout vs the rest | drafting | exp 13, 14, 15, 35, 38, 39, 41 | Turnover-sensitive construction bleeds the edge; weekly recompute wins |
|
||||
| 09 | The cost/turnover frontier | drafting | exp 26, 39, 41 | Cut turnover before adding signal; weekly rebalance is the proven lever |
|
||||
| 10 | Risk limits and gates that work | drafting | exp 18, 20, 40, 42 | Limits are a safety net, not alpha; gates churn without signal |
|
||||
| 11 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
|
||||
| 12 | Synthesis: how proved truth compounds | drafting | all | The scoreboard of what moved performance and why |
|
||||
| 12 | Synthesis: how proved truth compounds | drafting | all, exp 33–43 | Cost relief > signal; the Q-campaign scoreboard |
|
||||
|
||||
Status legend: `drafting` → `in-review` → `done`.
|
||||
|
||||
@@ -78,13 +78,15 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
| Dropping model-specific feature families (ou, hmm) improves the rank signal (RankIC 0.030→0.064) | `PROVEN` — exp 9 |
|
||||
| Adding moment/volatility families regresses the signal (exp 11), same failure mode as ou/hmm | `PROVEN` — exp 11 |
|
||||
| Adding OU mean-reversion (sp_ou_zscore) hurts on clean data | `PROVEN` — exp 25 |
|
||||
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3 runs, still panel-specific) |
|
||||
| Standalone 5d reversal (single feature sp_trend_slope_5) is not learnable — model trains positive IC | `PROVEN` — exp 43 (Q11) |
|
||||
| More features ≠ better signal on a small cross-section | `HYPOTHESIS` (supported by 3+ runs, still panel-specific) |
|
||||
|
||||
### 05 — Ensembles
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| 5-seed RankIC ensemble raises net-of-cost performance vs single model on the ablated set | `PROVEN` — exp 12 (pre-clean-lake), re-validated exp 22–24 |
|
||||
| Seed count is load-bearing: 2 seeds lose to 5 seeds on clean data | `PROVEN` — exp 28 |
|
||||
| 10 seeds raise rank breadth (RankIC 0.0671, L/S Sharpe 4.58) but the book stays negative net | `PROVEN` — exp 34 (Q02) |
|
||||
| Ensemble averaging's benefit is separable from feature expansion | `PROVEN` — exp 12 isolation design |
|
||||
|
||||
### 06 — Clean-lake reset
|
||||
@@ -97,31 +99,39 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
### 07 — Isolation runs
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 design |
|
||||
| Single-variable changes isolate what moved performance | `PROVEN` — exp 26→29/30/31 + exp 33–43 Q-runs design |
|
||||
| Multi-horizon momentum degrades the reference (net IR 0.21→-1.12) | `PROVEN` — exp 29 |
|
||||
| Risk-adjusted 22d Sharpe drift is promising on portfolio metrics, mixed on rank | `HYPOTHESIS` — exp 30 single run, unreproduced |
|
||||
| Risk-adjusted 22d Sharpe drift (M2): reproduced on the compact set by Q01 | `PROVEN` — exp 30 + exp 33 (Q01) |
|
||||
| GARCH(1,1) vol-regime features add no signal | `PROVEN` — exp 31 |
|
||||
| Longer labels raise IC monotonically but net worsens under daily turnover (10d/22d) | `PROVEN` — exp 36/37 (Q04/Q05) |
|
||||
| Standalone reversal feature does not reproduce | `PROVEN` — exp 43 (Q11) |
|
||||
|
||||
### 08 — Portfolio construction
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| TopkDropout beats stochastic-control OptimalStopControl on the ensemble signal | `PROVEN` — exp 13, 14 |
|
||||
| Stop-control constructions churn and bleed costs (cost drag ≈ −11.3pp) | `PROVEN` — exp 13 |
|
||||
| Fractional-Kelly sizing (exp 15) is unverified | `HYPOTHESIS` — run never finished |
|
||||
| Weekly rebalance recompute of the daily signal is the campaign's best construction (net +12.51%, IR 1.24) | `PROVEN` — exp 39 (Q07) |
|
||||
| Fractional-Kelly sizing (exp 15) is refuted on the clean lake (net +1.04%, IR 0.11) | `PROVEN` — exp 38 (Q06) |
|
||||
| Widening the book (topk 20) adds no edge; long-short top/bottom is destroyed by turnover | `PROVEN` — exp 35/41 (Q03/Q09) |
|
||||
|
||||
### 09 — Cost/turnover frontier
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| n_drop 2→1 flips net excess from −3.21% to +2.13% with identical signal metrics | `PROVEN` — exp 26 |
|
||||
| Cost drag is the binding constraint, not signal quality | `PROVEN` — exp 26 (IC/RankIC identical between n_drop variants) |
|
||||
| Weekly recompute cuts cost drag to ~1.1pp and unlocks +12.51% net | `PROVEN` — exp 39 (Q07) |
|
||||
| Long-short daily turnover costs 9.7% of NAV ($96.7k); fill rate 0.40 | `PROVEN` — exp 41 (Q09) |
|
||||
|
||||
### 10 — Risk limits
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| $5M liquidity floor improves net IR 0.81→0.98 and cuts drawdown 7.9%→5.4% | `PROVEN` — exp 18 (pre-clean-lake; see note in chapter) |
|
||||
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18 |
|
||||
| Size/concentration caps hurt by cutting deployed capital | `PROVEN` — exp 18; re-confirmed clean-lake exp 40 (Q08) |
|
||||
| Entry/risk gates are no-ops when the signal is the bottleneck | `PROVEN` — exp 20 (R2/R3 byte-identical) |
|
||||
| Exp-18 numbers are not comparable to post-reset runs due to env non-determinism | `PROVEN` — exp 20 R0 note |
|
||||
| Post-reset A/B: the floor binds but adds no IR edge; DD relief is pure defunding | `PROVEN` — exp 40 (Q08) |
|
||||
| HMM regime gate meets only the drawdown leg and churns | `PROVEN` — exp 42 (Q10) |
|
||||
|
||||
### 11 — Live execution and reconciliation
|
||||
| Claim | Expected status |
|
||||
@@ -133,8 +143,9 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
### 12 — Synthesis
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26 |
|
||||
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
|
||||
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26, 39 |
|
||||
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31, Q02–Q06, Q09–Q11) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
|
||||
| Turnover reduction is the dominant net-performance lever (weekly rebalance +12.51% vs daily −3.21%–+2.13%) | `PROVEN` — exp 26 vs 39 |
|
||||
| Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) |
|
||||
|
||||
## Repository layout
|
||||
@@ -153,7 +164,7 @@ book/
|
||||
|
||||
## Open questions for the desk
|
||||
|
||||
- `TODO(evidence-needed: reproduction of exp 30 M2 Sharpe-drift run on a second window)`
|
||||
- `TODO(evidence-needed: exp 15 Kelly sizing — run never finished; re-run on the clean lake)`
|
||||
- `TODO(evidence-needed: weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round)`
|
||||
- `TODO(evidence-needed: long-horizon label (10d/22d) paired with a low-turnover construction — the signal edge is proven, the cost kills daily churn)`
|
||||
- `TODO(evidence-needed: out-of-universe (non-ETF) validation of the compact stochastic feature set)`
|
||||
- `TODO(evidence-needed: reconciliation of exp 18 risk-limit spec on the post-reset reference signal)`
|
||||
- `TODO(evidence-needed: exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha)`
|
||||
Reference in New Issue
Block a user