book: add ch 11 walk-forward + 5 refuted guards (exp 52-56); EVIDENCE#043-047; rename 12→13-synthesis
This commit is contained in:
+20
-7
@@ -33,8 +33,9 @@ A quant-desk reader should be able to act on this book: replicate a signal pipel
|
||||
| 08 | Portfolio construction: dropout vs the rest | drafting | exp 13, 14, 15, 35, 38, 39, 41 | Turnover-sensitive construction bleeds the edge; weekly recompute wins |
|
||||
| 09 | The cost/turnover frontier | drafting | exp 26, 39, 41 | Cut turnover before adding signal; weekly rebalance is the proven lever |
|
||||
| 10 | Risk limits and gates that work | drafting | exp 18, 20, 40, 42 | Limits are a safety net, not alpha; gates churn without signal |
|
||||
| 11 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
|
||||
| 12 | Synthesis: how proved truth compounds | drafting | all, exp 33–43 | Cost relief > signal; the Q-campaign scoreboard |
|
||||
| 11 | Walk-forward re-validation and guard candidates | drafting | exp 52–56 | The edge is a 2025–2026 regime artifact; all 5 guards refuted |
|
||||
| 12 | Live execution and reconciliation | drafting | exp 27, round 3 | 4.54 bps slippage realized; funnel 10→10→10→9 |
|
||||
| 13 | Synthesis: how proved truth compounds | drafting | all, exp 33–43, 52–56 | Cost relief > signal; the Q-campaign scoreboard |
|
||||
|
||||
Status legend: `drafting` → `in-review` → `done`.
|
||||
|
||||
@@ -52,7 +53,7 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
### 01 — Metrics: the vocabulary of a price series
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–12) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) |
|
||||
| Every chapter claim reduces to a statistic computable on the lake (drift, jump, vol, regime, reversion, memory, risk, error, probability, timeline, decay) | `PROVEN` (chapters 03–13) + `HYPOTHESIS` (dataset-study magnitudes, chat-derived) |
|
||||
| Generic scale-free statistics beat model-specific machinery on a small daily panel | `PROVEN` (exp 23/24/25/29/31) + `HYPOTHESIS` (generality) |
|
||||
| A statistic is only as good as the falsification it survives (null z-scores, reproduction) | `PROVEN` (exp 21 detection playbook) + `REFERENCED` |
|
||||
| The strongest single-feature signal (OU z-score) can be worthless inside a rank model — the "OU paradox" | `PROVEN` (exp 25) + open mechanism `TODO(evidence-needed)` |
|
||||
@@ -133,19 +134,28 @@ Each chapter opens with its claims. The inventory below is the working contract:
|
||||
| Post-reset A/B: the floor binds but adds no IR edge; DD relief is pure defunding | `PROVEN` — exp 40 (Q08) |
|
||||
| HMM regime gate meets only the drawdown leg and churns | `PROVEN` — exp 42 (Q10) |
|
||||
|
||||
### 11 — Live execution and reconciliation
|
||||
### 11 — Walk-forward re-validation and guard candidates
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The headline results (weekly +12.51%, m2-sharpe22 +6.5%) are 2026-window-specific; walk-forward re-training across 2024/2025 is negative or flat | `PROVEN` — exp 52/53 |
|
||||
| A and C share identical predictions; the strategy layer alone decides the outcome | `PROVEN` — exp 52 |
|
||||
| No pre-deployment measurable gate (feature-PSI, label-regime PSI, streaming IC, window length, staleness) selects a profitable year | `PROVEN` — exp 52–56, all 5 guards refuted |
|
||||
| The edge is a 2025–2026 regime artifact; live capital must be cut until the regime returns | `PROVEN` (walk-forward) + `HYPOTHESIS` (forward-looking) |
|
||||
|
||||
### 12 — Live execution and reconciliation
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| Live funnel held: 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped | `PROVEN` — round 3 |
|
||||
| Realized slippage ≈ 4.54 bps, estimated cost ≈ $45, turnover 0.74 | `PROVEN` — round 3 metrics |
|
||||
| Live beats backtest: execution claims trace to round_id, not to backtest | `PROVEN` — methodology |
|
||||
|
||||
### 12 — Synthesis
|
||||
### 13 — Synthesis
|
||||
| Claim | Expected status |
|
||||
|-------|-----------------|
|
||||
| The largest performance deltas came from data quality, cost/turnover relief, feature pruning, and risk limits — not from adding features | `PROVEN` — composite of exp 9, 18, 21, 26, 39 |
|
||||
| The campaign's refuted runs (exp 11, 13, 14, 20, 25, 29, 31, Q02–Q06, Q09–Q11) were as valuable as wins | `REFERENCED` + `PROVEN` (they stopped wrong directions) |
|
||||
| Turnover reduction is the dominant net-performance lever (weekly rebalance +12.51% vs daily −3.21%–+2.13%) | `PROVEN` — exp 26 vs 39 |
|
||||
| The campaign's headline edges were a 2025–2026 regime artifact, not robust OOS | `PROVEN` — exp 52–56 |
|
||||
| Generalizability of the 50-ETF panel results is an open question | `HYPOTHESIS` — TODO(evidence-needed: out-of-panel universe) |
|
||||
|
||||
## Repository layout
|
||||
@@ -167,13 +177,16 @@ book/
|
||||
- `TODO(evidence-needed: a second live round beyond round 3, to confirm slippage and funnel hold under a different market regime)`
|
||||
- `TODO(evidence-needed: reconcile realized cost against the 5bp/15bp/$5 backtest model over a full position window)`
|
||||
- `TODO(evidence-needed: whether sp_sharpe_22 still helps when combined with the weekly-rebalance construction of ch. 08)`
|
||||
- `TODO(evidence-needed: purged walk-forward CV instead of single train/valid split on the exp-26 reference)`
|
||||
- `TODO(evidence-needed: automated lake-integrity check wired into every experiment run, not only on demand)`
|
||||
- `TODO(evidence-needed: live round under weekly-rebalance construction with risk-limit spec, to confirm safety-net behavior at higher deployed capital)`
|
||||
- `TODO(evidence-needed: a live window that matches the 2026 label regime, to test whether the edge returns when the regime returns)`
|
||||
- `TODO(evidence-needed: a causal (no-lookahead) regime-change detector that selects the 2026 window before the fact — none of the five guards did)`
|
||||
|
||||
### Settled open questions (no longer active)
|
||||
|
||||
- ~~`weekly-rebalance result (exp 39) reproduced on a second window before promotion to a live round`~~ — **ANSWERED (negatively):** Q13 (exp 45) tested weekly on 2025 OOS: net −4.21% IR −0.52. The edge is window-dependent, not robust. `EVIDENCE#037`.
|
||||
- ~~`long-horizon label (10d/22d) paired with a low-turnover construction`~~ — **ANSWERED:** Q12 (exp 44): 22d+weekly net −4.88% IR −0.566. Q21 (exp 51): 10d+weekly net +1.19% IR 0.148. Both below IR 0.5 acceptance. Weekly is a universal cost lever (~10pp improvement) but the5d label remains the sweet spot. `EVIDENCE#036/042`.
|
||||
- ~~`out-of-universe (non-ETF) validation of the compact stochastic feature set`~~ — **ANSWERED (negatively):** Q14 (exp 50): RankIC −0.02, ICIR −0.07 on 30 liquid single-stock names. Signal is noise outside the 50-ETF panel. `EVIDENCE#033`.
|
||||
- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`.
|
||||
- ~~`exp 18 risk-limit spec reconciliation — post-reset A/B (exp 40) shows it is a safety net, not alpha`~~ — **ANSWERED:** Q08 (exp 40): $5M floor binds but adds no IR edge (candidate 1.512 < baseline 1.580). DD relief is pure defunding. `EVIDENCE#029`.
|
||||
- ~~`do the headline results survive walk-forward re-training?`~~ — **ANSWERED (negatively):** exp 52/53/54 re-ran weekly, moments, ndrop2, and m2-sharpe22 across 2024–2026 (plus 2021/2023 label-regime matches). Only 2026 is profitable; all prior years negative or flat. Edge = 2025–2026 regime artifact. `EVIDENCE#043–045`.
|
||||
- ~~`is there a pre-deployment guard that isolates the profitable regime?`~~ — **ANSWERED (negatively):** feature-PSI, label-regime PSI, streaming IC (`ic_min_rankic`), adaptive short-window, and staleness guards all refuted. `EVIDENCE#043–047`.
|
||||
Reference in New Issue
Block a user