book: add ch 11 walk-forward + 5 refuted guards (exp 52-56); EVIDENCE#043-047; rename 12→13-synthesis

This commit is contained in:
zhaoli
2026-08-20 21:03:59 +00:00
parent c0d21115eb
commit 18b61bbb8f
8 changed files with 144 additions and 19 deletions
@@ -78,7 +78,7 @@ Two facts about risk matter throughout the book. First, the measurement floor: o
Error is the book's discipline: how wrong was the prediction, in a way that can be measured against a null? The canonical metrics on the clean lake are IC, ICIR, Rank IC, Rank ICIR (the rank-based signal quality) and net IR, net return, L/S Sharpe, max drawdown (the portfolio outcome). Pre-reset runs (exp 8–18) recorded a different schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ir_with_cost`) — the two schemas are never compared directly in this book `(EVIDENCE.md: metric-schema note)`.
Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); live results are reconciled with slippage and cost, not taken from the backtest (ch. 11). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26, round 3)`.
Error defines the book's truth tiers: a claim is PROVEN only when reproduced on the clean lake with the canonical schema; a backtest alone is not a promise (ch. 06); a single-window backtest is re-validated walk-forward before shipping (ch. 11); live results are reconciled with slippage and cost, not taken from the backtest (ch. 12). The metrics ladder that runs through the whole book is: IC/RankIC (does the signal exist?) → net IR/MDD (does it survive cost?) → walk-forward (does it survive re-training?) → reconciled live funnel (does it execute?) — `(PROVEN → exp 21/24/26/52–56, round 3)`.
## Probability and significance
@@ -101,7 +101,7 @@ The cycle this book runs on: **measure → hypothesize → pre-register → isol
3. **Pre-register** — the claim and its acceptance metric (IC/RankIC above reference, net IR above reference) are fixed before execution, to block post-hoc cherry-picking across the 31+ experiments.
4. **Isolate** — one variable changes per run; the reference book and its metrics are the control (exp 26 → 29/30/31).
5. **Prove or refute** — on the clean lake only. A reproduced improvement becomes PROVEN; a single un-reproduced run stays HYPOTHESIS (exp 30); a degradation is REFUTED and — critically — is recorded as a win for the discipline (exp 29, exp 31 stopped wrong directions).
6. **Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 11).
6. **Reconcile live** — the proved book runs a round; targets→decisions→fills and slippage/cost reconcile against intent (round 3, ch. 12).
Every metric in this chapter sits on this loop. The drift metric produced the momentum hypothesis and the reversal hypothesis; only one survived isolation. The error metrics are the loop's judge. The decay and risk metrics are why ch. 03 and ch. 09 exist at all. The rest of the book is the working-out of this cycle, claim by claim, with each claim traceable to `EVIDENCE.md` and a recorded run.