book: add ch 11 walk-forward + 5 refuted guards (exp 52-56); EVIDENCE#043-047; rename 12→13-synthesis
This commit is contained in:
+7
-2
@@ -24,7 +24,7 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
|
||||
| EVIDENCE#008 | Risk-limit A/B: $5M liquidity floor → net IR 0.81→0.98, cumDD 7.93%→5.44%; size cap 15% + conc 60% hurts (IR 0.816, ann 6.11%). | exp 18, run `28c7fa08…` (mlflow exp 21), branch `exp/18-risk-limit-control-on-the-reference-ense` | NOT usable as PROVEN — pre-clean-lake (idea: liquidity floor > concentration caps) |
|
||||
| EVIDENCE#009 | Improvement sweep (R1-R5): 4/5 refuted; R2 momentum gate and R3 HMM gate are byte-identical no-ops; R5 MA3/EWMA marginal (IR 0.049). Conclusion: signal quality is the bottleneck, not the execution/risk layer. | exp 20, run `958198a8…` (mlflow exp 21), branch `exp/20-improve-the-risk-limit-reference-signal` | NOT usable as PROVEN — pre-clean-lake (idea: gates are no-ops when signal is weak) |
|
||||
|
||||
## Post-reset period (exp 21–43) — canonical, current
|
||||
## Post-reset period (exp 21–56) — canonical, current
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
@@ -59,6 +59,11 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
|
||||
| EVIDENCE#040 | Q17 realized-moments features (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) on clean data: improves portfolio over baseline. Net +9.90% (IR 0.990), gross +14.61%, maxDD −6.49% vs baseline net +2.13% (IR 0.21). IC 0.039 (vs 0.050), RankIC 0.060 (vs 0.066) — signal metrics slightly lower but portfolio construction benefits from moment conditioning. Contradicts pre-reset EVIDENCE#004 (which was inflated by dirty data). Single run, unreproduced. | exp 48, run `e62ce326…` (mlflow exp 48), branch `exp/48-q17-moments` | yes — Q17 HYPOTHESIS (needs reproduction) |
|
||||
| EVIDENCE#041 | Q18 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) vs TopkDropout on clean data: same signal (IC 0.050, RankIC 0.066 — identical model), worse portfolio. Net −6.21% (IR −0.640) vs baseline +2.13% (IR 0.21). Cost drag ~8.3pp. Clean-lake confirmation of pre-reset EVIDENCE#006/#007. | exp 49, run `f140dcb8…` (mlflow exp 49), branch `exp/49-q18-optstop` | yes — Q18 FAIL (confirms EVIDENCE#006/#007 on clean data) |
|
||||
| EVIDENCE#042 | Q21 10d label + weekly rebalance: cost drag cut from 4.61pp (Q04 daily) to 1.05pp (weekly). Net flipped from −9.92% to +1.19% (IR 0.148, maxDD −4.78%). Signal identical to Q04 (IC 0.093, RankIC 0.096). Weekly rebalance delivers ~10pp improvement regardless of label horizon (5d: +10.38pp via Q07, 10d: +11.11pp via Q21). But IR 0.148 < 0.5 acceptance — 5d+weekly (Q07, IR 1.24) remains the best construction. | exp 51, run `046c93a6…` (mlflow exp 51), branch `exp/51-q21-test-10d-label--weekly-rebalance-q04` | yes — Q21 FAIL (below IR bar, but confirms weekly-rebalance universality) |
|
||||
| EVIDENCE#043 | Walk-forward 3×3 (3 best configs × 2024/2025/2026): A weekly n_drop1 = −18.1% (IR −1.39) / −4.2% (IR −0.52) / **+12.5% (IR 1.25)**; B moments n_drop1 = −16.1% (IR −1.91) / −8.9% (IR −1.00) / **+9.2% (IR 0.94)**; C base n_drop2 = −18.2% (IR −1.97) / −3.6% (IR −0.51) / **−1.4% (IR −0.13)**. Only 2026 is profitable, and only for A/B. A and C share identical predictions (byte-identical IC/RankIC) — the strategy layer alone decides the outcome. Run A-2025 exactly replicated exp 45 (`e5ac7a5d`). The edge is a 2026-window-specific regime artifact. | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` | yes — walk-forward REFUTED (edge window-specific) |
|
||||
| EVIDENCE#044 | m2-sharpe22 3-window: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the exp-33 reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — harness is reproducible; edge is recent-window-only. | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference `c7c12228` (exp 33) | yes — walk-forward REFUTED (edge recent-window-only) |
|
||||
| EVIDENCE#045 | Label-regime transfer (2021/2023 — the closest label-regime PSI matches to 2026): 2023 −26.0% (IR −2.04, maxDD −30.8%), 2021 −22.9% (IR −2.26, maxDD −27.0%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two closest matches both lose ≈ a quarter. Feature-PSI gate also fails: 2026 has the highest feature drift yet the best result (CSRankNorm'd ranks are scale-invariant). No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) | yes — guard candidates 1+2 REFUTED |
|
||||
| EVIDENCE#046 | Adaptive short-window retrain (1y/2y rolling windows): 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing 2016→prev-Aug window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length. | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` | yes — guard candidate 4 REFUTED |
|
||||
| EVIDENCE#047 | Window-staleness isolation: pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in EVERY bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%) — the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug, Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). No staleness threshold isolates the edge. Account-based cumulative excess vs SPY: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38% (blotter `return` field excludes initial cost — use `account`). | exp 56, staleness analysis on exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` | yes — guard candidate 5 REFUTED |
|
||||
|
||||
## Live execution trail
|
||||
|
||||
@@ -71,7 +76,7 @@ Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_wit
|
||||
|
||||
| ID | Claim | Source | Verified? |
|
||||
|----|-------|--------|-----------|
|
||||
| (none yet) | — | — | — |
|
||||
| EVIDENCE#048 | Streaming IC circuit-breaker (`ic_min_rankic`, `ICGateTopkDropoutStrategy` in `tac_qlib/contrib/strategy/ic_gate.py`) trip-rate study: with thresholds 0.02–0.06, the gate trips on 25–50% of days in every year (2021–2026), freezing TopkDropout's rotation out of losers. A gate that trips every year cannot separate good years from bad. Do not deploy live. | ad-hoc scripted study on exp 52/53 pred/label artifacts, `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py`, `tac_qlib/tac_qlib/risk_limits.py` | yes — guard candidate 3 REFUTED |
|
||||
|
||||
## External references (book/references/)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user