Files
tac-exp-dev/book/EVIDENCE.md
T

85 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Evidence Ledger
Every quantitative claim in the book lands here: id → claim → source (experiment/run/branch, round_id, script, citation) → verified?.
## Evidence boundary
**The clean-lake boundary (2026-08-18, exp 21) is the watermark.** `PROVEN` status in this book is reserved for the **Post-reset period** table (exp 21–31) and post-reset live rounds. The **Pre-clean-lake period** table below is **historical context and idea material only**: it was demonstrably inflated by lake data-quality problems (`EVIDENCE#010 → exp 21`). Pre-reset numbers may inform hypotheses but may never be cited as fact in the book.
## Key metric-schema note
Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ann_with_cost`, `excess_ir_with_cost`, `ls_ann_return`). Experiments 21+ use the canonical `IC / ICIR / Rank IC / Rank ICIR / net_IR / net_ann_return / gross_* / Long-Short_Ann_Sharpe / net_max_drawdown`. Do not compare schemas directly; chapter text states which schema a number comes from. Additionally, exp 20's R0 note states the exp-18 baseline is not comparable to post-reset runs due to environment non-determinism, and exp 21 invalidated all pre-clean-lake positive results.
## Pre-clean-lake period (exp 8–18) — historical context / idea material ONLY, superseded
| ID | Claim | Source | Verified? |
|----|-------|--------|-----------|
| EVIDENCE#001 | Baseline 1-day LGB signal weak on 2026 OOS: IC 0.017, ICIR 0.062, RankIC 0.040, RankICIR 0.161 (below 0.2 noise threshold). L/S ann +4.9%. | exp 8, run `e65cf1ec…` (mlflow exp 10), branch `exp/8-baseline-lightgbm-on-the-full-60etf-univ` | NOT usable as PROVEN — pre-clean-lake |
| EVIDENCE#002 | Costs erase most of the raw edge on baseline: excess +6.2% ann w/o cost (IR 0.31, MaxDD −20.4%) vs +1.6% ann after costs (IR 0.08). | exp 8 (same run) | NOT usable as PROVEN — pre-clean-lake |
| EVIDENCE#003 | Feature-family ablation: generic-only (jump,har,trend,hurst,signature,ret,max_move) beats all-24: RankIC 0.030→0.064, RankICIR 0.146→0.276, L/S Sharpe −0.83→+2.55, net excess −9.4%→+3.1%. | exp 9, run `7b1e7972…` (mlflow exp 11), branch `exp/9-sp5d-feature-family-ablation` | NOT usable as PROVEN — pre-clean-lake (idea: pruning generic beats model-specific) |
| EVIDENCE#004 | Adding 16 moment/volatility fields regresses every metric (RankIC 0.064→0.047, net excess −16.2% IR −1.57) — same failure mode as ou/hmm. | exp 11, run `a3f7d1d4…` (mlflow exp 12), branch `exp/11-sp5d-momentfeature-extension-after-exten` | NOT usable as PROVEN — pre-clean-lake (idea: panel width vs feature count) |
| EVIDENCE#005 | 5-seed RankIC ensemble on ablated generic features: RankIC 0.0586, RankICIR 0.224, net excess +7.8% (IR 0.79), L/S Sharpe 3.71, MDD −7.9%. Best pre-clean-lake net result. | exp 12, run `0cea66d9…` (mlflow exp 16), branch `exp/12-isolate-the-multiseed-rankic-ensemble-ef` | NOT usable as PROVEN — inflated by dirty lake (see EVIDENCE#010) |
| EVIDENCE#006 | OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) worse than TopkDropout: net excess −2.7% (IR −0.31) vs +7.8%; cost drag −11.3pp. | exp 13, run `4e1f77b4…` (mlflow exp 17), branch `exp/13-portfolioconstruction-variant-of-the-iso` | NOT usable as PROVEN — pre-clean-lake (idea: turnover-sensitive construction bleeds costs) |
| EVIDENCE#007 | OptimalStopControlV2 (turnover band/cooldown/cap) also refuted: net −6.9% (IR −0.72) vs TopkDropout +7.8% (IR 0.79). | exp 14, run `83d7e27e…` (mlflow exp 18), branch `exp/14-enhanced-stochasticcontrol-allocation-fo` | NOT usable as PROVEN — pre-clean-lake (idea only) |
| EVIDENCE#008 | Risk-limit A/B: $5M liquidity floor → net IR 0.81→0.98, cumDD 7.93%→5.44%; size cap 15% + conc 60% hurts (IR 0.816, ann 6.11%). | exp 18, run `28c7fa08…` (mlflow exp 21), branch `exp/18-risk-limit-control-on-the-reference-ense` | NOT usable as PROVEN — pre-clean-lake (idea: liquidity floor > concentration caps) |
| EVIDENCE#009 | Improvement sweep (R1-R5): 4/5 refuted; R2 momentum gate and R3 HMM gate are byte-identical no-ops; R5 MA3/EWMA marginal (IR 0.049). Conclusion: signal quality is the bottleneck, not the execution/risk layer. | exp 20, run `958198a8…` (mlflow exp 21), branch `exp/20-improve-the-risk-limit-reference-signal` | NOT usable as PROVEN — pre-clean-lake (idea: gates are no-ops when signal is weak) |
## Post-reset period (exp 21–56) — canonical, current
| ID | Claim | Source | Verified? |
|----|-------|--------|-----------|
| EVIDENCE#010 | Clean-lake re-execution of the reference collapsed: IC 0.0019 (vs ref 0.0354), RankIC 0.0259, net −20.6% (IR −2.70). Old lake data quality had inflated the signal. | exp 21, run `f1bd3c28…` (mlflow exp 23), branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` | yes |
| EVIDENCE#011 | Re-run after fixing feature routing: IC 0.0486, RankIC 0.0617, ICIR 0.235, RankICIR 0.243, L/S Sharpe 3.23. | exp 22, run `18db5bc1…` (mlflow exp 24), branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` | yes |
| EVIDENCE#012 | General stochastic features only (no TA/HMM/OU): IC 0.0728, ICIR 0.340, L/S Sharpe 4.56. | exp 23, run `be5cd314…` (mlflow exp 25), branch `exp/23-test-whether-the-5-day-rankic-ensemble-i` | yes |
| EVIDENCE#013 | Compact stochastic set (raw OHLCV + sp_ret, jump, RV1/5/22, vol ratios, trend slopes, logp, hurst, signature L1/L2): IC 0.0511, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54. | exp 24, run `fe469a19…` (mlflow exp 25), branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` | yes |
| EVIDENCE#014 | Adding sp_ou_zscore hurts on clean data: IC 0.0343 vs 0.0511, net −3.76% vs −3.21%. | exp 25, run `57450d1a…` (mlflow exp 25), branch `exp/25-test-the-clean-data-hypothesis-that-addi` | yes |
| EVIDENCE#015 | n_drop 2→1 on identical compact stochastic signal: gross +7.02%, net +2.13% (vs −3.21%), MDD −7.69%, IR 0.21. IC/RankIC identical to n_drop 2 — the gain is turnover/cost relief. | exp 26, run `21afc6af…` (mlflow exp 25), branch `exp/26-test-whether-reducing-topkdropout-daily` | yes — best result of the campaign |
| EVIDENCE#016 | 2-seed ensemble loses to 5-seed on clean data: RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21). Seed count is load-bearing. | exp 28, run `c4ab1d01…` (mlflow exp 27), branch `exp/28-isolate-the-seed-count-effect-on-the-ndr` | yes |
| EVIDENCE#017 | Multi-horizon momentum bundle refuted: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13%. | exp 29, run `b4586675…` (mlflow exp 28), branch `exp/29-isolation-run-m1-does-adding-multi-horiz` | yes |
| EVIDENCE#018 | Risk-adjusted 22d Sharpe drift: mixed — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio strong (net +6.53% IR 0.62 vs +2.13% IR 0.21). Single run, unreproduced. | exp 30, run `d5d775f9…` (mlflow exp 29), branch `exp/30-isolation-run-m2-does-adding-risk-adjust` | yes — mark HYPOTHESIS in text |
| EVIDENCE#019 | GARCH(1,1) vol-regime trio refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13). | exp 31, run `514cb523…` (mlflow exp 30), branch `exp/31-isolation-run-m3-does-adding-garch11-vol` | yes |
| EVIDENCE#022 | Q01 M2 repro: adding sp_sharpe_22 to the compact 25-field set reproduces exp-30 exactly (IC 0.0464, ICIR 0.211, RankIC 0.0578, RankICIR 0.231; net +6.53% IR 0.623, maxDD −8.0%, gross +11.41%). M2 Sharpe-drift edge confirmed on the compact set. | exp 33, run `c7c12228…` (mlflow exp 32), branch `exp/33-q01-m2-reproduction-add-spsharpe22-to-th` | yes — Q01 PASS |
| EVIDENCE#023 | Q02 10-seed ensemble: breadth improves signal (RankIC 0.0671 vs 0.0579, RankICIR 0.259 vs 0.231, L/S Sharpe 4.58) but book stays negative net of cost (−0.93%, IR −0.089, maxDD −8.80%). | exp 34, run `ce49e4e0…` (mlflow exp 33), branch `exp/34-q02-seed10-10-seed-rankicensemble-vs-ref` | yes — Q02 FAIL (signal up, net down) |
| EVIDENCE#024 | Q03 topk20: widening the book to 20 names cuts vol (std 0.0048 vs 0.0065) but adds no edge net of cost (−1.88%, IR −0.253, gross +0.64%, maxDD −8.80%). | exp 35, run `2a844c02…` (mlflow exp 34), branch `exp/35-q03-topk20-widen-topkdropout-portfolio-f` | yes — Q03 FAIL |
| EVIDENCE#025 | Q04 10d label: strongest IC of label series (IC 0.0925, ICIR 0.422, RankIC 0.0960, L/S Sharpe 5.89) but does not survive daily-rebalance cost (−9.92% net, IR −1.152, gross −5.31%). | exp 36, run `ef211826…` (mlflow exp 35), branch `exp/36-q04-label10d-10d-forward-return-label-vs` | yes — Q04 FAIL (horizon signal, daily churn) |
| EVIDENCE#026 | Q05 22d label: best signal of all 11 (IC 0.0970, ICIR 0.526, RankIC 0.1165, RankICIR 0.507, L/S Sharpe 8.35) but book flat gross (−0.03%) / negative net (−4.60%, IR −0.588). Horizon gains never monetize under daily turnover. | exp 37, run `daad5042…` (mlflow exp 36), branch `exp/37-q05-label22d-22d-forward-return-label-vs` | yes — Q05 FAIL |
| EVIDENCE#027 | Q06 Fractional-Kelly sizing (cap_frac 0.5): turns negative book mildly positive (+1.04% net, IR 0.112, maxDD −7.13%) and trims drawdown below the 7.69% bar, but far below the 0.21 net-IR acceptance. | exp 38, run `afca4b80…` (mlflow exp 37), branch `exp/38-q06-kelly-sizing-score-magnitude-fractio` | yes — Q06 FAIL (below bar) |
| EVIDENCE#028 | Q07 weekly rebalance: weekly recompute of the same daily signal is the campaign's best result — net +12.51% (IR 1.243), maxDD −4.13%, cost drag only ~1.1pp (gross +13.59%). Same IC/RankIC as exp 26. | exp 39, run `eb38588c…` (mlflow exp 38), branch `exp/39-q07-weekly-rebalance-recompute-topkdropo` | yes — Q07 PASS, wins chapter |
| EVIDENCE#029 | Q08 risk-limit A/B on the exp-26 pred: gates bind ($5M floor drops DBA,DBC,ESPO,FDN,REM,TAN,UNG,XAR) but no IR edge — candidate IR 1.512 < baseline 1.580; drawdown cut (−0.65% vs −6.91%) is pure defunding (size_cap×conc folds risk_degree to ~0.0095, ~$9.5k deployed of $1M). exp-18's floor improvement NOT reproduced on clean data. | exp 40 (manual MLflow run `4667984187…`, mlflow exp 43 `tac-rd-q08-risklimit`), branch `exp/40-q08-risk-limit-ab-on-exp-26-reference-si`, `book/data/evidence/q08-risklimit/risk_calibration.json` | yes — Q08 REFUTED (safety net only) |
| EVIDENCE#030 | Q09 long-short top10/bottom10: real pre-cost edge (gross +6.57%, IR 0.656) destroyed by daily L/S turnover — total_cost $96,721 (≈9.7% of $1M), 2485 trades/150d, fill rate 0.401; net −8.38%, IR −0.834, maxDD −11.22%. | exp 41, run `0647eadd…` (mlflow exp 39), branch `exp/41-q09-long-short-market-neutral-long-top-1` | yes — Q09 FAIL (turnover kills) |
| EVIDENCE#031 | Q10 HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5 overlay): meets only the DD leg (−7.38% maxDD) — churns 276 trades/150d, cost ~6.3pp erases +2.02% gross; net −4.26%, IR −0.382. Regime-overlay hypothesis refuted. | exp 42, run `436acd01…` (mlflow exp 40), branch `exp/42-q10-hmm-regime-overlay-entry-gate-on-sph` | yes — Q10 FAIL |
| EVIDENCE#032 | Q11 standalone 5d reversal (single feature sp_trend_slope_5): IC is slightly positive (+0.0023), so the model did NOT learn reversal — the pooled trend-slope reversal beta does not reproduce standalone. Gross −10.36%, net −15.22% (IR −1.572). Cost is not the culprit. | exp 43, run `e859adfe…` (mlflow exp 41), branch `exp/43-q11-standalone-5d-reversal-single-featur` | yes — Q11 FAIL (no reversal learned) |
| EVIDENCE#033 | Q14 out-of-universe validation: compact stochastic set on 30 liquid single-stock names (AAPL,MSFT,NVDA,…). RankIC −0.0198 (needed >0.03), ICIR −0.073 (needed >0.15) — signal is noise on this universe. Net P&L positive (+10.02% ann, IR 0.668, maxDD −6.67%) but that is top-10 concentration luck, not predictive signal. Train RankIC 0.316 shows the model overfits to the 50-ETF panel. | exp 50, run `809ff460…` (mlflow exp 50 `tac-rd-q14-out-of-universe`), branch `exp/50-q14-compact-stochastic-set-generalizes-t` | yes — Q14 FAIL (signal does not generalize cross-universe) |
| EVIDENCE#034 | Q19 variance-ratio study (Lo-MacKinlay robust VR): 71-ETF panel, 2015–2026. Median VR < 1 at all horizons — 5d: 0.925, 10d: 0.900, 20d: 0.884. 37–47% of ETFs have VR < 1 with |z| > 2 (significant mean-reversion). Only 1–3% show significant momentum. Assets are mean-reverting at short horizons on the clean lake. Note: pooled trend_slope_5 beta is strongly positive (+3.80, t=237) — the cross-sectional signal does NOT capture time-series mean-reversion. | scripted study, `book/data/evidence/q19-vr/vr_study.py`, VR_stats.csv, VR_summary.json | yes — Q19 PROVEN (market-structure claim) |
| EVIDENCE#035 | Q20 effective independent names: eigenvalue analysis on 71-ETF correlation matrix (test window 2026-01-04 to 2026-08-10). Participation ratio = 4.46. Top-4 eigenvalues explain 66.8% of variance. 4 eigenvalues above Marchenko-Pastur bound (2.86). The 50-ETF book has ≈4.5 effective independent names — confirming the chat-derived claim. This explains why topk 10→20 adds no breadth (EVIDENCE#024). | scripted study, `book/data/evidence/q20-effective-names/eigenanalysis.py`, eigenanalysis_50etf.csv, eigen_summary_50etf.json | yes — Q20 PROVEN (diversification claim) |
| EVIDENCE#036 | Q12 22d label + weekly rebalance: same IC/RankIC as Q05 (IC 0.097, RankIC 0.117 — identical training), but weekly recompute cannot rescue the stale signal. Net −4.88% (IR −0.566), gross +1.46%, maxDD −10.49%. The 22d label's problem is not daily turnover alone — the signal itself is stale. | exp 44, run `aed45c54…` (mlflow exp 44), branch `exp/44-q12-label22d-weekly` | yes — Q12 FAIL (redundant with Q05, confirms signal-stale hypothesis) |
| EVIDENCE#037 | Q13 weekly rebalance on 2025 OOS window (train→2024-08-30, test 2025-01-02..2025-12-31): edge is window-dependent. IC 0.031 (vs Q07's 0.050), RankIC 0.073 (vs 0.066), L/S Sharpe 1.19 (vs 4.54). Net −4.21% (IR −0.523), maxDD −10.66%. Q07's +12.51% (IR 1.24) was specific to the 2026-01-04..2026-08-10 window. Weekly rebalance is not a robust edge. | exp 45, run `e5ac7a5d…` (mlflow exp 45), branch `exp/45-q13-weekly-oos` | yes — Q13 FAIL (limits Q07's generalizability) |
| EVIDENCE#038 | Q15 single-seed vs 5-seed: 1 seed loses to 5 seeds on every metric. RankIC 0.044 vs 0.066, RankICIR 0.160 vs 0.255, net −2.89% (IR −0.278) vs +12.51% (IR 1.24). Clean-lake confirmation of EVIDENCE#016 (2-seed < 5-seed). Seed count is load-bearing. | exp 46, run `8d49e0be…` (mlflow exp 46), branch `exp/46-q15-single-seed` | yes — Q15 FAIL (confirms EVIDENCE#016) |
| EVIDENCE#039 | Q16 HMM features (sp_hmm_p_regime1, sp_hmm_state) on clean data: degrades both signal and portfolio. IC 0.030 (vs 0.050 baseline), RankIC 0.048 (vs 0.066), net −6.06% (IR −0.623), L/S Sharpe 1.23 (vs 4.54). HMM regime detection adds noise, not signal. | exp 47, run `ff092e1c…` (mlflow exp 47), branch `exp/47-q16-hmm` | yes — Q16 FAIL (HMM refuted on clean data) |
| EVIDENCE#040 | Q17 realized-moments features (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) on clean data: improves portfolio over baseline. Net +9.90% (IR 0.990), gross +14.61%, maxDD −6.49% vs baseline net +2.13% (IR 0.21). IC 0.039 (vs 0.050), RankIC 0.060 (vs 0.066) — signal metrics slightly lower but portfolio construction benefits from moment conditioning. Contradicts pre-reset EVIDENCE#004 (which was inflated by dirty data). Single run, unreproduced. | exp 48, run `e62ce326…` (mlflow exp 48), branch `exp/48-q17-moments` | yes — Q17 HYPOTHESIS (needs reproduction) |
| EVIDENCE#041 | Q18 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) vs TopkDropout on clean data: same signal (IC 0.050, RankIC 0.066 — identical model), worse portfolio. Net −6.21% (IR −0.640) vs baseline +2.13% (IR 0.21). Cost drag ~8.3pp. Clean-lake confirmation of pre-reset EVIDENCE#006/#007. | exp 49, run `f140dcb8…` (mlflow exp 49), branch `exp/49-q18-optstop` | yes — Q18 FAIL (confirms EVIDENCE#006/#007 on clean data) |
| EVIDENCE#042 | Q21 10d label + weekly rebalance: cost drag cut from 4.61pp (Q04 daily) to 1.05pp (weekly). Net flipped from −9.92% to +1.19% (IR 0.148, maxDD −4.78%). Signal identical to Q04 (IC 0.093, RankIC 0.096). Weekly rebalance delivers ~10pp improvement regardless of label horizon (5d: +10.38pp via Q07, 10d: +11.11pp via Q21). But IR 0.148 < 0.5 acceptance — 5d+weekly (Q07, IR 1.24) remains the best construction. | exp 51, run `046c93a6…` (mlflow exp 51), branch `exp/51-q21-test-10d-label--weekly-rebalance-q04` | yes — Q21 FAIL (below IR bar, but confirms weekly-rebalance universality) |
| EVIDENCE#043 | Walk-forward 3×3 (3 best configs × 2024/2025/2026): A weekly n_drop1 = −18.1% (IR −1.39) / −4.2% (IR −0.52) / **+12.5% (IR 1.25)**; B moments n_drop1 = −16.1% (IR −1.91) / −8.9% (IR −1.00) / **+9.2% (IR 0.94)**; C base n_drop2 = −18.2% (IR −1.97) / −3.6% (IR −0.51) / **−1.4% (IR −0.13)**. Only 2026 is profitable, and only for A/B. A and C share identical predictions (byte-identical IC/RankIC) — the strategy layer alone decides the outcome. Run A-2025 exactly replicated exp 45 (`e5ac7a5d`). The edge is a 2026-window-specific regime artifact. | exp 52, mlflow exp 52 `tac-rd-bt-3x3-windows` (9 runs: `9f98ea5c` A-2026, `fe967416` A-2025, `71ed5bfa` A-2024; `163c01ce` B-2026, `4a85d68e` B-2025, `1e49b8e8` B-2024; `e3e06a24` C-2026, `353fff8f` C-2025, `13a9bbdf` C-2024), branch `exp/52-walk-forward-re-validation-of-the-3-best` | yes — walk-forward REFUTED (edge window-specific) |
| EVIDENCE#044 | m2-sharpe22 3-window: 2026 **+6.5%** (IR 0.623, maxDD −8.0%), 2025 **+0.4%** (IR 0.05), 2024 **−26.4%** (IR −2.11, maxDD −32.4%). The 2026 window reproduces the exp-33 reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — harness is reproducible; edge is recent-window-only. | exp 53, mlflow exp 53 `tac-rd-bt-m2-sharpe22-3windows` (runs `7464c3e7` 2026, `061f558b` 2025, `b49c6845` 2024), branch `exp/53-walk-forward-re-validation-of-m2-sharpe2`; reference `c7c12228` (exp 33) | yes — walk-forward REFUTED (edge recent-window-only) |
| EVIDENCE#045 | Label-regime transfer (2021/2023 — the closest label-regime PSI matches to 2026): 2023 −26.0% (IR −2.04, maxDD −30.8%), 2021 −22.9% (IR −2.26, maxDD −27.0%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two closest matches both lose ≈ a quarter. Feature-PSI gate also fails: 2026 has the highest feature drift yet the best result (CSRankNorm'd ranks are scale-invariant). No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. | exp 54, mlflow exp 56 `tac-rd-bt-m2-sharpe22-2021-2023` (runs `4e0700dd` 2021, `8ca46e55` 2023), branch `exp/54-walk-forward-transfer-test-m2-sharpe22-o`; feature/label-regime PSI study (exp 53 follow-up) | yes — guard candidates 1+2 REFUTED |
| EVIDENCE#046 | Adaptive short-window retrain (1y/2y rolling windows): 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing 2016→prev-Aug window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length. | exp 55, mlflow exp 57/58 `tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}`, branch `exp/55-adaptive-short-window-retrain-test-the-4` | yes — guard candidate 4 REFUTED |
| EVIDENCE#047 | Window-staleness isolation: pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in EVERY bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%) — the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug, Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). No staleness threshold isolates the edge. Account-based cumulative excess vs SPY: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38% (blotter `return` field excludes initial cost — use `account`). | exp 56, staleness analysis on exp 53/54 pred/label artifacts, branch `exp/56-window-staleness-isolation-the-m2-sharpe` | yes — guard candidate 5 REFUTED |
## Live execution trail
| ID | Claim | Source | Verified? |
|----|-------|--------|-----------|
| EVIDENCE#020 | Live round 3 (target 2026-08-17): retrained exp-26 n_drop=1 config on rolling 4y window; Topk10/n_drop1 with risk limits (liq floor $5M dropped 8, size cap 12%, conc 95%, drawdown pause 10%); funnel 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped (SLV delta_zero); invested $74,202.85, slippage 4.54 bps, est. cost ~$45. | round 3 (`tac-rd-book`), trace 27, run `721ef257…` (mlflow exp 26), branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac` | yes — settled, reconcile available |
| EVIDENCE#021 | Scheduled retrain on 2026-08-14 (pre-reset reference): 10 buys + 6 sells placed, 0 cancelled by sentiment gate; sized on live equity $99,999.93. | trace 16, run `3b858b2b…` (mlflow exp 13), branch `exp/16-scheduled-algo-retrain-on-20260814-tacrd` | yes — historical, pre-reset signal |
## Ad-hoc scripts (book/data/)
| ID | Claim | Source | Verified? |
|----|-------|--------|-----------|
| EVIDENCE#048 | Streaming IC circuit-breaker (`ic_min_rankic`, `ICGateTopkDropoutStrategy` in `tac_qlib/contrib/strategy/ic_gate.py`) trip-rate study: with thresholds 0.02–0.06, the gate trips on 25–50% of days in every year (2021–2026), freezing TopkDropout's rotation out of losers. A gate that trips every year cannot separate good years from bad. Do not deploy live. | ad-hoc scripted study on exp 52/53 pred/label artifacts, `tac_qlib/tac_qlib/contrib/strategy/ic_gate.py`, `tac_qlib/tac_qlib/risk_limits.py` | yes — guard candidate 3 REFUTED |
## External references (book/references/)
| ID | Claim | Source | Verified? |
|----|-------|--------|-----------|
| (none yet) | — | — | — |