Files
tac-exp-dev/book/EVIDENCE.md
T
zhaoli 06fb1e8ee9 book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches
- CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims
- README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign
- new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated
- Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
2026-08-20 01:23:21 +00:00

12 KiB
Raw Blame History

Evidence Ledger

Every quantitative claim in the book lands here: id → claim → source (experiment/run/branch, round_id, script, citation) → verified?.

Evidence boundary

The clean-lake boundary (2026-08-18, exp 21) is the watermark. PROVEN status in this book is reserved for the Post-reset period table (exp 21–31) and post-reset live rounds. The Pre-clean-lake period table below is historical context and idea material only: it was demonstrably inflated by lake data-quality problems (EVIDENCE#010 → exp 21). Pre-reset numbers may inform hypotheses but may never be cited as fact in the book.

Key metric-schema note

Experiments 8–18 record metrics under a legacy schema (ls_sharpe, maxdd_with_cost, excess_ann_with_cost, excess_ir_with_cost, ls_ann_return). Experiments 21+ use the canonical IC / ICIR / Rank IC / Rank ICIR / net_IR / net_ann_return / gross_* / Long-Short_Ann_Sharpe / net_max_drawdown. Do not compare schemas directly; chapter text states which schema a number comes from. Additionally, exp 20's R0 note states the exp-18 baseline is not comparable to post-reset runs due to environment non-determinism, and exp 21 invalidated all pre-clean-lake positive results.

Pre-clean-lake period (exp 8–18) — historical context / idea material ONLY, superseded

ID Claim Source Verified?
EVIDENCE#001 Baseline 1-day LGB signal weak on 2026 OOS: IC 0.017, ICIR 0.062, RankIC 0.040, RankICIR 0.161 (below 0.2 noise threshold). L/S ann +4.9%. exp 8, run e65cf1ec… (mlflow exp 10), branch exp/8-baseline-lightgbm-on-the-full-60etf-univ NOT usable as PROVEN — pre-clean-lake
EVIDENCE#002 Costs erase most of the raw edge on baseline: excess +6.2% ann w/o cost (IR 0.31, MaxDD −20.4%) vs +1.6% ann after costs (IR 0.08). exp 8 (same run) NOT usable as PROVEN — pre-clean-lake
EVIDENCE#003 Feature-family ablation: generic-only (jump,har,trend,hurst,signature,ret,max_move) beats all-24: RankIC 0.030→0.064, RankICIR 0.146→0.276, L/S Sharpe −0.83→+2.55, net excess −9.4%→+3.1%. exp 9, run 7b1e7972… (mlflow exp 11), branch exp/9-sp5d-feature-family-ablation NOT usable as PROVEN — pre-clean-lake (idea: pruning generic beats model-specific)
EVIDENCE#004 Adding 16 moment/volatility fields regresses every metric (RankIC 0.064→0.047, net excess −16.2% IR −1.57) — same failure mode as ou/hmm. exp 11, run a3f7d1d4… (mlflow exp 12), branch exp/11-sp5d-momentfeature-extension-after-exten NOT usable as PROVEN — pre-clean-lake (idea: panel width vs feature count)
EVIDENCE#005 5-seed RankIC ensemble on ablated generic features: RankIC 0.0586, RankICIR 0.224, net excess +7.8% (IR 0.79), L/S Sharpe 3.71, MDD −7.9%. Best pre-clean-lake net result. exp 12, run 0cea66d9… (mlflow exp 16), branch exp/12-isolate-the-multiseed-rankic-ensemble-ef NOT usable as PROVEN — inflated by dirty lake (see EVIDENCE#010)
EVIDENCE#006 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) worse than TopkDropout: net excess −2.7% (IR −0.31) vs +7.8%; cost drag −11.3pp. exp 13, run 4e1f77b4… (mlflow exp 17), branch exp/13-portfolioconstruction-variant-of-the-iso NOT usable as PROVEN — pre-clean-lake (idea: turnover-sensitive construction bleeds costs)
EVIDENCE#007 OptimalStopControlV2 (turnover band/cooldown/cap) also refuted: net −6.9% (IR −0.72) vs TopkDropout +7.8% (IR 0.79). exp 14, run 83d7e27e… (mlflow exp 18), branch exp/14-enhanced-stochasticcontrol-allocation-fo NOT usable as PROVEN — pre-clean-lake (idea only)
EVIDENCE#008 Risk-limit A/B: $5M liquidity floor → net IR 0.81→0.98, cumDD 7.93%→5.44%; size cap 15% + conc 60% hurts (IR 0.816, ann 6.11%). exp 18, run 28c7fa08… (mlflow exp 21), branch exp/18-risk-limit-control-on-the-reference-ense NOT usable as PROVEN — pre-clean-lake (idea: liquidity floor > concentration caps)
EVIDENCE#009 Improvement sweep (R1-R5): 4/5 refuted; R2 momentum gate and R3 HMM gate are byte-identical no-ops; R5 MA3/EWMA marginal (IR 0.049). Conclusion: signal quality is the bottleneck, not the execution/risk layer. exp 20, run 958198a8… (mlflow exp 21), branch exp/20-improve-the-risk-limit-reference-signal NOT usable as PROVEN — pre-clean-lake (idea: gates are no-ops when signal is weak)

Post-reset period (exp 21–43) — canonical, current

ID Claim Source Verified?
EVIDENCE#010 Clean-lake re-execution of the reference collapsed: IC 0.0019 (vs ref 0.0354), RankIC 0.0259, net −20.6% (IR −2.70). Old lake data quality had inflated the signal. exp 21, run f1bd3c28… (mlflow exp 23), branch exp/21-clean-lake-re-execution-of-the-tac-rd-ra yes
EVIDENCE#011 Re-run after fixing feature routing: IC 0.0486, RankIC 0.0617, ICIR 0.235, RankICIR 0.243, L/S Sharpe 3.23. exp 22, run 18db5bc1… (mlflow exp 24), branch exp/22-re-run-experiment-16s-5-day-rankic-ensem yes
EVIDENCE#012 General stochastic features only (no TA/HMM/OU): IC 0.0728, ICIR 0.340, L/S Sharpe 4.56. exp 23, run be5cd314… (mlflow exp 25), branch exp/23-test-whether-the-5-day-rankic-ensemble-i yes
EVIDENCE#013 Compact stochastic set (raw OHLCV + sp_ret, jump, RV1/5/22, vol ratios, trend slopes, logp, hurst, signature L1/L2): IC 0.0511, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54. exp 24, run fe469a19… (mlflow exp 25), branch exp/24-run-the-rankic-ensemble-in-mlflow-experi yes
EVIDENCE#014 Adding sp_ou_zscore hurts on clean data: IC 0.0343 vs 0.0511, net −3.76% vs −3.21%. exp 25, run 57450d1a… (mlflow exp 25), branch exp/25-test-the-clean-data-hypothesis-that-addi yes
EVIDENCE#015 n_drop 2→1 on identical compact stochastic signal: gross +7.02%, net +2.13% (vs −3.21%), MDD −7.69%, IR 0.21. IC/RankIC identical to n_drop 2 — the gain is turnover/cost relief. exp 26, run 21afc6af… (mlflow exp 25), branch exp/26-test-whether-reducing-topkdropout-daily yes — best result of the campaign
EVIDENCE#016 2-seed ensemble loses to 5-seed on clean data: RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21). Seed count is load-bearing. exp 28, run c4ab1d01… (mlflow exp 27), branch exp/28-isolate-the-seed-count-effect-on-the-ndr yes
EVIDENCE#017 Multi-horizon momentum bundle refuted: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13%. exp 29, run b4586675… (mlflow exp 28), branch exp/29-isolation-run-m1-does-adding-multi-horiz yes
EVIDENCE#018 Risk-adjusted 22d Sharpe drift: mixed — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio strong (net +6.53% IR 0.62 vs +2.13% IR 0.21). Single run, unreproduced. exp 30, run d5d775f9… (mlflow exp 29), branch exp/30-isolation-run-m2-does-adding-risk-adjust yes — mark HYPOTHESIS in text
EVIDENCE#019 GARCH(1,1) vol-regime trio refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13). exp 31, run 514cb523… (mlflow exp 30), branch exp/31-isolation-run-m3-does-adding-garch11-vol yes
EVIDENCE#022 Q01 M2 repro: adding sp_sharpe_22 to the compact 25-field set reproduces exp-30 exactly (IC 0.0464, ICIR 0.211, RankIC 0.0578, RankICIR 0.231; net +6.53% IR 0.623, maxDD −8.0%, gross +11.41%). M2 Sharpe-drift edge confirmed on the compact set. exp 33, run c7c12228… (mlflow exp 32), branch exp/33-q01-m2-reproduction-add-spsharpe22-to-th yes — Q01 PASS
EVIDENCE#023 Q02 10-seed ensemble: breadth improves signal (RankIC 0.0671 vs 0.0579, RankICIR 0.259 vs 0.231, L/S Sharpe 4.58) but book stays negative net of cost (−0.93%, IR −0.089, maxDD −8.80%). exp 34, run ce49e4e0… (mlflow exp 33), branch exp/34-q02-seed10-10-seed-rankicensemble-vs-ref yes — Q02 FAIL (signal up, net down)
EVIDENCE#024 Q03 topk20: widening the book to 20 names cuts vol (std 0.0048 vs 0.0065) but adds no edge net of cost (−1.88%, IR −0.253, gross +0.64%, maxDD −8.80%). exp 35, run 2a844c02… (mlflow exp 34), branch exp/35-q03-topk20-widen-topkdropout-portfolio-f yes — Q03 FAIL
EVIDENCE#025 Q04 10d label: strongest IC of label series (IC 0.0925, ICIR 0.422, RankIC 0.0960, L/S Sharpe 5.89) but does not survive daily-rebalance cost (−9.92% net, IR −1.152, gross −5.31%). exp 36, run ef211826… (mlflow exp 35), branch exp/36-q04-label10d-10d-forward-return-label-vs yes — Q04 FAIL (horizon signal, daily churn)
EVIDENCE#026 Q05 22d label: best signal of all 11 (IC 0.0970, ICIR 0.526, RankIC 0.1165, RankICIR 0.507, L/S Sharpe 8.35) but book flat gross (−0.03%) / negative net (−4.60%, IR −0.588). Horizon gains never monetize under daily turnover. exp 37, run daad5042… (mlflow exp 36), branch exp/37-q05-label22d-22d-forward-return-label-vs yes — Q05 FAIL
EVIDENCE#027 Q06 Fractional-Kelly sizing (cap_frac 0.5): turns negative book mildly positive (+1.04% net, IR 0.112, maxDD −7.13%) and trims drawdown below the 7.69% bar, but far below the 0.21 net-IR acceptance. exp 38, run afca4b80… (mlflow exp 37), branch exp/38-q06-kelly-sizing-score-magnitude-fractio yes — Q06 FAIL (below bar)
EVIDENCE#028 Q07 weekly rebalance: weekly recompute of the same daily signal is the campaign's best result — net +12.51% (IR 1.243), maxDD −4.13%, cost drag only ~1.1pp (gross +13.59%). Same IC/RankIC as exp 26. exp 39, run eb38588c… (mlflow exp 38), branch exp/39-q07-weekly-rebalance-recompute-topkdropo yes — Q07 PASS, wins chapter
EVIDENCE#029 Q08 risk-limit A/B on the exp-26 pred: gates bind ($5M floor drops DBA,DBC,ESPO,FDN,REM,TAN,UNG,XAR) but no IR edge — candidate IR 1.512 < baseline 1.580; drawdown cut (−0.65% vs −6.91%) is pure defunding (size_cap×conc folds risk_degree to ~0.0095, ~$9.5k deployed of $1M). exp-18's floor improvement NOT reproduced on clean data. exp 40 (manual MLflow run 4667984187…, mlflow exp 43 tac-rd-q08-risklimit), branch exp/40-q08-risk-limit-ab-on-exp-26-reference-si, book/data/evidence/q08-risklimit/risk_calibration.json yes — Q08 REFUTED (safety net only)
EVIDENCE#030 Q09 long-short top10/bottom10: real pre-cost edge (gross +6.57%, IR 0.656) destroyed by daily L/S turnover — total_cost $96,721 (≈9.7% of $1M), 2485 trades/150d, fill rate 0.401; net −8.38%, IR −0.834, maxDD −11.22%. exp 41, run 0647eadd… (mlflow exp 39), branch exp/41-q09-long-short-market-neutral-long-top-1 yes — Q09 FAIL (turnover kills)
EVIDENCE#031 Q10 HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5 overlay): meets only the DD leg (−7.38% maxDD) — churns 276 trades/150d, cost ~6.3pp erases +2.02% gross; net −4.26%, IR −0.382. Regime-overlay hypothesis refuted. exp 42, run 436acd01… (mlflow exp 40), branch exp/42-q10-hmm-regime-overlay-entry-gate-on-sph yes — Q10 FAIL
EVIDENCE#032 Q11 standalone 5d reversal (single feature sp_trend_slope_5): IC is slightly positive (+0.0023), so the model did NOT learn reversal — the pooled trend-slope reversal beta does not reproduce standalone. Gross −10.36%, net −15.22% (IR −1.572). Cost is not the culprit. exp 43, run e859adfe… (mlflow exp 41), branch exp/43-q11-standalone-5d-reversal-single-featur yes — Q11 FAIL (no reversal learned)

Live execution trail

ID Claim Source Verified?
EVIDENCE#020 Live round 3 (target 2026-08-17): retrained exp-26 n_drop=1 config on rolling 4y window; Topk10/n_drop1 with risk limits (liq floor $5M dropped 8, size cap 12%, conc 95%, drawdown pause 10%); funnel 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped (SLV delta_zero); invested $74,202.85, slippage 4.54 bps, est. cost ~$45. round 3 (tac-rd-book), trace 27, run 721ef257… (mlflow exp 26), branch exp/27-scheduled-algo-retrain-on-2026-08-17-tac yes — settled, reconcile available
EVIDENCE#021 Scheduled retrain on 2026-08-14 (pre-reset reference): 10 buys + 6 sells placed, 0 cancelled by sentiment gate; sized on live equity $99,999.93. trace 16, run 3b858b2b… (mlflow exp 13), branch exp/16-scheduled-algo-retrain-on-20260814-tacrd yes — historical, pre-reset signal

Ad-hoc scripts (book/data/)

ID Claim Source Verified?
(none yet) — — —

External references (book/references/)

ID Claim Source Verified?
(none yet) — — —