80 lines
18 KiB
Markdown
80 lines
18 KiB
Markdown
# Evidence Ledger
|
||
|
||
Every quantitative claim in the book lands here: id → claim → source (experiment/run/branch, round_id, script, citation) → verified?.
|
||
|
||
## Evidence boundary
|
||
|
||
**The clean-lake boundary (2026-08-18, exp 21) is the watermark.** `PROVEN` status in this book is reserved for the **Post-reset period** table (exp 21–31) and post-reset live rounds. The **Pre-clean-lake period** table below is **historical context and idea material only**: it was demonstrably inflated by lake data-quality problems (`EVIDENCE#010 → exp 21`). Pre-reset numbers may inform hypotheses but may never be cited as fact in the book.
|
||
|
||
## Key metric-schema note
|
||
|
||
Experiments 8–18 record metrics under a legacy schema (`ls_sharpe`, `maxdd_with_cost`, `excess_ann_with_cost`, `excess_ir_with_cost`, `ls_ann_return`). Experiments 21+ use the canonical `IC / ICIR / Rank IC / Rank ICIR / net_IR / net_ann_return / gross_* / Long-Short_Ann_Sharpe / net_max_drawdown`. Do not compare schemas directly; chapter text states which schema a number comes from. Additionally, exp 20's R0 note states the exp-18 baseline is not comparable to post-reset runs due to environment non-determinism, and exp 21 invalidated all pre-clean-lake positive results.
|
||
|
||
## Pre-clean-lake period (exp 8–18) — historical context / idea material ONLY, superseded
|
||
|
||
| ID | Claim | Source | Verified? |
|
||
|----|-------|--------|-----------|
|
||
| EVIDENCE#001 | Baseline 1-day LGB signal weak on 2026 OOS: IC 0.017, ICIR 0.062, RankIC 0.040, RankICIR 0.161 (below 0.2 noise threshold). L/S ann +4.9%. | exp 8, run `e65cf1ec…` (mlflow exp 10), branch `exp/8-baseline-lightgbm-on-the-full-60etf-univ` | NOT usable as PROVEN — pre-clean-lake |
|
||
| EVIDENCE#002 | Costs erase most of the raw edge on baseline: excess +6.2% ann w/o cost (IR 0.31, MaxDD −20.4%) vs +1.6% ann after costs (IR 0.08). | exp 8 (same run) | NOT usable as PROVEN — pre-clean-lake |
|
||
| EVIDENCE#003 | Feature-family ablation: generic-only (jump,har,trend,hurst,signature,ret,max_move) beats all-24: RankIC 0.030→0.064, RankICIR 0.146→0.276, L/S Sharpe −0.83→+2.55, net excess −9.4%→+3.1%. | exp 9, run `7b1e7972…` (mlflow exp 11), branch `exp/9-sp5d-feature-family-ablation` | NOT usable as PROVEN — pre-clean-lake (idea: pruning generic beats model-specific) |
|
||
| EVIDENCE#004 | Adding 16 moment/volatility fields regresses every metric (RankIC 0.064→0.047, net excess −16.2% IR −1.57) — same failure mode as ou/hmm. | exp 11, run `a3f7d1d4…` (mlflow exp 12), branch `exp/11-sp5d-momentfeature-extension-after-exten` | NOT usable as PROVEN — pre-clean-lake (idea: panel width vs feature count) |
|
||
| EVIDENCE#005 | 5-seed RankIC ensemble on ablated generic features: RankIC 0.0586, RankICIR 0.224, net excess +7.8% (IR 0.79), L/S Sharpe 3.71, MDD −7.9%. Best pre-clean-lake net result. | exp 12, run `0cea66d9…` (mlflow exp 16), branch `exp/12-isolate-the-multiseed-rankic-ensemble-ef` | NOT usable as PROVEN — inflated by dirty lake (see EVIDENCE#010) |
|
||
| EVIDENCE#006 | OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) worse than TopkDropout: net excess −2.7% (IR −0.31) vs +7.8%; cost drag −11.3pp. | exp 13, run `4e1f77b4…` (mlflow exp 17), branch `exp/13-portfolioconstruction-variant-of-the-iso` | NOT usable as PROVEN — pre-clean-lake (idea: turnover-sensitive construction bleeds costs) |
|
||
| EVIDENCE#007 | OptimalStopControlV2 (turnover band/cooldown/cap) also refuted: net −6.9% (IR −0.72) vs TopkDropout +7.8% (IR 0.79). | exp 14, run `83d7e27e…` (mlflow exp 18), branch `exp/14-enhanced-stochasticcontrol-allocation-fo` | NOT usable as PROVEN — pre-clean-lake (idea only) |
|
||
| EVIDENCE#008 | Risk-limit A/B: $5M liquidity floor → net IR 0.81→0.98, cumDD 7.93%→5.44%; size cap 15% + conc 60% hurts (IR 0.816, ann 6.11%). | exp 18, run `28c7fa08…` (mlflow exp 21), branch `exp/18-risk-limit-control-on-the-reference-ense` | NOT usable as PROVEN — pre-clean-lake (idea: liquidity floor > concentration caps) |
|
||
| EVIDENCE#009 | Improvement sweep (R1-R5): 4/5 refuted; R2 momentum gate and R3 HMM gate are byte-identical no-ops; R5 MA3/EWMA marginal (IR 0.049). Conclusion: signal quality is the bottleneck, not the execution/risk layer. | exp 20, run `958198a8…` (mlflow exp 21), branch `exp/20-improve-the-risk-limit-reference-signal` | NOT usable as PROVEN — pre-clean-lake (idea: gates are no-ops when signal is weak) |
|
||
|
||
## Post-reset period (exp 21–43) — canonical, current
|
||
|
||
| ID | Claim | Source | Verified? |
|
||
|----|-------|--------|-----------|
|
||
| EVIDENCE#010 | Clean-lake re-execution of the reference collapsed: IC 0.0019 (vs ref 0.0354), RankIC 0.0259, net −20.6% (IR −2.70). Old lake data quality had inflated the signal. | exp 21, run `f1bd3c28…` (mlflow exp 23), branch `exp/21-clean-lake-re-execution-of-the-tac-rd-ra` | yes |
|
||
| EVIDENCE#011 | Re-run after fixing feature routing: IC 0.0486, RankIC 0.0617, ICIR 0.235, RankICIR 0.243, L/S Sharpe 3.23. | exp 22, run `18db5bc1…` (mlflow exp 24), branch `exp/22-re-run-experiment-16s-5-day-rankic-ensem` | yes |
|
||
| EVIDENCE#012 | General stochastic features only (no TA/HMM/OU): IC 0.0728, ICIR 0.340, L/S Sharpe 4.56. | exp 23, run `be5cd314…` (mlflow exp 25), branch `exp/23-test-whether-the-5-day-rankic-ensemble-i` | yes |
|
||
| EVIDENCE#013 | Compact stochastic set (raw OHLCV + sp_ret, jump, RV1/5/22, vol ratios, trend slopes, logp, hurst, signature L1/L2): IC 0.0511, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54. | exp 24, run `fe469a19…` (mlflow exp 25), branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` | yes |
|
||
| EVIDENCE#014 | Adding sp_ou_zscore hurts on clean data: IC 0.0343 vs 0.0511, net −3.76% vs −3.21%. | exp 25, run `57450d1a…` (mlflow exp 25), branch `exp/25-test-the-clean-data-hypothesis-that-addi` | yes |
|
||
| EVIDENCE#015 | n_drop 2→1 on identical compact stochastic signal: gross +7.02%, net +2.13% (vs −3.21%), MDD −7.69%, IR 0.21. IC/RankIC identical to n_drop 2 — the gain is turnover/cost relief. | exp 26, run `21afc6af…` (mlflow exp 25), branch `exp/26-test-whether-reducing-topkdropout-daily` | yes — best result of the campaign |
|
||
| EVIDENCE#016 | 2-seed ensemble loses to 5-seed on clean data: RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21). Seed count is load-bearing. | exp 28, run `c4ab1d01…` (mlflow exp 27), branch `exp/28-isolate-the-seed-count-effect-on-the-ndr` | yes |
|
||
| EVIDENCE#017 | Multi-horizon momentum bundle refuted: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13%. | exp 29, run `b4586675…` (mlflow exp 28), branch `exp/29-isolation-run-m1-does-adding-multi-horiz` | yes |
|
||
| EVIDENCE#018 | Risk-adjusted 22d Sharpe drift: mixed — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio strong (net +6.53% IR 0.62 vs +2.13% IR 0.21). Single run, unreproduced. | exp 30, run `d5d775f9…` (mlflow exp 29), branch `exp/30-isolation-run-m2-does-adding-risk-adjust` | yes — mark HYPOTHESIS in text |
|
||
| EVIDENCE#019 | GARCH(1,1) vol-regime trio refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13). | exp 31, run `514cb523…` (mlflow exp 30), branch `exp/31-isolation-run-m3-does-adding-garch11-vol` | yes |
|
||
| EVIDENCE#022 | Q01 M2 repro: adding sp_sharpe_22 to the compact 25-field set reproduces exp-30 exactly (IC 0.0464, ICIR 0.211, RankIC 0.0578, RankICIR 0.231; net +6.53% IR 0.623, maxDD −8.0%, gross +11.41%). M2 Sharpe-drift edge confirmed on the compact set. | exp 33, run `c7c12228…` (mlflow exp 32), branch `exp/33-q01-m2-reproduction-add-spsharpe22-to-th` | yes — Q01 PASS |
|
||
| EVIDENCE#023 | Q02 10-seed ensemble: breadth improves signal (RankIC 0.0671 vs 0.0579, RankICIR 0.259 vs 0.231, L/S Sharpe 4.58) but book stays negative net of cost (−0.93%, IR −0.089, maxDD −8.80%). | exp 34, run `ce49e4e0…` (mlflow exp 33), branch `exp/34-q02-seed10-10-seed-rankicensemble-vs-ref` | yes — Q02 FAIL (signal up, net down) |
|
||
| EVIDENCE#024 | Q03 topk20: widening the book to 20 names cuts vol (std 0.0048 vs 0.0065) but adds no edge net of cost (−1.88%, IR −0.253, gross +0.64%, maxDD −8.80%). | exp 35, run `2a844c02…` (mlflow exp 34), branch `exp/35-q03-topk20-widen-topkdropout-portfolio-f` | yes — Q03 FAIL |
|
||
| EVIDENCE#025 | Q04 10d label: strongest IC of label series (IC 0.0925, ICIR 0.422, RankIC 0.0960, L/S Sharpe 5.89) but does not survive daily-rebalance cost (−9.92% net, IR −1.152, gross −5.31%). | exp 36, run `ef211826…` (mlflow exp 35), branch `exp/36-q04-label10d-10d-forward-return-label-vs` | yes — Q04 FAIL (horizon signal, daily churn) |
|
||
| EVIDENCE#026 | Q05 22d label: best signal of all 11 (IC 0.0970, ICIR 0.526, RankIC 0.1165, RankICIR 0.507, L/S Sharpe 8.35) but book flat gross (−0.03%) / negative net (−4.60%, IR −0.588). Horizon gains never monetize under daily turnover. | exp 37, run `daad5042…` (mlflow exp 36), branch `exp/37-q05-label22d-22d-forward-return-label-vs` | yes — Q05 FAIL |
|
||
| EVIDENCE#027 | Q06 Fractional-Kelly sizing (cap_frac 0.5): turns negative book mildly positive (+1.04% net, IR 0.112, maxDD −7.13%) and trims drawdown below the 7.69% bar, but far below the 0.21 net-IR acceptance. | exp 38, run `afca4b80…` (mlflow exp 37), branch `exp/38-q06-kelly-sizing-score-magnitude-fractio` | yes — Q06 FAIL (below bar) |
|
||
| EVIDENCE#028 | Q07 weekly rebalance: weekly recompute of the same daily signal is the campaign's best result — net +12.51% (IR 1.243), maxDD −4.13%, cost drag only ~1.1pp (gross +13.59%). Same IC/RankIC as exp 26. | exp 39, run `eb38588c…` (mlflow exp 38), branch `exp/39-q07-weekly-rebalance-recompute-topkdropo` | yes — Q07 PASS, wins chapter |
|
||
| EVIDENCE#029 | Q08 risk-limit A/B on the exp-26 pred: gates bind ($5M floor drops DBA,DBC,ESPO,FDN,REM,TAN,UNG,XAR) but no IR edge — candidate IR 1.512 < baseline 1.580; drawdown cut (−0.65% vs −6.91%) is pure defunding (size_cap×conc folds risk_degree to ~0.0095, ~$9.5k deployed of $1M). exp-18's floor improvement NOT reproduced on clean data. | exp 40 (manual MLflow run `4667984187…`, mlflow exp 43 `tac-rd-q08-risklimit`), branch `exp/40-q08-risk-limit-ab-on-exp-26-reference-si`, `book/data/evidence/q08-risklimit/risk_calibration.json` | yes — Q08 REFUTED (safety net only) |
|
||
| EVIDENCE#030 | Q09 long-short top10/bottom10: real pre-cost edge (gross +6.57%, IR 0.656) destroyed by daily L/S turnover — total_cost $96,721 (≈9.7% of $1M), 2485 trades/150d, fill rate 0.401; net −8.38%, IR −0.834, maxDD −11.22%. | exp 41, run `0647eadd…` (mlflow exp 39), branch `exp/41-q09-long-short-market-neutral-long-top-1` | yes — Q09 FAIL (turnover kills) |
|
||
| EVIDENCE#031 | Q10 HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5 overlay): meets only the DD leg (−7.38% maxDD) — churns 276 trades/150d, cost ~6.3pp erases +2.02% gross; net −4.26%, IR −0.382. Regime-overlay hypothesis refuted. | exp 42, run `436acd01…` (mlflow exp 40), branch `exp/42-q10-hmm-regime-overlay-entry-gate-on-sph` | yes — Q10 FAIL |
|
||
| EVIDENCE#032 | Q11 standalone 5d reversal (single feature sp_trend_slope_5): IC is slightly positive (+0.0023), so the model did NOT learn reversal — the pooled trend-slope reversal beta does not reproduce standalone. Gross −10.36%, net −15.22% (IR −1.572). Cost is not the culprit. | exp 43, run `e859adfe…` (mlflow exp 41), branch `exp/43-q11-standalone-5d-reversal-single-featur` | yes — Q11 FAIL (no reversal learned) |
|
||
| EVIDENCE#033 | Q14 out-of-universe validation: compact stochastic set on 30 liquid single-stock names (AAPL,MSFT,NVDA,…). RankIC −0.0198 (needed >0.03), ICIR −0.073 (needed >0.15) — signal is noise on this universe. Net P&L positive (+10.02% ann, IR 0.668, maxDD −6.67%) but that is top-10 concentration luck, not predictive signal. Train RankIC 0.316 shows the model overfits to the 50-ETF panel. | exp 50, run `809ff460…` (mlflow exp 50 `tac-rd-q14-out-of-universe`), branch `exp/50-q14-compact-stochastic-set-generalizes-t` | yes — Q14 FAIL (signal does not generalize cross-universe) |
|
||
| EVIDENCE#034 | Q19 variance-ratio study (Lo-MacKinlay robust VR): 71-ETF panel, 2015–2026. Median VR < 1 at all horizons — 5d: 0.925, 10d: 0.900, 20d: 0.884. 37–47% of ETFs have VR < 1 with |z| > 2 (significant mean-reversion). Only 1–3% show significant momentum. Assets are mean-reverting at short horizons on the clean lake. Note: pooled trend_slope_5 beta is strongly positive (+3.80, t=237) — the cross-sectional signal does NOT capture time-series mean-reversion. | scripted study, `book/data/evidence/q19-vr/vr_study.py`, VR_stats.csv, VR_summary.json | yes — Q19 PROVEN (market-structure claim) |
|
||
| EVIDENCE#035 | Q20 effective independent names: eigenvalue analysis on 71-ETF correlation matrix (test window 2026-01-04 to 2026-08-10). Participation ratio = 4.46. Top-4 eigenvalues explain 66.8% of variance. 4 eigenvalues above Marchenko-Pastur bound (2.86). The 50-ETF book has ≈4.5 effective independent names — confirming the chat-derived claim. This explains why topk 10→20 adds no breadth (EVIDENCE#024). | scripted study, `book/data/evidence/q20-effective-names/eigenanalysis.py`, eigenanalysis_50etf.csv, eigen_summary_50etf.json | yes — Q20 PROVEN (diversification claim) |
|
||
| EVIDENCE#036 | Q12 22d label + weekly rebalance: same IC/RankIC as Q05 (IC 0.097, RankIC 0.117 — identical training), but weekly recompute cannot rescue the stale signal. Net −4.88% (IR −0.566), gross +1.46%, maxDD −10.49%. The 22d label's problem is not daily turnover alone — the signal itself is stale. | exp 44, run `aed45c54…` (mlflow exp 44), branch `exp/44-q12-label22d-weekly` | yes — Q12 FAIL (redundant with Q05, confirms signal-stale hypothesis) |
|
||
| EVIDENCE#037 | Q13 weekly rebalance on 2025 OOS window (train→2024-08-30, test 2025-01-02..2025-12-31): edge is window-dependent. IC 0.031 (vs Q07's 0.050), RankIC 0.073 (vs 0.066), L/S Sharpe 1.19 (vs 4.54). Net −4.21% (IR −0.523), maxDD −10.66%. Q07's +12.51% (IR 1.24) was specific to the 2026-01-04..2026-08-10 window. Weekly rebalance is not a robust edge. | exp 45, run `e5ac7a5d…` (mlflow exp 45), branch `exp/45-q13-weekly-oos` | yes — Q13 FAIL (limits Q07's generalizability) |
|
||
| EVIDENCE#038 | Q15 single-seed vs 5-seed: 1 seed loses to 5 seeds on every metric. RankIC 0.044 vs 0.066, RankICIR 0.160 vs 0.255, net −2.89% (IR −0.278) vs +12.51% (IR 1.24). Clean-lake confirmation of EVIDENCE#016 (2-seed < 5-seed). Seed count is load-bearing. | exp 46, run `8d49e0be…` (mlflow exp 46), branch `exp/46-q15-single-seed` | yes — Q15 FAIL (confirms EVIDENCE#016) |
|
||
| EVIDENCE#039 | Q16 HMM features (sp_hmm_p_regime1, sp_hmm_state) on clean data: degrades both signal and portfolio. IC 0.030 (vs 0.050 baseline), RankIC 0.048 (vs 0.066), net −6.06% (IR −0.623), L/S Sharpe 1.23 (vs 4.54). HMM regime detection adds noise, not signal. | exp 47, run `ff092e1c…` (mlflow exp 47), branch `exp/47-q16-hmm` | yes — Q16 FAIL (HMM refuted on clean data) |
|
||
| EVIDENCE#040 | Q17 realized-moments features (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) on clean data: improves portfolio over baseline. Net +9.90% (IR 0.990), gross +14.61%, maxDD −6.49% vs baseline net +2.13% (IR 0.21). IC 0.039 (vs 0.050), RankIC 0.060 (vs 0.066) — signal metrics slightly lower but portfolio construction benefits from moment conditioning. Contradicts pre-reset EVIDENCE#004 (which was inflated by dirty data). Single run, unreproduced. | exp 48, run `e62ce326…` (mlflow exp 48), branch `exp/48-q17-moments` | yes — Q17 HYPOTHESIS (needs reproduction) |
|
||
| EVIDENCE#041 | Q18 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) vs TopkDropout on clean data: same signal (IC 0.050, RankIC 0.066 — identical model), worse portfolio. Net −6.21% (IR −0.640) vs baseline +2.13% (IR 0.21). Cost drag ~8.3pp. Clean-lake confirmation of pre-reset EVIDENCE#006/#007. | exp 49, run `f140dcb8…` (mlflow exp 49), branch `exp/49-q18-optstop` | yes — Q18 FAIL (confirms EVIDENCE#006/#007 on clean data) |
|
||
| EVIDENCE#042 | Q21 10d label + weekly rebalance: cost drag cut from 4.61pp (Q04 daily) to 1.05pp (weekly). Net flipped from −9.92% to +1.19% (IR 0.148, maxDD −4.78%). Signal identical to Q04 (IC 0.093, RankIC 0.096). Weekly rebalance delivers ~10pp improvement regardless of label horizon (5d: +10.38pp via Q07, 10d: +11.11pp via Q21). But IR 0.148 < 0.5 acceptance — 5d+weekly (Q07, IR 1.24) remains the best construction. | exp 51, run `046c93a6…` (mlflow exp 51), branch `exp/51-q21-test-10d-label--weekly-rebalance-q04` | yes — Q21 FAIL (below IR bar, but confirms weekly-rebalance universality) |
|
||
|
||
## Live execution trail
|
||
|
||
| ID | Claim | Source | Verified? |
|
||
|----|-------|--------|-----------|
|
||
| EVIDENCE#020 | Live round 3 (target 2026-08-17): retrained exp-26 n_drop=1 config on rolling 4y window; Topk10/n_drop1 with risk limits (liq floor $5M dropped 8, size cap 12%, conc 95%, drawdown pause 10%); funnel 10 targets → 10 decided → 10 placed → 9 filled, 1 cancelled, 1 skipped (SLV delta_zero); invested $74,202.85, slippage 4.54 bps, est. cost ~$45. | round 3 (`tac-rd-book`), trace 27, run `721ef257…` (mlflow exp 26), branch `exp/27-scheduled-algo-retrain-on-2026-08-17-tac` | yes — settled, reconcile available |
|
||
| EVIDENCE#021 | Scheduled retrain on 2026-08-14 (pre-reset reference): 10 buys + 6 sells placed, 0 cancelled by sentiment gate; sized on live equity $99,999.93. | trace 16, run `3b858b2b…` (mlflow exp 13), branch `exp/16-scheduled-algo-retrain-on-20260814-tacrd` | yes — historical, pre-reset signal |
|
||
|
||
## Ad-hoc scripts (book/data/)
|
||
|
||
| ID | Claim | Source | Verified? |
|
||
|----|-------|--------|-----------|
|
||
| (none yet) | — | — | — |
|
||
|
||
## External references (book/references/)
|
||
|
||
| ID | Claim | Source | Verified? |
|
||
|----|-------|--------|-----------|
|
||
| (none yet) | — | — | — | |