| EVIDENCE#010 |
Clean-lake re-execution of the reference collapsed: IC 0.0019 (vs ref 0.0354), RankIC 0.0259, net −20.6% (IR −2.70). Old lake data quality had inflated the signal. |
exp 21, run f1bd3c28… (mlflow exp 23), branch exp/21-clean-lake-re-execution-of-the-tac-rd-ra |
yes |
| EVIDENCE#011 |
Re-run after fixing feature routing: IC 0.0486, RankIC 0.0617, ICIR 0.235, RankICIR 0.243, L/S Sharpe 3.23. |
exp 22, run 18db5bc1… (mlflow exp 24), branch exp/22-re-run-experiment-16s-5-day-rankic-ensem |
yes |
| EVIDENCE#012 |
General stochastic features only (no TA/HMM/OU): IC 0.0728, ICIR 0.340, L/S Sharpe 4.56. |
exp 23, run be5cd314… (mlflow exp 25), branch exp/23-test-whether-the-5-day-rankic-ensemble-i |
yes |
| EVIDENCE#013 |
Compact stochastic set (raw OHLCV + sp_ret, jump, RV1/5/22, vol ratios, trend slopes, logp, hurst, signature L1/L2): IC 0.0511, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54. |
exp 24, run fe469a19… (mlflow exp 25), branch exp/24-run-the-rankic-ensemble-in-mlflow-experi |
yes |
| EVIDENCE#014 |
Adding sp_ou_zscore hurts on clean data: IC 0.0343 vs 0.0511, net −3.76% vs −3.21%. |
exp 25, run 57450d1a… (mlflow exp 25), branch exp/25-test-the-clean-data-hypothesis-that-addi |
yes |
| EVIDENCE#015 |
n_drop 2→1 on identical compact stochastic signal: gross +7.02%, net +2.13% (vs −3.21%), MDD −7.69%, IR 0.21. IC/RankIC identical to n_drop 2 — the gain is turnover/cost relief. |
exp 26, run 21afc6af… (mlflow exp 25), branch exp/26-test-whether-reducing-topkdropout-daily |
yes — best result of the campaign |
| EVIDENCE#016 |
2-seed ensemble loses to 5-seed on clean data: RankIC 0.0579 vs 0.0663, net −1.49% (IR −0.14) vs +2.13% (IR 0.21). Seed count is load-bearing. |
exp 28, run c4ab1d01… (mlflow exp 27), branch exp/28-isolate-the-seed-count-effect-on-the-ndr |
yes |
| EVIDENCE#017 |
Multi-horizon momentum bundle refuted: IC 0.0337 vs 0.0511, net −13.35% (IR −1.12) vs +2.13%. |
exp 29, run b4586675… (mlflow exp 28), branch exp/29-isolation-run-m1-does-adding-multi-horiz |
yes |
| EVIDENCE#018 |
Risk-adjusted 22d Sharpe drift: mixed — rank metrics lower (RankIC 0.0576 vs 0.0663) but portfolio strong (net +6.53% IR 0.62 vs +2.13% IR 0.21). Single run, unreproduced. |
exp 30, run d5d775f9… (mlflow exp 29), branch exp/30-isolation-run-m2-does-adding-risk-adjust |
yes — mark HYPOTHESIS in text |
| EVIDENCE#019 |
GARCH(1,1) vol-regime trio refuted: IC 0.0415 vs 0.0511, RankICIR 0.179 vs 0.255, net +1.36% (IR 0.13). |
exp 31, run 514cb523… (mlflow exp 30), branch exp/31-isolation-run-m3-does-adding-garch11-vol |
yes |
| EVIDENCE#022 |
Q01 M2 repro: adding sp_sharpe_22 to the compact 25-field set reproduces exp-30 exactly (IC 0.0464, ICIR 0.211, RankIC 0.0578, RankICIR 0.231; net +6.53% IR 0.623, maxDD −8.0%, gross +11.41%). M2 Sharpe-drift edge confirmed on the compact set. |
exp 33, run c7c12228… (mlflow exp 32), branch exp/33-q01-m2-reproduction-add-spsharpe22-to-th |
yes — Q01 PASS |
| EVIDENCE#023 |
Q02 10-seed ensemble: breadth improves signal (RankIC 0.0671 vs 0.0579, RankICIR 0.259 vs 0.231, L/S Sharpe 4.58) but book stays negative net of cost (−0.93%, IR −0.089, maxDD −8.80%). |
exp 34, run ce49e4e0… (mlflow exp 33), branch exp/34-q02-seed10-10-seed-rankicensemble-vs-ref |
yes — Q02 FAIL (signal up, net down) |
| EVIDENCE#024 |
Q03 topk20: widening the book to 20 names cuts vol (std 0.0048 vs 0.0065) but adds no edge net of cost (−1.88%, IR −0.253, gross +0.64%, maxDD −8.80%). |
exp 35, run 2a844c02… (mlflow exp 34), branch exp/35-q03-topk20-widen-topkdropout-portfolio-f |
yes — Q03 FAIL |
| EVIDENCE#025 |
Q04 10d label: strongest IC of label series (IC 0.0925, ICIR 0.422, RankIC 0.0960, L/S Sharpe 5.89) but does not survive daily-rebalance cost (−9.92% net, IR −1.152, gross −5.31%). |
exp 36, run ef211826… (mlflow exp 35), branch exp/36-q04-label10d-10d-forward-return-label-vs |
yes — Q04 FAIL (horizon signal, daily churn) |
| EVIDENCE#026 |
Q05 22d label: best signal of all 11 (IC 0.0970, ICIR 0.526, RankIC 0.1165, RankICIR 0.507, L/S Sharpe 8.35) but book flat gross (−0.03%) / negative net (−4.60%, IR −0.588). Horizon gains never monetize under daily turnover. |
exp 37, run daad5042… (mlflow exp 36), branch exp/37-q05-label22d-22d-forward-return-label-vs |
yes — Q05 FAIL |
| EVIDENCE#027 |
Q06 Fractional-Kelly sizing (cap_frac 0.5): turns negative book mildly positive (+1.04% net, IR 0.112, maxDD −7.13%) and trims drawdown below the 7.69% bar, but far below the 0.21 net-IR acceptance. |
exp 38, run afca4b80… (mlflow exp 37), branch exp/38-q06-kelly-sizing-score-magnitude-fractio |
yes — Q06 FAIL (below bar) |
| EVIDENCE#028 |
Q07 weekly rebalance: weekly recompute of the same daily signal is the campaign's best result — net +12.51% (IR 1.243), maxDD −4.13%, cost drag only ~1.1pp (gross +13.59%). Same IC/RankIC as exp 26. |
exp 39, run eb38588c… (mlflow exp 38), branch exp/39-q07-weekly-rebalance-recompute-topkdropo |
yes — Q07 PASS, wins chapter |
| EVIDENCE#029 |
Q08 risk-limit A/B on the exp-26 pred: gates bind ($5M floor drops DBA,DBC,ESPO,FDN,REM,TAN,UNG,XAR) but no IR edge — candidate IR 1.512 < baseline 1.580; drawdown cut (−0.65% vs −6.91%) is pure defunding (size_cap×conc folds risk_degree to ~0.0095, ~$9.5k deployed of $1M). exp-18's floor improvement NOT reproduced on clean data. |
exp 40 (manual MLflow run 4667984187…, mlflow exp 43 tac-rd-q08-risklimit), branch exp/40-q08-risk-limit-ab-on-exp-26-reference-si, book/data/evidence/q08-risklimit/risk_calibration.json |
yes — Q08 REFUTED (safety net only) |
| EVIDENCE#030 |
Q09 long-short top10/bottom10: real pre-cost edge (gross +6.57%, IR 0.656) destroyed by daily L/S turnover — total_cost $96,721 (≈9.7% of $1M), 2485 trades/150d, fill rate 0.401; net −8.38%, IR −0.834, maxDD −11.22%. |
exp 41, run 0647eadd… (mlflow exp 39), branch exp/41-q09-long-short-market-neutral-long-top-1 |
yes — Q09 FAIL (turnover kills) |
| EVIDENCE#031 |
Q10 HMM regime entry gate (sp_hmm_p_regime1 ≥ 0.5 overlay): meets only the DD leg (−7.38% maxDD) — churns 276 trades/150d, cost ~6.3pp erases +2.02% gross; net −4.26%, IR −0.382. Regime-overlay hypothesis refuted. |
exp 42, run 436acd01… (mlflow exp 40), branch exp/42-q10-hmm-regime-overlay-entry-gate-on-sph |
yes — Q10 FAIL |
| EVIDENCE#032 |
Q11 standalone 5d reversal (single feature sp_trend_slope_5): IC is slightly positive (+0.0023), so the model did NOT learn reversal — the pooled trend-slope reversal beta does not reproduce standalone. Gross −10.36%, net −15.22% (IR −1.572). Cost is not the culprit. |
exp 43, run e859adfe… (mlflow exp 41), branch exp/43-q11-standalone-5d-reversal-single-featur |
yes — Q11 FAIL (no reversal learned) |
| EVIDENCE#033 |
Q14 out-of-universe validation: compact stochastic set on 30 liquid single-stock names (AAPL,MSFT,NVDA,…). RankIC −0.0198 (needed >0.03), ICIR −0.073 (needed >0.15) — signal is noise on this universe. Net P&L positive (+10.02% ann, IR 0.668, maxDD −6.67%) but that is top-10 concentration luck, not predictive signal. Train RankIC 0.316 shows the model overfits to the 50-ETF panel. |
exp 50, run 809ff460… (mlflow exp 50 tac-rd-q14-out-of-universe), branch exp/50-q14-compact-stochastic-set-generalizes-t |
yes — Q14 FAIL (signal does not generalize cross-universe) |
| EVIDENCE#034 |
Q19 variance-ratio study (Lo-MacKinlay robust VR): 71-ETF panel, 2015–2026. Median VR < 1 at all horizons — 5d: 0.925, 10d: 0.900, 20d: 0.884. 37–47% of ETFs have VR < 1 with |
z |
> 2 (significant mean-reversion). Only 1–3% show significant momentum. Assets are mean-reverting at short horizons on the clean lake. Note: pooled trend_slope_5 beta is strongly positive (+3.80, t=237) — the cross-sectional signal does NOT capture time-series mean-reversion. |
| EVIDENCE#035 |
Q20 effective independent names: eigenvalue analysis on 71-ETF correlation matrix (test window 2026-01-04 to 2026-08-10). Participation ratio = 4.46. Top-4 eigenvalues explain 66.8% of variance. 4 eigenvalues above Marchenko-Pastur bound (2.86). The 50-ETF book has ≈4.5 effective independent names — confirming the chat-derived claim. This explains why topk 10→20 adds no breadth (EVIDENCE#024). |
scripted study, book/data/evidence/q20-effective-names/eigenanalysis.py, eigenanalysis_50etf.csv, eigen_summary_50etf.json |
yes — Q20 PROVEN (diversification claim) |
| EVIDENCE#036 |
Q12 22d label + weekly rebalance: same IC/RankIC as Q05 (IC 0.097, RankIC 0.117 — identical training), but weekly recompute cannot rescue the stale signal. Net −4.88% (IR −0.566), gross +1.46%, maxDD −10.49%. The 22d label's problem is not daily turnover alone — the signal itself is stale. |
exp 44, run aed45c54… (mlflow exp 44), branch exp/44-q12-label22d-weekly |
yes — Q12 FAIL (redundant with Q05, confirms signal-stale hypothesis) |
| EVIDENCE#037 |
Q13 weekly rebalance on 2025 OOS window (train→2024-08-30, test 2025-01-02..2025-12-31): edge is window-dependent. IC 0.031 (vs Q07's 0.050), RankIC 0.073 (vs 0.066), L/S Sharpe 1.19 (vs 4.54). Net −4.21% (IR −0.523), maxDD −10.66%. Q07's +12.51% (IR 1.24) was specific to the 2026-01-04..2026-08-10 window. Weekly rebalance is not a robust edge. |
exp 45, run e5ac7a5d… (mlflow exp 45), branch exp/45-q13-weekly-oos |
yes — Q13 FAIL (limits Q07's generalizability) |
| EVIDENCE#038 |
Q15 single-seed vs 5-seed: 1 seed loses to 5 seeds on every metric. RankIC 0.044 vs 0.066, RankICIR 0.160 vs 0.255, net −2.89% (IR −0.278) vs +12.51% (IR 1.24). Clean-lake confirmation of EVIDENCE#016 (2-seed < 5-seed). Seed count is load-bearing. |
exp 46, run 8d49e0be… (mlflow exp 46), branch exp/46-q15-single-seed |
yes — Q15 FAIL (confirms EVIDENCE#016) |
| EVIDENCE#039 |
Q16 HMM features (sp_hmm_p_regime1, sp_hmm_state) on clean data: degrades both signal and portfolio. IC 0.030 (vs 0.050 baseline), RankIC 0.048 (vs 0.066), net −6.06% (IR −0.623), L/S Sharpe 1.23 (vs 4.54). HMM regime detection adds noise, not signal. |
exp 47, run ff092e1c… (mlflow exp 47), branch exp/47-q16-hmm |
yes — Q16 FAIL (HMM refuted on clean data) |
| EVIDENCE#040 |
Q17 realized-moments features (sp_rskew_5/22, sp_rkurt_5/22, sp_dsv_5/22) on clean data: improves portfolio over baseline. Net +9.90% (IR 0.990), gross +14.61%, maxDD −6.49% vs baseline net +2.13% (IR 0.21). IC 0.039 (vs 0.050), RankIC 0.060 (vs 0.066) — signal metrics slightly lower but portfolio construction benefits from moment conditioning. Contradicts pre-reset EVIDENCE#004 (which was inflated by dirty data). Single run, unreproduced. |
exp 48, run e62ce326… (mlflow exp 48), branch exp/48-q17-moments |
yes — Q17 HYPOTHESIS (needs reproduction) |
| EVIDENCE#041 |
Q18 OptimalStopControl (entry 0.85/exit 0.7/hold 10/sl −0.08) vs TopkDropout on clean data: same signal (IC 0.050, RankIC 0.066 — identical model), worse portfolio. Net −6.21% (IR −0.640) vs baseline +2.13% (IR 0.21). Cost drag ~8.3pp. Clean-lake confirmation of pre-reset EVIDENCE#006/#007. |
exp 49, run f140dcb8… (mlflow exp 49), branch exp/49-q18-optstop |
yes — Q18 FAIL (confirms EVIDENCE#006/#007 on clean data) |
| EVIDENCE#042 |
Q21 10d label + weekly rebalance: cost drag cut from 4.61pp (Q04 daily) to 1.05pp (weekly). Net flipped from −9.92% to +1.19% (IR 0.148, maxDD −4.78%). Signal identical to Q04 (IC 0.093, RankIC 0.096). Weekly rebalance delivers ~10pp improvement regardless of label horizon (5d: +10.38pp via Q07, 10d: +11.11pp via Q21). But IR 0.148 < 0.5 acceptance — 5d+weekly (Q07, IR 1.24) remains the best construction. |
exp 51, run 046c93a6… (mlflow exp 51), branch exp/51-q21-test-10d-label--weekly-rebalance-q04 |
yes — Q21 FAIL (below IR bar, but confirms weekly-rebalance universality) |
| EVIDENCE#043 |
Walk-forward 3×3 (3 best configs × 2024/2025/2026): A weekly n_drop1 = −18.1% (IR −1.39) / −4.2% (IR −0.52) / +12.5% (IR 1.25); B moments n_drop1 = −16.1% (IR −1.91) / −8.9% (IR −1.00) / +9.2% (IR 0.94); C base n_drop2 = −18.2% (IR −1.97) / −3.6% (IR −0.51) / −1.4% (IR −0.13). Only 2026 is profitable, and only for A/B. A and C share identical predictions (byte-identical IC/RankIC) — the strategy layer alone decides the outcome. Run A-2025 exactly replicated exp 45 (e5ac7a5d). The edge is a 2026-window-specific regime artifact. |
exp 52, mlflow exp 52 tac-rd-bt-3x3-windows (9 runs: 9f98ea5c A-2026, fe967416 A-2025, 71ed5bfa A-2024; 163c01ce B-2026, 4a85d68e B-2025, 1e49b8e8 B-2024; e3e06a24 C-2026, 353fff8f C-2025, 13a9bbdf C-2024), branch exp/52-walk-forward-re-validation-of-the-3-best |
yes — walk-forward REFUTED (edge window-specific) |
| EVIDENCE#044 |
m2-sharpe22 3-window: 2026 +6.5% (IR 0.623, maxDD −8.0%), 2025 +0.4% (IR 0.05), 2024 −26.4% (IR −2.11, maxDD −32.4%). The 2026 window reproduces the exp-33 reference almost exactly (IC 0.0464 vs 0.0464, RankIC 0.0578 vs 0.0578) — harness is reproducible; edge is recent-window-only. |
exp 53, mlflow exp 53 tac-rd-bt-m2-sharpe22-3windows (runs 7464c3e7 2026, 061f558b 2025, b49c6845 2024), branch exp/53-walk-forward-re-validation-of-m2-sharpe2; reference c7c12228 (exp 33) |
yes — walk-forward REFUTED (edge recent-window-only) |
| EVIDENCE#045 |
Label-regime transfer (2021/2023 — the closest label-regime PSI matches to 2026): 2023 −26.0% (IR −2.04, maxDD −30.8%), 2021 −22.9% (IR −2.26, maxDD −27.0%). Label-regime PSI similarity to 2026 ranks 2023 (0.028) > 2025 (0.035) > 2021 (0.039) — the two closest matches both lose ≈ a quarter. Feature-PSI gate also fails: 2026 has the highest feature drift yet the best result (CSRankNorm'd ranks are scale-invariant). No pre-deployment measurable gate — feature PSI, label-regime PSI, or drift — selects a profitable year. |
exp 54, mlflow exp 56 tac-rd-bt-m2-sharpe22-2021-2023 (runs 4e0700dd 2021, 8ca46e55 2023), branch exp/54-walk-forward-transfer-test-m2-sharpe22-o; feature/label-regime PSI study (exp 53 follow-up) |
yes — guard candidates 1+2 REFUTED |
| EVIDENCE#046 |
Adaptive short-window retrain (1y/2y rolling windows): 1y and 2y put every test year negative (2021 −15%/−18%, 2023 −20%/−23%, 2024 −14%/−19%, 2025 −6%/−3%, 2026 −10%/−6%); only the growing 2016→prev-Aug window ever went positive (2025 +0.4%, 2026 +6.5% IR 0.62). Short windows shave losses in bad years (2024 −26.4%→−13.9%) but destroy the 2026 edge (+6.5%→−9.6%). Mean annual excess ≈ −13% for every window length. |
exp 55, mlflow exp 57/58 tac-rd-bt-m2-sharpe22-adaptive-{1y,2y}, branch exp/55-adaptive-short-window-retrain-test-the-4 |
yes — guard candidate 4 REFUTED |
| EVIDENCE#047 |
Window-staleness isolation: pooled monthly excess (account vs SPY) by 90-day staleness bucket is negative in EVERY bucket (90d −17.4%, 180d −30.1%, 270d −17.7%, 360d −13.9%, 450d −9.7%) — the freshest bucket is the most negative. The 2026 edge is NOT concentrated in low-staleness days (best month Mar +8.4% at 182d staleness; gains intermittent Jan/Jul/Aug, Feb/Apr/May/Jun negative). 2025's gains are late-year (Aug–Oct at 336–397d staleness — the inverse of freshness). No staleness threshold isolates the edge. Account-based cumulative excess vs SPY: 2021 −27.9%, 2023 −30.4%, 2024 −31.6%, 2025 +0.25%, 2026 +4.38% (blotter return field excludes initial cost — use account). |
exp 56, staleness analysis on exp 53/54 pred/label artifacts, branch exp/56-window-staleness-isolation-the-m2-sharpe |
yes — guard candidate 5 REFUTED |