6.3 KiB
6.3 KiB
Chat-Mined Ideas & Hypotheses
Source: opencode chat transcripts under book/data/chat_mining/ (historical context, pre-clean-lake). Per the evidence contract these are idea material only — none may be cited as PROVEN. Each idea below is a hypothesis to be tested on the clean lake (exp 21+).
Data-quality failure classes (feed ch. 06)
These are the classes of failure documented across exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt. Durable lessons even though exact numbers are pre-reset.
- Silent column-dropping via provider path mismatch.
LakeFeatureProviderreadfeatures/market=*/timeframe=*/symbol=*.parquet, but the lake stored features under afamily=ta|sppartition — that path never existed, so workflows silently loadedsp_*/ta_*as NaN andDropAllNaNdropped them; models trained on OHLCV only. Smoke test: all-NaN pred before fix, real values after. - Silent NaN-drop during feature regeneration. Regenerating
sp_*features without theharfamily dropped 5 columns (sp_rv1/5/22,sp_vol_ratio_1_22/5_22) from 71 of 72 parquet files. A model trained on 25 features silently became a 20-feature model. - Schema fragmentation. 4 different feature schemas across 72 files (24/53/58/66 columns) — column panels not homogeneous across the lake.
- Stale coverage / truncated feature range.
get_lake_spdefaultedstartto end-minus-30-days: SPY had 2669 bar rows but only 20 feature rows withsp_rv1. - Mid-experiment regeneration. Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files.
- Detection playbook (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines.
Market-structure hypotheses (feed ch. 04/07; from martingale study + clean-data study)
- Submartingale at long horizons, mean-reverting at short horizons. Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation.
- 5-day momentum strongly reverses (pooled regression:
sp_trend_slope_5β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever. - Peso problem in commodities. USO/UNG apparent drift (+0.94/+0.55 ann) is spike-regime compensation, not carry. Trend-follow the spikes, don't hold the reversion stanza.
- HMM regime gating as an overlay, not a feature. Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested.
- Edge is long-short, not long-only (drift is mostly common/market-wide).
Feature methodology hypotheses (feed ch. 04/07)
- Panel width vs feature count: three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM.
- Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model.
sp_ou_zscorewas the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed). - RankIC vs IC vs per-symbol IC are different objects — never mix them (SigAnaRecord vs PortAnaRecord).
- Scale-free features required to survive CSRankNorm; scale-free was necessary but insufficient (moments still regressed).
Model / training hypotheses
- Train/valid RankIC gap as a regime/overfit diagnostic. Proposed bands: ratio <2x underfit, 2–4x healthy, >5x overfitting risk. Hypothesis, untested.
- Sign accuracy, IC hit rate, IC half-life as standard evaluation metrics (bridge from RankIC to traded edge). Proposed, not implemented.
- Equal-weight seed blend > rolling-IC adaptive blending (adaptive weights overfit noise).
- Calibration for rank strategy:
calibrated_pred = pred / Tshrinks prediction spread without changing rankings. Untested.
Strategy / cost hypotheses
- Turnover is the binding constraint (~$60k on $1M over ~7 months at topk10/n_drop2; ~20% daily book turnover). Reductions: n_drop 1 (→ proved on clean data, exp 26), weekly rebalance, no-trade buffer bands, notional-vs-qty orders.
- Kelly sizing is a sizing rule, not a strategy — current equal-weight × risk_degree throws away edge-magnitude information.
- Lower topk increases concentration/drawdown risk — prefer
topk: 20totopk: 5if diversifying. Proposed, untested.
Open questions surfaced by the chats
- OU paradox: why does the strongest single-feature predictor degrade the model?
- Is 5-day reversal a standalone tradable strategy net of costs? (Unisolated.)
- Why does
sp_sharpe_22(M2) improve net IR (0.21→0.62 on clean data) while degrading IC? Mechanism unexplained. - Does the 5-seed ensemble win by variance reduction or by diversification of model families?
- Purged/walk-forward CV instead of single train/valid split — recommended, not implemented.
- Macro/drift overlays (SPY>200d MA regime gate, momentum tilt, macro surprise indices) — proposed; macro needs a new data pipeline.
- PSI-based drift-aware retraining cadence — proposed; rolling retrain exists (exp 27) but no PSI gate.
- Per-symbol calibration of HMM regime posterior — needed before any overlay use.
- Non-overlapping longer horizons (10d/22d labels) to test true trend-following — 5d label can't see 1–12m drift.
Live/ops lessons
- Long MCP runs time out but continue — poll
rd_exp_get_run/rd_exp_list; onlyFINISHEDis final. - Run experiments sequentially, never concurrently (concurrent runs hung for 2h).
- Trace ID ≠ MLflow experiment ID (trace 23 → mlflow exp 25).
trace.sh finishhard-resets the branch and wipes intermediate commits — re-commit after.- Repo and venv copies of custom model code must stay in sync.
- Backtest risk block reports gross equity — a tooling trap; reconcile net separately.
- Model artifact persistence broken on clean runs (no LightGBM booster saved) — fix for inspectability.