Files

6.3 KiB
Raw Permalink Blame History

Chat-Mined Ideas & Hypotheses

Source: opencode chat transcripts under book/data/chat_mining/ (historical context, pre-clean-lake). Per the evidence contract these are idea material only — none may be cited as PROVEN. Each idea below is a hypothesis to be tested on the clean lake (exp 21+).

Data-quality failure classes (feed ch. 05)

These are the classes of failure documented across exp-polluted-lake.txt, exp-dirty-lake.txt, cleaned-lake.txt. Durable lessons even though exact numbers are pre-reset.

  1. Silent column-dropping via provider path mismatch. LakeFeatureProvider read features/market=*/timeframe=*/symbol=*.parquet, but the lake stored features under a family=ta|sp partition — that path never existed, so workflows silently loaded sp_*/ta_* as NaN and DropAllNaN dropped them; models trained on OHLCV only. Smoke test: all-NaN pred before fix, real values after.
  2. Silent NaN-drop during feature regeneration. Regenerating sp_* features without the har family dropped 5 columns (sp_rv1/5/22, sp_vol_ratio_1_22/5_22) from 71 of 72 parquet files. A model trained on 25 features silently became a 20-feature model.
  3. Schema fragmentation. 4 different feature schemas across 72 files (24/53/58/66 columns) — column panels not homogeneous across the lake.
  4. Stale coverage / truncated feature range. get_lake_sp defaulted start to end-minus-30-days: SPY had 2669 bar rows but only 20 feature rows with sp_rv1.
  5. Mid-experiment regeneration. Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files.
  6. Detection playbook (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines.

Market-structure hypotheses (feed ch. 03/06; from martingale study + clean-data study)

  • Submartingale at long horizons, mean-reverting at short horizons. Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation.
  • 5-day momentum strongly reverses (pooled regression: sp_trend_slope_5 β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever.
  • Peso problem in commodities. USO/UNG apparent drift (+0.94/+0.55 ann) is spike-regime compensation, not carry. Trend-follow the spikes, don't hold the reversion stanza.
  • HMM regime gating as an overlay, not a feature. Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested.
  • Edge is long-short, not long-only (drift is mostly common/market-wide).

Feature methodology hypotheses (feed ch. 03/06)

  • Panel width vs feature count: three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM.
  • Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model. sp_ou_zscore was the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed).
  • RankIC vs IC vs per-symbol IC are different objects — never mix them (SigAnaRecord vs PortAnaRecord).
  • Scale-free features required to survive CSRankNorm; scale-free was necessary but insufficient (moments still regressed).

Model / training hypotheses

  • Train/valid RankIC gap as a regime/overfit diagnostic. Proposed bands: ratio <2x underfit, 2–4x healthy, >5x overfitting risk. Hypothesis, untested.
  • Sign accuracy, IC hit rate, IC half-life as standard evaluation metrics (bridge from RankIC to traded edge). Proposed, not implemented.
  • Equal-weight seed blend > rolling-IC adaptive blending (adaptive weights overfit noise).
  • Calibration for rank strategy: calibrated_pred = pred / T shrinks prediction spread without changing rankings. Untested.

Strategy / cost hypotheses

  • Turnover is the binding constraint (~$60k on $1M over ~7 months at topk10/n_drop2; ~20% daily book turnover). Reductions: n_drop 1 (→ proved on clean data, exp 26), weekly rebalance, no-trade buffer bands, notional-vs-qty orders.
  • Kelly sizing is a sizing rule, not a strategy — current equal-weight × risk_degree throws away edge-magnitude information.
  • Lower topk increases concentration/drawdown risk — prefer topk: 20 to topk: 5 if diversifying. Proposed, untested.

Open questions surfaced by the chats

  • OU paradox: why does the strongest single-feature predictor degrade the model?
  • Is 5-day reversal a standalone tradable strategy net of costs? (Unisolated.)
  • Why does sp_sharpe_22 (M2) improve net IR (0.21→0.62 on clean data) while degrading IC? Mechanism unexplained.
  • Does the 5-seed ensemble win by variance reduction or by diversification of model families?
  • Purged/walk-forward CV instead of single train/valid split — recommended, not implemented.
  • Macro/drift overlays (SPY>200d MA regime gate, momentum tilt, macro surprise indices) — proposed; macro needs a new data pipeline.
  • PSI-based drift-aware retraining cadence — proposed; rolling retrain exists (exp 27) but no PSI gate.
  • Per-symbol calibration of HMM regime posterior — needed before any overlay use.
  • Non-overlapping longer horizons (10d/22d labels) to test true trend-following — 5d label can't see 1–12m drift.

Live/ops lessons

  • Long MCP runs time out but continue — poll rd_exp_get_run/rd_exp_list; only FINISHED is final.
  • Run experiments sequentially, never concurrently (concurrent runs hung for 2h).
  • Trace ID ≠ MLflow experiment ID (trace 23 → mlflow exp 25).
  • trace.sh finish hard-resets the branch and wipes intermediate commits — re-commit after.
  • Repo and venv copies of custom model code must stay in sync.
  • Backtest risk block reports gross equity — a tooling trap; reconcile net separately.
  • Model artifact persistence broken on clean runs (no LightGBM booster saved) — fix for inspectability.