Files
book-tac/book/references/chat-ideas.md
T

64 lines
6.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chat-Mined Ideas & Hypotheses
Source: opencode chat transcripts under `book/data/chat_mining/` (historical context, pre-clean-lake). Per the evidence contract these are **idea material only** — none may be cited as `PROVEN`. Each idea below is a hypothesis to be tested on the clean lake (exp 21+).
## Data-quality failure classes (feed ch. 05)
These are the *classes* of failure documented across `exp-polluted-lake.txt`, `exp-dirty-lake.txt`, `cleaned-lake.txt`. Durable lessons even though exact numbers are pre-reset.
1. **Silent column-dropping via provider path mismatch.** `LakeFeatureProvider` read `features/market=*/timeframe=*/symbol=*.parquet`, but the lake stored features under a `family=ta|sp` partition — that path never existed, so workflows silently loaded `sp_*`/`ta_*` as NaN and `DropAllNaN` dropped them; models trained on OHLCV only. Smoke test: all-NaN pred before fix, real values after.
2. **Silent NaN-drop during feature regeneration.** Regenerating `sp_*` features without the `har` family dropped 5 columns (`sp_rv1/5/22`, `sp_vol_ratio_1_22/5_22`) from 71 of 72 parquet files. A model trained on 25 features silently became a 20-feature model.
3. **Schema fragmentation.** 4 different feature schemas across 72 files (24/53/58/66 columns) — column panels not homogeneous across the lake.
4. **Stale coverage / truncated feature range.** `get_lake_sp` defaulted `start` to end-minus-30-days: SPY had 2669 bar rows but only 20 feature rows with `sp_rv1`.
5. **Mid-experiment regeneration.** Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files.
6. **Detection playbook** (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines.
## Market-structure hypotheses (feed ch. 03/06; from martingale study + clean-data study)
- **Submartingale at long horizons, mean-reverting at short horizons.** Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation.
- **5-day momentum strongly reverses** (pooled regression: `sp_trend_slope_5` β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever.
- **Peso problem in commodities.** USO/UNG apparent drift (+0.94/+0.55 ann) is spike-regime compensation, not carry. Trend-follow the spikes, don't hold the reversion stanza.
- **HMM regime gating as an overlay, not a feature.** Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested.
- **Edge is long-short, not long-only** (drift is mostly common/market-wide).
## Feature methodology hypotheses (feed ch. 03/06)
- **Panel width vs feature count:** three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM.
- **Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model.** `sp_ou_zscore` was the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed).
- **RankIC vs IC vs per-symbol IC are different objects** — never mix them (SigAnaRecord vs PortAnaRecord).
- **Scale-free features required** to survive CSRankNorm; scale-free was necessary but insufficient (moments still regressed).
## Model / training hypotheses
- **Train/valid RankIC gap as a regime/overfit diagnostic.** Proposed bands: ratio <2x underfit, 2–4x healthy, >5x overfitting risk. Hypothesis, untested.
- **Sign accuracy, IC hit rate, IC half-life** as standard evaluation metrics (bridge from RankIC to traded edge). Proposed, not implemented.
- **Equal-weight seed blend > rolling-IC adaptive blending** (adaptive weights overfit noise).
- **Calibration for rank strategy:** `calibrated_pred = pred / T` shrinks prediction spread without changing rankings. Untested.
## Strategy / cost hypotheses
- **Turnover is the binding constraint** (~$60k on $1M over ~7 months at topk10/n_drop2; ~20% daily book turnover). Reductions: n_drop 1 (→ proved on clean data, exp 26), weekly rebalance, no-trade buffer bands, notional-vs-qty orders.
- **Kelly sizing is a sizing rule, not a strategy** — current equal-weight × risk_degree throws away edge-magnitude information.
- **Lower topk increases concentration/drawdown risk** — prefer `topk: 20` to `topk: 5` if diversifying. Proposed, untested.
## Open questions surfaced by the chats
- OU paradox: why does the strongest single-feature predictor degrade the model?
- Is 5-day reversal a standalone tradable strategy net of costs? (Unisolated.)
- Why does `sp_sharpe_22` (M2) improve net IR (0.21→0.62 on clean data) while degrading IC? Mechanism unexplained.
- Does the 5-seed ensemble win by variance reduction or by diversification of model families?
- Purged/walk-forward CV instead of single train/valid split — recommended, not implemented.
- Macro/drift overlays (SPY>200d MA regime gate, momentum tilt, macro surprise indices) — proposed; macro needs a new data pipeline.
- PSI-based drift-aware retraining cadence — proposed; rolling retrain exists (exp 27) but no PSI gate.
- Per-symbol calibration of HMM regime posterior — needed before any overlay use.
- Non-overlapping longer horizons (10d/22d labels) to test true trend-following — 5d label can't see 1–12m drift.
## Live/ops lessons
- Long MCP runs time out but continue — poll `rd_exp_get_run`/`rd_exp_list`; only `FINISHED` is final.
- Run experiments sequentially, never concurrently (concurrent runs hung for 2h).
- Trace ID ≠ MLflow experiment ID (trace 23 → mlflow exp 25).
- `trace.sh finish` hard-resets the branch and wipes intermediate commits — re-commit after.
- Repo and venv copies of custom model code must stay in sync.
- Backtest risk block reports gross equity — a tooling trap; reconcile net separately.
- Model artifact persistence broken on clean runs (no LightGBM booster saved) — fix for inspectability.