book: insert ch01 metrics vocabulary + ch02 research loop; renumber 03/06 — evidence exp 21-31, martingale study
This commit is contained in:
@@ -2,7 +2,7 @@
|
||||
|
||||
Source: opencode chat transcripts under `book/data/chat_mining/` (historical context, pre-clean-lake). Per the evidence contract these are **idea material only** — none may be cited as `PROVEN`. Each idea below is a hypothesis to be tested on the clean lake (exp 21+).
|
||||
|
||||
## Data-quality failure classes (feed ch. 05)
|
||||
## Data-quality failure classes (feed ch. 06)
|
||||
|
||||
These are the *classes* of failure documented across `exp-polluted-lake.txt`, `exp-dirty-lake.txt`, `cleaned-lake.txt`. Durable lessons even though exact numbers are pre-reset.
|
||||
|
||||
@@ -13,7 +13,7 @@ These are the *classes* of failure documented across `exp-polluted-lake.txt`, `e
|
||||
5. **Mid-experiment regeneration.** Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files.
|
||||
6. **Detection playbook** (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines.
|
||||
|
||||
## Market-structure hypotheses (feed ch. 03/06; from martingale study + clean-data study)
|
||||
## Market-structure hypotheses (feed ch. 04/07; from martingale study + clean-data study)
|
||||
|
||||
- **Submartingale at long horizons, mean-reverting at short horizons.** Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation.
|
||||
- **5-day momentum strongly reverses** (pooled regression: `sp_trend_slope_5` β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever.
|
||||
@@ -21,7 +21,7 @@ These are the *classes* of failure documented across `exp-polluted-lake.txt`, `e
|
||||
- **HMM regime gating as an overlay, not a feature.** Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested.
|
||||
- **Edge is long-short, not long-only** (drift is mostly common/market-wide).
|
||||
|
||||
## Feature methodology hypotheses (feed ch. 03/06)
|
||||
## Feature methodology hypotheses (feed ch. 04/07)
|
||||
|
||||
- **Panel width vs feature count:** three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM.
|
||||
- **Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model.** `sp_ou_zscore` was the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed).
|
||||
|
||||
Reference in New Issue
Block a user