book: insert ch01 metrics vocabulary + ch02 research loop; renumber 03/06 — evidence exp 21-31, martingale study

This commit is contained in:
TradeAC Book Agent
2026-08-18 23:04:10 +00:00
parent 3562b7f776
commit 436692a620
7 changed files with 186 additions and 34 deletions
+3 -3
View File
@@ -2,7 +2,7 @@
Source: opencode chat transcripts under `book/data/chat_mining/` (historical context, pre-clean-lake). Per the evidence contract these are **idea material only** — none may be cited as `PROVEN`. Each idea below is a hypothesis to be tested on the clean lake (exp 21+).
## Data-quality failure classes (feed ch. 05)
## Data-quality failure classes (feed ch. 06)
These are the *classes* of failure documented across `exp-polluted-lake.txt`, `exp-dirty-lake.txt`, `cleaned-lake.txt`. Durable lessons even though exact numbers are pre-reset.
@@ -13,7 +13,7 @@ These are the *classes* of failure documented across `exp-polluted-lake.txt`, `e
5. **Mid-experiment regeneration.** Feature parquet mtimes showed regeneration at 00:56 and 02:50 (Aug 17) — after exp-18 but before R0 — so reference and R0 ran on different feature files.
6. **Detection playbook** (the valuable part): byte-identical-config reproduction; prediction-distribution comparison (pred_std, rank correlation, top-10 overlap); null-baseline IC z-scores (daily RankIC null std = 1/√(N−1) ≈ 0.143 for 50 names); per-day IC outlier fingerprints (3–4σ single-day ICs are contamination, not signal); feature-vs-bar alignment checks; file-mtime forensics; same-environment baselines.
## Market-structure hypotheses (feed ch. 03/06; from martingale study + clean-data study)
## Market-structure hypotheses (feed ch. 04/07; from martingale study + clean-data study)
- **Submartingale at long horizons, mean-reverting at short horizons.** Drift compounds but explains ~0.5% of daily variance; short-horizon reversal (VR<1 at 5–20d for ~32/72 assets) is the tradable deviation.
- **5-day momentum strongly reverses** (pooled regression: `sp_trend_slope_5` β = −0.53, t = −24). Fade 5-day strength; the repo's 5-day label is the best IC lever.
@@ -21,7 +21,7 @@ These are the *classes* of failure documented across `exp-polluted-lake.txt`, `e
- **HMM regime gating as an overlay, not a feature.** Regime flags failed as model features (exp 9, exp 25) but the long-only/regime-gate overlay idea survives untested.
- **Edge is long-short, not long-only** (drift is mostly common/market-wide).
## Feature methodology hypotheses (feed ch. 03/06)
## Feature methodology hypotheses (feed ch. 04/07)
- **Panel width vs feature count:** three independent feature expansions (ou/hmm, realized moments, TA) regressed; the minimal generic set won repeatedly. Hypothesis: on ~50-name daily panels, cross-sectional features dilute CSRankNorm+LGBM.
- **Single-feature time-series IC ≠ marginal contribution in a cross-sectional rank model.** `sp_ou_zscore` was the strongest stable single-feature predictor (IC −0.15/−0.13) yet hurt the model (IC 0.051→0.034). Measurement mismatch unresolved. TODO(evidence-needed).