queue: pre-register Series 2 (Q12-Q20) targeting unproven book hypotheses

Series 1 (Q01-Q11, exp 33-43) executed and folded into book/CLAIMS.md.
Series 2 covers the remaining HYPOTHESIS rows and open questions:
  Q12 22d label + weekly recompute (untested combo)
  Q13 weekly rebalance reproduction on a 2nd OOS window
  Q14 out-of-universe validation (single-stock panel, needs lake backfill)
  Q15 5-seed vs single-seed clean A/B
  Q16 hmm family as features
  Q17 realized-moments family as features
  Q18 OptimalStopControl clean re-test
  Q19/Q20 martingale-VR + effective-names scripted studies

Each workflow pins one-variable change vs exp-26 reference and acceptance.
This commit is contained in:
zhaoli
2026-08-20 03:51:58 +00:00
parent c0eb65efa7
commit 7124ef5d8e
17 changed files with 502 additions and 627 deletions
+47 -56
View File
@@ -1,78 +1,69 @@
# TradeAC Experiment Queue — hypotheses that would prove "better trading performance"
# TradeAC Experiment Queue — Series 2 (Q12+)
**Purpose.** A staging queue of experiment runs, each designed to PROVE (or
REFUTE) one hypothesis about how to achieve better trading performance on the
TradeAC stack. Every item is pre-registered: hypothesis, change-vs-reference,
and acceptance metric are fixed BEFORE the run (book ch.02 isolation + falsification
discipline). Nothing here is executed yet — each entry carries its execution
command and can be run by tracing first (`rd_trace_start` → `rd_run_workflow` /
`rd_risk_calibrate` → `rd_trace_finish`).
**Purpose.** The next pre-registered batch of experiments, continuing Series 1
(Q01–Q11, exp 33–43, all executed and folded into `book/CLAIMS.md` /
`book/EVIDENCE.md`). Each entry targets a still-unproven `HYPOTHESIS` from the
book or an open question flagged in `CLAIMS.md`/`book/README.md`, and follows the
Series-1 discipline: one variable changed vs the exp-26 reference, acceptance
fixed BEFORE the run, sequential execution, trace-first, verify-then-close.
**Source.** Mined from the `book` branch of this repo (`book/CLAIMS.md`,
`book/EVIDENCE.md`, `book/chapters/*`, `book/references/chat-ideas.md`). Only
clean-lake (exp 21+) facts are cited as reference numbers; pre-clean-lake claims
are idea material that the queue is designed to test.
**Reference / control (MUST reproduce first).** exp 26 (`21afc6af…`, mlflow exp
25) is the campaign baseline; exp 39 (Q07, weekly rebalance) is the best
construction. Reference config is byte-reproduced in `workflows/exp26/` on the
`exp/26-…` branch and in this dir's `workflows/*.yaml`.
## Reference / control (MUST reproduce first)
The exp-26 reference — the campaign's best clean-lake result (EVIDENCE#015, run
`21afc6af…`, mlflow exp 25, branch `exp/26-test-whether-reducing-topkdropout-daily`):
| Config element | Reference value |
| Config element | exp-26 reference value |
|---|---|
| Universe | 50-ETF panel (same `UNIVERSE` list as exp-24/26) |
| Features | compact stochastic: `$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead` |
| Universe | 50-ETF panel (`UNIVERSE` below) |
| Features | compact stochastic 25-field set (no ou/hmm/moments/garch) |
| Label | `Ref($close,-6)/Ref($close,-1)-1` (5d) |
| Model | `RankICEnsembleLGBModel` (tac_qlib.contrib.model.rank_ensemble), seeds `42,7,2026,99,123`, lr 0.02, num_leaves 31, 3000 rounds, early_stop 200, min_data_in_leaf 20, lambda_l2 0.5, colsample/subsample 0.8 |
| Train / valid / test | 2016-01-04..2025-09-01 / 2025-09-03..2026-01-03 / 2026-01-04..2026-08-10 |
| Model | `RankICEnsembleLGBModel`, seeds `42,7,2026,99,123`, lr 0.02, leaves 31, 3000 rounds, ES 200 |
| Segments | train 2016-01-04..2025-09-01 / valid 2025-09-03..2026-01-03 / test 2026-01-04..2026-08-10 |
| Strategy | TopkDropout, topk 10, n_drop 1, risk_degree 0.95 |
| Costs / benchmark | open 0.0005 / close 0.0015 / min $5, deal $close, SPY, $1M |
| Costs | open 0.0005 / close 0.0015 / min $5, deal $close, SPY benchmark, $1M |
Reference metrics to beat (EVIDENCE#015): **net_ann_return +2.13%, net_IR 0.21,
gross +7.02%, net_max_drawdown −7.69%, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54.**
**Reference metrics to beat (EVIDENCE#015):** net_ann +2.13%, net_IR 0.21, gross
+7.02%, maxDD −7.69%, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54. Weekly
(Q07, EVIDENCE#028): net +12.51%, IR 1.24, maxDD −4.13%, ~1.1pp cost drag.
## The queue (ordered by value × feasibility)
| ID | Title / hypothesis | Change vs reference (ONE var) | Acceptance | Config | Ready? |
|----|--------------------|-------------------------------|------------|--------|--------|
| Q01 | **M2 Sharpe-drift reproduction** — adding `sp_sharpe_22` (risk-adjusted 22d drift) improves net perf (exp 30: +6.53% IR 0.62, unreproduced → promote HYPOTHESIS) | +`sp_sharpe_22` to features | net_IR > 0.21, net_ann > +2.13% | `workflows/q01_m2_sharpe22_repro.yaml` | ✅ |
| Q02 | **Seed count 10 vs 5** — more seeds → higher ICIR/net; tests whether averaging saturates (exp 28 proved 5>2) | seeds → 10 | net_IR ≥ 0.21, ICIR/RankICIR ≥ ref | `workflows/q02_seed10.yaml` | ✅ |
| Q03 | **topk 20 diversification** — effective book is ~4 independent names; wider book cuts drawdown without hurting weak signal | topk 10→20 | net_IR > 0.21, MDD < 7.69%, net_ann ≥ +2.13% | `workflows/q03_topk20.yaml` | ✅ |
| Q04 | **10-day non-overlapping label** — longer horizon captures trend/reversal 5d blurs, lowers churn | label → 10d | net_IR > 0.21, net_ann > +2.13% | `workflows/q04_label10d.yaml` | ✅ |
| Q05 | **22-day label** — true trend-following; 5d can't see 1–12m drift (submartingale) | label → 22d | net_IR > 0.21, net_ann > +2.13%, cost ≤ ref | `workflows/q05_label22d.yaml` | ✅ |
| Q06 | **Fractional-Kelly sizing** (re-run exp 15 on clean lake) — sizing by edge magnitude beats equal-weight net of costs | custom strategy (sizing) | net_IR > 0.21, net_ann > +2.13%, cost ≤ ref | `designs/q06_kelly_sizing.md` | ⚠️ needs `kelly_dropout.py` |
| Q07 | **Weekly rebalance** — next turnover lever after n_drop 1; cut forced churn at same signal | custom strategy (weekly) | cost/turnover ↓ AND net_IR > 0.21, net_ann > +2.13% | `designs/q07_weekly_rebalance.md` | ⚠️ needs `weekly_rebalance.py` |
| Q08 | **Risk-limit re-validation** — $5M liquidity floor improves net IR / cuts DD on post-reset signal (exp 18 pre-clean-lake) | `rd_risk_calibrate` A/B on exp-26 pred | net_IR > 0.21, MDD < 7.69% vs no-limit | `designs/q08_risk_limit_ab.md` | ✅ tool-only |
| Q09 | **Long-short construction** — the L/S edge (Sharpe 4.54) realizes more net of costs than long-only | custom strategy (top+bottom) | net_IR > 0.21, net_ann > +2.13%, cost ≤ 2× ref | `designs/q09_long_short.md` | ⚠️ needs `top_bottom.py` |
| Q10 | **HMM regime overlay** — regime as overlay (not feature) cuts drawdown; exp 25 proved features fail, overlay untested | custom strategy (regime gate) | MDD < 7.69%, net_IR ≥ 0.21 | `designs/q10_hmm_regime_overlay.md` | ⚠️ needs `regime_gate.py` + `get_lake_sp` |
| Q11 | **Standalone 5-day reversal** — reversal (β −0.53, t −24) tradable net of 20bp round-trip; unisolated | single-feature model/backtest | net_ann > 0 standalone | `designs/q11_standalone_reversal.md` | ⚠️ partial |
| Q12 | **22d label + weekly recompute** — the untested combo: Q05's label edge (IC 0.097, RankIC 0.117) with Q07's cost relief | label → 22d AND strategy → weekly (two coupled, explicitly pre-registered) | net_IR > 0.5, net_ann > +5%, cost drag ≤ 2pp | `workflows/q12_label22d_weekly.yaml` | ✅ |
| Q13 | **Weekly rebalance reproduction on a 2nd window** — Q07 was a single OOS window; reproduce on test 2025-01-02..2025-12-31 before promoting to a live round | segments only (shifted) | net_IR > 0.21, net_ann > +2.13% on the new window | `workflows/q13_weekly_second_window.yaml` | ✅ |
| Q14 | **Out-of-universe validation** — compact stochastic set generalizes off the 50-ETF panel to a single-stock universe | universe → 30 liquid single names | RankIC > 0.03, ICIR > 0.15, net IR > 0 on stocks | `workflows/q14_out_of_universe.yaml` | ⚠️ needs stock-lake backfill (see design) |
| Q15 | **5-seed vs single-model clean A/B** — seed-count claim (exp 12 idea, re-validated exp 22–24, never a clean A/B) | seeds → 1 (`2026`) | single-model RankIC/IR < 5-seed ref; net_IR ≥ 0.21 acceptable if ≥ single | `workflows/q15_single_seed.yaml` | ✅ |
| Q16 | **HMM family added as features** — settles "dropping model-specific (ou,hmm) improves signal" (exp 25 tested OU; hmm-as-feature untested) | features += `sp_hmm_p_regime1,sp_hmm_state` | no improvement: RankIC ≤ 0.0663, net_IR ≤ 0.21 | `workflows/q16_hmm_features.yaml` | ✅ |
| Q17 | **Realized-moments family added** — settles "moment/volatility families regress" (exp 11 idea, never clean A/B) | features += `sp_rskew_5,sp_rskew_22,sp_rkurt_5,sp_rkurt_22,sp_dsv_5,sp_dsv_22` | no improvement: RankIC ≤ 0.0663, net_IR ≤ 0.21 | `workflows/q17_moments_features.yaml` | ✅ |
| Q18 | **OptimalStopControl clean re-test** — exp 13/14 claim (TopkDropout > stop-control) never re-tested post-reset | strategy → `OptimalStopControl` (exp-13 params) | TopkDropout net_IR ≥ stop-control net_IR; document cost drag | `workflows/q18_optstop.yaml` | ✅ (module verified in venv) |
| Q19 | **Martingale / variance-ratio study close-out** — exp 19 never closed; VR<1 at 5–20d on clean lake | ad-hoc script (no qrun) | VR stats + drift decomposition on 50-ETF panel | `designs/q19_martingale_vr.md` | ✅ script |
| Q20 | **Effective independent names (≈4)** — eigenvalue analysis on clean-lake covariance | ad-hoc script | eigenvalue spectrum + effective-rank count | `designs/q20_effective_names.md` | ✅ script |
### Deferred (methodology / infra, P3)
- **Q12 Purged / walk-forward CV** on the exp-26 reference (book ch.02 open
question) — methodology improvement, not a direct alpha lever.
- **Q13 Out-of-universe validation** — non-ETF universe for the compact
stochastic feature set (book README open question; needs new lake symbols).
- Purged / walk-forward CV (was queue's old Q12) — methodology, not an alpha lever.
- PSI-based drift-aware retraining cadence — needs a drift-gate module + a retrain decision rule.
- No-trade buffer band / notional-vs-qty sizing — siblings of Q12/Q13; queue only if weekly reproduces.
- Macro/drift overlays (SPY>200d regime gate, momentum tilt) — needs new data pipeline.
## Execution protocol (per queued run)
1. **Validate the lake first** (`validate_lake_dataset` + `rd_status`) — the
clean-lake lesson: silent NaN-drops and hollow coverage invalidate a run.
2. **Trace before running** (`rd_trace_start` with the hypothesis as `rational`,
`evolved_from=auto` for lineage → it will fork from the closest prior
experiment). Use a FRESH experiment name per run, e.g. `tac-rd-q01-m2-...`.
3. **Run** `rd_run_workflow config_path=<abs path to the queue YAML>
experiment_name=<fresh name>` — use `wait=false`, poll `rd_exp_get_run` until
`FINISHED` (4-year trains outlive the MCP call).
1. **Validate the lake first** (`validate_lake_dataset` + `rd_status`) — clean-lake lesson: silent NaN-drops and hollow coverage invalidate a run. Q14 additionally requires backfilling the single-stock universe (bars + sp/ta features, full range, explicit `start`/`end`).
2. **Trace before running** (`rd_trace_start` with the hypothesis as `rational`, fresh `experiment_name`, `evolved_from=auto`).
3. **Run** `rd_run_workflow config_path=<abs path to the queue YAML> experiment_name=<fresh name>` — `wait=false`, poll `rd_exp_get_run` until `FINISHED`.
4. **Verify against acceptance** via `rd_exp_result` (headline + backtest risk).
5. **Finish the trace** (`rd_trace_finish` with `metrics` + `evaluation`),
snapshot any new custom modules (`rd_trace_snapshot`).
6. **Report to the book** — on PROVE, update `book/CLAIMS.md`/`EVIDENCE.md`;
on REFUTE, record the negative (falsification is the output).
5. **Finish the trace** (`rd_trace_finish` with `metrics` + `evaluation`), snapshot any changed contrib modules.
6. **Report to the book** — PROVE/REFUTE → update `book/CLAIMS.md` + `book/EVIDENCE.md`.
Sequential execution only (concurrent runs hang — chat-ideas.md ops lesson).
Sequential execution only (concurrent runs hang — chat-ideas.md ops lesson). Any
custom strategy/module changed here must be copied into the venv site-packages
snapshot before `rd_run_workflow` can import it (see `/app/AGENTS.md`). As of
2026-08-20 `WeeklyRebalanceDropoutStrategy` and `OptimalStopControl` are verified
in sync with the venv snapshot; the lake already persists the `sp_hmm_*` and
`sp_moments` families on the 50-ETF panel.
## Provenance
Mined 2026-08-19 from `book/` on the `book` branch (HEAD `436692a`). Reference
config reproduced byte-for-byte from the exp-26 run artifact config
(`/home/data/lake/mlruns/25/21afc6afdb674a399b59dd76c97628ce/artifacts/config`).
Mined 2026-08-20 from `book/CLAIMS.md`, `book/EVIDENCE.md`, `book/README.md`,
`book/references/chat-ideas.md`, and Series-1 `queue/` (Q01–Q11, executed exp
33–43). Reference numbers are post-clean-lake (exp 21+).