Files
tac-exp-dev/queue/README.md
T
zhaoli c0eb65efa7 queue: pre-registered experiment backlog to prove better-trading-performance hypotheses
Mined from book/ on the 'book' branch (HEAD 436692a). 11 queued runs,
each = hypothesis + one-variable change vs the exp-26 reference + acceptance
metric, per the ch.02 isolation/falsification discipline.

- workflows/: 5 runnable config-only YAMLs (Q01 M2 repro, Q02 seed10, Q03 topk20,
  Q04 label10d, Q05 label22d) byte-derived from the exp-26 reference
- designs/: 6 design docs needing custom strategy modules or tool-only A/B
  (Q06 Kelly, Q07 weekly rebalance, Q08 risk-limit A/B, Q09 long-short,
   Q10 HMM overlay, Q11 standalone reversal)
2026-08-19 15:30:46 +00:00

78 lines
6.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TradeAC Experiment Queue — hypotheses that would prove "better trading performance"
**Purpose.** A staging queue of experiment runs, each designed to PROVE (or
REFUTE) one hypothesis about how to achieve better trading performance on the
TradeAC stack. Every item is pre-registered: hypothesis, change-vs-reference,
and acceptance metric are fixed BEFORE the run (book ch.02 isolation + falsification
discipline). Nothing here is executed yet — each entry carries its execution
command and can be run by tracing first (`rd_trace_start` → `rd_run_workflow` /
`rd_risk_calibrate` → `rd_trace_finish`).
**Source.** Mined from the `book` branch of this repo (`book/CLAIMS.md`,
`book/EVIDENCE.md`, `book/chapters/*`, `book/references/chat-ideas.md`). Only
clean-lake (exp 21+) facts are cited as reference numbers; pre-clean-lake claims
are idea material that the queue is designed to test.
## Reference / control (MUST reproduce first)
The exp-26 reference — the campaign's best clean-lake result (EVIDENCE#015, run
`21afc6af…`, mlflow exp 25, branch `exp/26-test-whether-reducing-topkdropout-daily`):
| Config element | Reference value |
|---|---|
| Universe | 50-ETF panel (same `UNIVERSE` list as exp-24/26) |
| Features | compact stochastic: `$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead` |
| Label | `Ref($close,-6)/Ref($close,-1)-1` (5d) |
| Model | `RankICEnsembleLGBModel` (tac_qlib.contrib.model.rank_ensemble), seeds `42,7,2026,99,123`, lr 0.02, num_leaves 31, 3000 rounds, early_stop 200, min_data_in_leaf 20, lambda_l2 0.5, colsample/subsample 0.8 |
| Train / valid / test | 2016-01-04..2025-09-01 / 2025-09-03..2026-01-03 / 2026-01-04..2026-08-10 |
| Strategy | TopkDropout, topk 10, n_drop 1, risk_degree 0.95 |
| Costs / benchmark | open 0.0005 / close 0.0015 / min $5, deal $close, SPY, $1M |
Reference metrics to beat (EVIDENCE#015): **net_ann_return +2.13%, net_IR 0.21,
gross +7.02%, net_max_drawdown −7.69%, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54.**
## The queue (ordered by value × feasibility)
| ID | Title / hypothesis | Change vs reference (ONE var) | Acceptance | Config | Ready? |
|----|--------------------|-------------------------------|------------|--------|--------|
| Q01 | **M2 Sharpe-drift reproduction** — adding `sp_sharpe_22` (risk-adjusted 22d drift) improves net perf (exp 30: +6.53% IR 0.62, unreproduced → promote HYPOTHESIS) | +`sp_sharpe_22` to features | net_IR > 0.21, net_ann > +2.13% | `workflows/q01_m2_sharpe22_repro.yaml` | ✅ |
| Q02 | **Seed count 10 vs 5** — more seeds → higher ICIR/net; tests whether averaging saturates (exp 28 proved 5>2) | seeds → 10 | net_IR ≥ 0.21, ICIR/RankICIR ≥ ref | `workflows/q02_seed10.yaml` | ✅ |
| Q03 | **topk 20 diversification** — effective book is ~4 independent names; wider book cuts drawdown without hurting weak signal | topk 10→20 | net_IR > 0.21, MDD < 7.69%, net_ann ≥ +2.13% | `workflows/q03_topk20.yaml` | ✅ |
| Q04 | **10-day non-overlapping label** — longer horizon captures trend/reversal 5d blurs, lowers churn | label → 10d | net_IR > 0.21, net_ann > +2.13% | `workflows/q04_label10d.yaml` | ✅ |
| Q05 | **22-day label** — true trend-following; 5d can't see 1–12m drift (submartingale) | label → 22d | net_IR > 0.21, net_ann > +2.13%, cost ≤ ref | `workflows/q05_label22d.yaml` | ✅ |
| Q06 | **Fractional-Kelly sizing** (re-run exp 15 on clean lake) — sizing by edge magnitude beats equal-weight net of costs | custom strategy (sizing) | net_IR > 0.21, net_ann > +2.13%, cost ≤ ref | `designs/q06_kelly_sizing.md` | ⚠️ needs `kelly_dropout.py` |
| Q07 | **Weekly rebalance** — next turnover lever after n_drop 1; cut forced churn at same signal | custom strategy (weekly) | cost/turnover ↓ AND net_IR > 0.21, net_ann > +2.13% | `designs/q07_weekly_rebalance.md` | ⚠️ needs `weekly_rebalance.py` |
| Q08 | **Risk-limit re-validation** — $5M liquidity floor improves net IR / cuts DD on post-reset signal (exp 18 pre-clean-lake) | `rd_risk_calibrate` A/B on exp-26 pred | net_IR > 0.21, MDD < 7.69% vs no-limit | `designs/q08_risk_limit_ab.md` | ✅ tool-only |
| Q09 | **Long-short construction** — the L/S edge (Sharpe 4.54) realizes more net of costs than long-only | custom strategy (top+bottom) | net_IR > 0.21, net_ann > +2.13%, cost ≤ 2× ref | `designs/q09_long_short.md` | ⚠️ needs `top_bottom.py` |
| Q10 | **HMM regime overlay** — regime as overlay (not feature) cuts drawdown; exp 25 proved features fail, overlay untested | custom strategy (regime gate) | MDD < 7.69%, net_IR ≥ 0.21 | `designs/q10_hmm_regime_overlay.md` | ⚠️ needs `regime_gate.py` + `get_lake_sp` |
| Q11 | **Standalone 5-day reversal** — reversal (β −0.53, t −24) tradable net of 20bp round-trip; unisolated | single-feature model/backtest | net_ann > 0 standalone | `designs/q11_standalone_reversal.md` | ⚠️ partial |
### Deferred (methodology / infra, P3)
- **Q12 Purged / walk-forward CV** on the exp-26 reference (book ch.02 open
question) — methodology improvement, not a direct alpha lever.
- **Q13 Out-of-universe validation** — non-ETF universe for the compact
stochastic feature set (book README open question; needs new lake symbols).
## Execution protocol (per queued run)
1. **Validate the lake first** (`validate_lake_dataset` + `rd_status`) — the
clean-lake lesson: silent NaN-drops and hollow coverage invalidate a run.
2. **Trace before running** (`rd_trace_start` with the hypothesis as `rational`,
`evolved_from=auto` for lineage → it will fork from the closest prior
experiment). Use a FRESH experiment name per run, e.g. `tac-rd-q01-m2-...`.
3. **Run** `rd_run_workflow config_path=<abs path to the queue YAML>
experiment_name=<fresh name>` — use `wait=false`, poll `rd_exp_get_run` until
`FINISHED` (4-year trains outlive the MCP call).
4. **Verify against acceptance** via `rd_exp_result` (headline + backtest risk).
5. **Finish the trace** (`rd_trace_finish` with `metrics` + `evaluation`),
snapshot any new custom modules (`rd_trace_snapshot`).
6. **Report to the book** — on PROVE, update `book/CLAIMS.md`/`EVIDENCE.md`;
on REFUTE, record the negative (falsification is the output).
Sequential execution only (concurrent runs hang — chat-ideas.md ops lesson).
## Provenance
Mined 2026-08-19 from `book/` on the `book` branch (HEAD `436692a`). Reference
config reproduced byte-for-byte from the exp-26 run artifact config
(`/home/data/lake/mlruns/25/21afc6afdb674a399b59dd76c97628ce/artifacts/config`).