Files
tac-exp-dev/queue
zhaoli c0eb65efa7 queue: pre-registered experiment backlog to prove better-trading-performance hypotheses
Mined from book/ on the 'book' branch (HEAD 436692a). 11 queued runs,
each = hypothesis + one-variable change vs the exp-26 reference + acceptance
metric, per the ch.02 isolation/falsification discipline.

- workflows/: 5 runnable config-only YAMLs (Q01 M2 repro, Q02 seed10, Q03 topk20,
  Q04 label10d, Q05 label22d) byte-derived from the exp-26 reference
- designs/: 6 design docs needing custom strategy modules or tool-only A/B
  (Q06 Kelly, Q07 weekly rebalance, Q08 risk-limit A/B, Q09 long-short,
   Q10 HMM overlay, Q11 standalone reversal)
2026-08-19 15:30:46 +00:00
..

TradeAC Experiment Queue — hypotheses that would prove "better trading performance"

Purpose. A staging queue of experiment runs, each designed to PROVE (or REFUTE) one hypothesis about how to achieve better trading performance on the TradeAC stack. Every item is pre-registered: hypothesis, change-vs-reference, and acceptance metric are fixed BEFORE the run (book ch.02 isolation + falsification discipline). Nothing here is executed yet — each entry carries its execution command and can be run by tracing first (rd_trace_start → rd_run_workflow / rd_risk_calibrate → rd_trace_finish).

Source. Mined from the book branch of this repo (book/CLAIMS.md, book/EVIDENCE.md, book/chapters/*, book/references/chat-ideas.md). Only clean-lake (exp 21+) facts are cited as reference numbers; pre-clean-lake claims are idea material that the queue is designed to test.

Reference / control (MUST reproduce first)

The exp-26 reference — the campaign's best clean-lake result (EVIDENCE#015, run 21afc6af…, mlflow exp 25, branch exp/26-test-whether-reducing-topkdropout-daily):

Config element Reference value
Universe 50-ETF panel (same UNIVERSE list as exp-24/26)
Features compact stochastic: $open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead
Label Ref($close,-6)/Ref($close,-1)-1 (5d)
Model RankICEnsembleLGBModel (tac_qlib.contrib.model.rank_ensemble), seeds 42,7,2026,99,123, lr 0.02, num_leaves 31, 3000 rounds, early_stop 200, min_data_in_leaf 20, lambda_l2 0.5, colsample/subsample 0.8
Train / valid / test 2016-01-04..2025-09-01 / 2025-09-03..2026-01-03 / 2026-01-04..2026-08-10
Strategy TopkDropout, topk 10, n_drop 1, risk_degree 0.95
Costs / benchmark open 0.0005 / close 0.0015 / min $5, deal $close, SPY, $1M

Reference metrics to beat (EVIDENCE#015): net_ann_return +2.13%, net_IR 0.21, gross +7.02%, net_max_drawdown −7.69%, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54.

The queue (ordered by value × feasibility)

ID Title / hypothesis Change vs reference (ONE var) Acceptance Config Ready?
Q01 M2 Sharpe-drift reproduction — adding sp_sharpe_22 (risk-adjusted 22d drift) improves net perf (exp 30: +6.53% IR 0.62, unreproduced → promote HYPOTHESIS) +sp_sharpe_22 to features net_IR > 0.21, net_ann > +2.13% workflows/q01_m2_sharpe22_repro.yaml ✅
Q02 Seed count 10 vs 5 — more seeds → higher ICIR/net; tests whether averaging saturates (exp 28 proved 5>2) seeds → 10 net_IR ≥ 0.21, ICIR/RankICIR ≥ ref workflows/q02_seed10.yaml ✅
Q03 topk 20 diversification — effective book is ~4 independent names; wider book cuts drawdown without hurting weak signal topk 10→20 net_IR > 0.21, MDD < 7.69%, net_ann ≥ +2.13% workflows/q03_topk20.yaml ✅
Q04 10-day non-overlapping label — longer horizon captures trend/reversal 5d blurs, lowers churn label → 10d net_IR > 0.21, net_ann > +2.13% workflows/q04_label10d.yaml ✅
Q05 22-day label — true trend-following; 5d can't see 1–12m drift (submartingale) label → 22d net_IR > 0.21, net_ann > +2.13%, cost ≤ ref workflows/q05_label22d.yaml ✅
Q06 Fractional-Kelly sizing (re-run exp 15 on clean lake) — sizing by edge magnitude beats equal-weight net of costs custom strategy (sizing) net_IR > 0.21, net_ann > +2.13%, cost ≤ ref designs/q06_kelly_sizing.md ⚠️ needs kelly_dropout.py
Q07 Weekly rebalance — next turnover lever after n_drop 1; cut forced churn at same signal custom strategy (weekly) cost/turnover ↓ AND net_IR > 0.21, net_ann > +2.13% designs/q07_weekly_rebalance.md ⚠️ needs weekly_rebalance.py
Q08 Risk-limit re-validation — $5M liquidity floor improves net IR / cuts DD on post-reset signal (exp 18 pre-clean-lake) rd_risk_calibrate A/B on exp-26 pred net_IR > 0.21, MDD < 7.69% vs no-limit designs/q08_risk_limit_ab.md ✅ tool-only
Q09 Long-short construction — the L/S edge (Sharpe 4.54) realizes more net of costs than long-only custom strategy (top+bottom) net_IR > 0.21, net_ann > +2.13%, cost ≤ 2× ref designs/q09_long_short.md ⚠️ needs top_bottom.py
Q10 HMM regime overlay — regime as overlay (not feature) cuts drawdown; exp 25 proved features fail, overlay untested custom strategy (regime gate) MDD < 7.69%, net_IR ≥ 0.21 designs/q10_hmm_regime_overlay.md ⚠️ needs regime_gate.py + get_lake_sp
Q11 Standalone 5-day reversal — reversal (β −0.53, t −24) tradable net of 20bp round-trip; unisolated single-feature model/backtest net_ann > 0 standalone designs/q11_standalone_reversal.md ⚠️ partial

Deferred (methodology / infra, P3)

  • Q12 Purged / walk-forward CV on the exp-26 reference (book ch.02 open question) — methodology improvement, not a direct alpha lever.
  • Q13 Out-of-universe validation — non-ETF universe for the compact stochastic feature set (book README open question; needs new lake symbols).

Execution protocol (per queued run)

  1. Validate the lake first (validate_lake_dataset + rd_status) — the clean-lake lesson: silent NaN-drops and hollow coverage invalidate a run.
  2. Trace before running (rd_trace_start with the hypothesis as rational, evolved_from=auto for lineage → it will fork from the closest prior experiment). Use a FRESH experiment name per run, e.g. tac-rd-q01-m2-....
  3. Run rd_run_workflow config_path=<abs path to the queue YAML> experiment_name=<fresh name> — use wait=false, poll rd_exp_get_run until FINISHED (4-year trains outlive the MCP call).
  4. Verify against acceptance via rd_exp_result (headline + backtest risk).
  5. Finish the trace (rd_trace_finish with metrics + evaluation), snapshot any new custom modules (rd_trace_snapshot).
  6. Report to the book — on PROVE, update book/CLAIMS.md/EVIDENCE.md; on REFUTE, record the negative (falsification is the output).

Sequential execution only (concurrent runs hang — chat-ideas.md ops lesson).

Provenance

Mined 2026-08-19 from book/ on the book branch (HEAD 436692a). Reference config reproduced byte-for-byte from the exp-26 run artifact config (/home/data/lake/mlruns/25/21afc6afdb674a399b59dd76c97628ce/artifacts/config).