queue: pre-registered experiment backlog to prove better-trading-performance hypotheses
Mined from book/ on the 'book' branch (HEAD 436692a). 11 queued runs,
each = hypothesis + one-variable change vs the exp-26 reference + acceptance
metric, per the ch.02 isolation/falsification discipline.
- workflows/: 5 runnable config-only YAMLs (Q01 M2 repro, Q02 seed10, Q03 topk20,
Q04 label10d, Q05 label22d) byte-derived from the exp-26 reference
- designs/: 6 design docs needing custom strategy modules or tool-only A/B
(Q06 Kelly, Q07 weekly rebalance, Q08 risk-limit A/B, Q09 long-short,
Q10 HMM overlay, Q11 standalone reversal)
This commit is contained in:
@@ -0,0 +1,78 @@
|
|||||||
|
# TradeAC Experiment Queue — hypotheses that would prove "better trading performance"
|
||||||
|
|
||||||
|
**Purpose.** A staging queue of experiment runs, each designed to PROVE (or
|
||||||
|
REFUTE) one hypothesis about how to achieve better trading performance on the
|
||||||
|
TradeAC stack. Every item is pre-registered: hypothesis, change-vs-reference,
|
||||||
|
and acceptance metric are fixed BEFORE the run (book ch.02 isolation + falsification
|
||||||
|
discipline). Nothing here is executed yet — each entry carries its execution
|
||||||
|
command and can be run by tracing first (`rd_trace_start` → `rd_run_workflow` /
|
||||||
|
`rd_risk_calibrate` → `rd_trace_finish`).
|
||||||
|
|
||||||
|
**Source.** Mined from the `book` branch of this repo (`book/CLAIMS.md`,
|
||||||
|
`book/EVIDENCE.md`, `book/chapters/*`, `book/references/chat-ideas.md`). Only
|
||||||
|
clean-lake (exp 21+) facts are cited as reference numbers; pre-clean-lake claims
|
||||||
|
are idea material that the queue is designed to test.
|
||||||
|
|
||||||
|
## Reference / control (MUST reproduce first)
|
||||||
|
|
||||||
|
The exp-26 reference — the campaign's best clean-lake result (EVIDENCE#015, run
|
||||||
|
`21afc6af…`, mlflow exp 25, branch `exp/26-test-whether-reducing-topkdropout-daily`):
|
||||||
|
|
||||||
|
| Config element | Reference value |
|
||||||
|
|---|---|
|
||||||
|
| Universe | 50-ETF panel (same `UNIVERSE` list as exp-24/26) |
|
||||||
|
| Features | compact stochastic: `$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead` |
|
||||||
|
| Label | `Ref($close,-6)/Ref($close,-1)-1` (5d) |
|
||||||
|
| Model | `RankICEnsembleLGBModel` (tac_qlib.contrib.model.rank_ensemble), seeds `42,7,2026,99,123`, lr 0.02, num_leaves 31, 3000 rounds, early_stop 200, min_data_in_leaf 20, lambda_l2 0.5, colsample/subsample 0.8 |
|
||||||
|
| Train / valid / test | 2016-01-04..2025-09-01 / 2025-09-03..2026-01-03 / 2026-01-04..2026-08-10 |
|
||||||
|
| Strategy | TopkDropout, topk 10, n_drop 1, risk_degree 0.95 |
|
||||||
|
| Costs / benchmark | open 0.0005 / close 0.0015 / min $5, deal $close, SPY, $1M |
|
||||||
|
|
||||||
|
Reference metrics to beat (EVIDENCE#015): **net_ann_return +2.13%, net_IR 0.21,
|
||||||
|
gross +7.02%, net_max_drawdown −7.69%, RankIC 0.0663, RankICIR 0.2545, L/S Sharpe 4.54.**
|
||||||
|
|
||||||
|
## The queue (ordered by value × feasibility)
|
||||||
|
|
||||||
|
| ID | Title / hypothesis | Change vs reference (ONE var) | Acceptance | Config | Ready? |
|
||||||
|
|----|--------------------|-------------------------------|------------|--------|--------|
|
||||||
|
| Q01 | **M2 Sharpe-drift reproduction** — adding `sp_sharpe_22` (risk-adjusted 22d drift) improves net perf (exp 30: +6.53% IR 0.62, unreproduced → promote HYPOTHESIS) | +`sp_sharpe_22` to features | net_IR > 0.21, net_ann > +2.13% | `workflows/q01_m2_sharpe22_repro.yaml` | ✅ |
|
||||||
|
| Q02 | **Seed count 10 vs 5** — more seeds → higher ICIR/net; tests whether averaging saturates (exp 28 proved 5>2) | seeds → 10 | net_IR ≥ 0.21, ICIR/RankICIR ≥ ref | `workflows/q02_seed10.yaml` | ✅ |
|
||||||
|
| Q03 | **topk 20 diversification** — effective book is ~4 independent names; wider book cuts drawdown without hurting weak signal | topk 10→20 | net_IR > 0.21, MDD < 7.69%, net_ann ≥ +2.13% | `workflows/q03_topk20.yaml` | ✅ |
|
||||||
|
| Q04 | **10-day non-overlapping label** — longer horizon captures trend/reversal 5d blurs, lowers churn | label → 10d | net_IR > 0.21, net_ann > +2.13% | `workflows/q04_label10d.yaml` | ✅ |
|
||||||
|
| Q05 | **22-day label** — true trend-following; 5d can't see 1–12m drift (submartingale) | label → 22d | net_IR > 0.21, net_ann > +2.13%, cost ≤ ref | `workflows/q05_label22d.yaml` | ✅ |
|
||||||
|
| Q06 | **Fractional-Kelly sizing** (re-run exp 15 on clean lake) — sizing by edge magnitude beats equal-weight net of costs | custom strategy (sizing) | net_IR > 0.21, net_ann > +2.13%, cost ≤ ref | `designs/q06_kelly_sizing.md` | ⚠️ needs `kelly_dropout.py` |
|
||||||
|
| Q07 | **Weekly rebalance** — next turnover lever after n_drop 1; cut forced churn at same signal | custom strategy (weekly) | cost/turnover ↓ AND net_IR > 0.21, net_ann > +2.13% | `designs/q07_weekly_rebalance.md` | ⚠️ needs `weekly_rebalance.py` |
|
||||||
|
| Q08 | **Risk-limit re-validation** — $5M liquidity floor improves net IR / cuts DD on post-reset signal (exp 18 pre-clean-lake) | `rd_risk_calibrate` A/B on exp-26 pred | net_IR > 0.21, MDD < 7.69% vs no-limit | `designs/q08_risk_limit_ab.md` | ✅ tool-only |
|
||||||
|
| Q09 | **Long-short construction** — the L/S edge (Sharpe 4.54) realizes more net of costs than long-only | custom strategy (top+bottom) | net_IR > 0.21, net_ann > +2.13%, cost ≤ 2× ref | `designs/q09_long_short.md` | ⚠️ needs `top_bottom.py` |
|
||||||
|
| Q10 | **HMM regime overlay** — regime as overlay (not feature) cuts drawdown; exp 25 proved features fail, overlay untested | custom strategy (regime gate) | MDD < 7.69%, net_IR ≥ 0.21 | `designs/q10_hmm_regime_overlay.md` | ⚠️ needs `regime_gate.py` + `get_lake_sp` |
|
||||||
|
| Q11 | **Standalone 5-day reversal** — reversal (β −0.53, t −24) tradable net of 20bp round-trip; unisolated | single-feature model/backtest | net_ann > 0 standalone | `designs/q11_standalone_reversal.md` | ⚠️ partial |
|
||||||
|
|
||||||
|
### Deferred (methodology / infra, P3)
|
||||||
|
- **Q12 Purged / walk-forward CV** on the exp-26 reference (book ch.02 open
|
||||||
|
question) — methodology improvement, not a direct alpha lever.
|
||||||
|
- **Q13 Out-of-universe validation** — non-ETF universe for the compact
|
||||||
|
stochastic feature set (book README open question; needs new lake symbols).
|
||||||
|
|
||||||
|
## Execution protocol (per queued run)
|
||||||
|
|
||||||
|
1. **Validate the lake first** (`validate_lake_dataset` + `rd_status`) — the
|
||||||
|
clean-lake lesson: silent NaN-drops and hollow coverage invalidate a run.
|
||||||
|
2. **Trace before running** (`rd_trace_start` with the hypothesis as `rational`,
|
||||||
|
`evolved_from=auto` for lineage → it will fork from the closest prior
|
||||||
|
experiment). Use a FRESH experiment name per run, e.g. `tac-rd-q01-m2-...`.
|
||||||
|
3. **Run** `rd_run_workflow config_path=<abs path to the queue YAML>
|
||||||
|
experiment_name=<fresh name>` — use `wait=false`, poll `rd_exp_get_run` until
|
||||||
|
`FINISHED` (4-year trains outlive the MCP call).
|
||||||
|
4. **Verify against acceptance** via `rd_exp_result` (headline + backtest risk).
|
||||||
|
5. **Finish the trace** (`rd_trace_finish` with `metrics` + `evaluation`),
|
||||||
|
snapshot any new custom modules (`rd_trace_snapshot`).
|
||||||
|
6. **Report to the book** — on PROVE, update `book/CLAIMS.md`/`EVIDENCE.md`;
|
||||||
|
on REFUTE, record the negative (falsification is the output).
|
||||||
|
|
||||||
|
Sequential execution only (concurrent runs hang — chat-ideas.md ops lesson).
|
||||||
|
|
||||||
|
## Provenance
|
||||||
|
|
||||||
|
Mined 2026-08-19 from `book/` on the `book` branch (HEAD `436692a`). Reference
|
||||||
|
config reproduced byte-for-byte from the exp-26 run artifact config
|
||||||
|
(`/home/data/lake/mlruns/25/21afc6afdb674a399b59dd76c97628ce/artifacts/config`).
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# QUEUE-06 — Fractional-Kelly sizing vs equal-weight top-k (re-run exp 15 on clean lake)
|
||||||
|
|
||||||
|
**Status:** QUEUED · **Priority:** P1 · **Effort:** custom strategy module + run
|
||||||
|
|
||||||
|
## Hypothesis (prove)
|
||||||
|
Fractional-Kelly sizing — sizing each name by the edge magnitude of its score
|
||||||
|
instead of equal-weight × risk_degree — is a sizing rule (not a strategy) that
|
||||||
|
throws away less edge and beats equal-weight top-k **net of costs** on the clean
|
||||||
|
lake. Source: `book/README.md` open questions (exp 15 run never finished),
|
||||||
|
`book/references/chat-ideas.md` ("Kelly sizing is a sizing rule, not a strategy").
|
||||||
|
|
||||||
|
## Change vs exp-26 reference (ONE variable)
|
||||||
|
- **Strategy**: equal-weight `TopkDropoutStrategy` (topk 10, n_drop 1) →
|
||||||
|
custom `FractionalKellyDropoutStrategy` (same topk/n_drop selection, sizing ∝
|
||||||
|
score magnitude, capped at a fraction f of the equal-weight notional; f as a
|
||||||
|
parameter, e.g. 0.5).
|
||||||
|
- All signal/config unchanged (compact stochastic features, 5-seed RankIC
|
||||||
|
ensemble, 5d label, train/valid/test, SPY benchmark, 5bp/15bp/$5 costs).
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
- `net_IR > 0.21` AND `net_ann_return > +2.13%` (exp-26 reference), with
|
||||||
|
`total_cost` not higher than the reference book.
|
||||||
|
- If sizing flattens the book (over-concentration) and net degrades → REFUTED
|
||||||
|
(recorded negative; equal-weight stays canonical).
|
||||||
|
|
||||||
|
## Execution prerequisites
|
||||||
|
1. New contrib module `tac_qlib/contrib/strategy/kelly_dropout.py`
|
||||||
|
(`FractionalKellyDropoutStrategy` subclassing
|
||||||
|
`qlib.contrib.strategy.signal_strategy.TopkDropoutStrategy`), copy to the
|
||||||
|
venv site-packages copy (`/opt/venv/lib/python3.12/site-packages/tac_qlib/...`).
|
||||||
|
2. Workflow YAML with `strategy.class=FractionalKellyDropoutStrategy`,
|
||||||
|
`module_path=tac_qlib.contrib.strategy.kelly_dropout`.
|
||||||
|
3. Trace (rd_trace_start → run → rd_trace_finish), snapshot the new module.
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
# QUEUE-07 — Turnover relief: weekly rebalance vs daily (next cost lever after n_drop 1)
|
||||||
|
|
||||||
|
**Status:** QUEUED · **Priority:** P1 · **Effort:** custom strategy module + run
|
||||||
|
|
||||||
|
## Hypothesis (prove)
|
||||||
|
n_drop 2→1 proved the cost/turnover frontier is the binding constraint
|
||||||
|
(EVIDENCE#015, ch.03/ch.09: identical IC/RankIC, net flips −3.21% → +2.13%).
|
||||||
|
The next lever in the same direction: rebalance the TopkDropout book only
|
||||||
|
**weekly** (e.g. on Mondays) instead of daily — cutting forced churn further
|
||||||
|
should lift net performance at the same signal quality.
|
||||||
|
|
||||||
|
Source: `book/references/chat-ideas.md` ("weekly rebalance" among the turnover
|
||||||
|
reduction ideas), ch.09 claim inventory.
|
||||||
|
|
||||||
|
## Change vs exp-26 reference (ONE variable)
|
||||||
|
- **Strategy**: daily TopkDropout (topk 10, n_drop 1) → custom
|
||||||
|
`WeeklyRebalanceDropoutStrategy` that recomputes the target book once per
|
||||||
|
week and otherwise holds (no-trade buffer band for small deltas).
|
||||||
|
- All signal/config unchanged.
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
- `total_cost`/turnover strictly below the reference AND `net_IR > 0.21` AND
|
||||||
|
`net_ann_return > +2.13%`.
|
||||||
|
- Reference numbers to beat: turnover ~0.74 (round-3 live), est. ~20% daily
|
||||||
|
book turnover at topk10/n_drop2 (pre-clean-lake estimate).
|
||||||
|
|
||||||
|
## Execution prerequisites
|
||||||
|
1. New contrib module `tac_qlib/contrib/strategy/weekly_rebalance.py`
|
||||||
|
(`WeeklyRebalanceDropoutStrategy` subclassing `TopkDropoutStrategy`, trade
|
||||||
|
only when the trade calendar day is the week's first trading day), copy to
|
||||||
|
the venv site-packages copy.
|
||||||
|
2. Workflow YAML wiring the strategy.
|
||||||
|
3. Trace + run + snapshot.
|
||||||
|
|
||||||
|
## Sibling (deferred)
|
||||||
|
No-trade buffer band and notional-vs-qty order sizing are variants of the same
|
||||||
|
cost lever; queue them only if Q07 reproduces positively.
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
# QUEUE-08 — Risk-limit A/B re-validation: $5M liquidity floor on the exp-26 reference
|
||||||
|
|
||||||
|
**Status:** QUEUED · **Priority:** P1 · **Effort:** tool-only (no new code)
|
||||||
|
|
||||||
|
## Hypothesis (prove)
|
||||||
|
The $5M liquidity floor improves net IR and cuts drawdown on the **post-reset**
|
||||||
|
reference signal (pre-reset exp 18, EVIDENCE#008: net IR 0.81→0.98, cumDD
|
||||||
|
7.93%→5.44%), while size/concentration caps hurt by cutting deployed capital.
|
||||||
|
Needs re-validation on the exp-26 lineage because exp 18 is pre-clean-lake and
|
||||||
|
not comparable (EVIDENCE#009/010). Source: `book/CLAIMS.md` open question +
|
||||||
|
`book/README.md` `TODO(evidence-needed: reconciliation of exp 18 risk-limit spec
|
||||||
|
on the post-reset reference signal)`.
|
||||||
|
|
||||||
|
## Change vs exp-26 reference (ONE variable)
|
||||||
|
- Reference: the saved exp-26 prediction (run `21afc6af…`, mlflow exp 25).
|
||||||
|
- A/B via `rd_risk_calibrate` (runs limit-vs-no-limit A/B + sensitivity grid
|
||||||
|
over size_cap_pct, concentration_cap_pct, liquidity_floor_adv) and/or
|
||||||
|
`rd_backtest` with `risk_limits` on the SAME saved `pred.pkl`:
|
||||||
|
- baseline: no limits (this must reproduce the exp-26 net +2.13% / IR 0.21);
|
||||||
|
- candidate: `{"liquidity_floor_adv": 5000000, "size_cap_pct": 0.12,
|
||||||
|
"concentration_cap_pct": 0.95, "drawdown_pause_pct": 0.10}` (round-3 spec).
|
||||||
|
- Pick the spec (B2 calibration) that keeps live ≈ backtest.
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
- Candidate spec: `net_IR > 0.21` AND `net_max_drawdown < 7.69%` vs no-limit on
|
||||||
|
the same pred. Size/concentration caps expected to REDUCE deployed capital
|
||||||
|
(record the direction as confirmation of exp 18).
|
||||||
|
- If the floor is a no-op (gates don't bind at this signal) → report that gates
|
||||||
|
are no-ops when the signal is the bottleneck (exp 20 pattern) as a PROVEN
|
||||||
|
clean-lake result.
|
||||||
|
|
||||||
|
## Execution prerequisites
|
||||||
|
- None (uses saved pred + `rd_risk_calibrate`/`rd_backtest`). Trace the A/B as
|
||||||
|
an experiment; record the spec chosen for the next live round.
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
# QUEUE-09 — Long-short construction: capture the long-short edge net of costs
|
||||||
|
|
||||||
|
**Status:** QUEUED · **Priority:** P2 · **Effort:** custom strategy module + run
|
||||||
|
|
||||||
|
## Hypothesis (prove)
|
||||||
|
The compact stochastic signal's long-short spread is the real edge (L/S ann
|
||||||
|
Sharpe 4.54, exp 24; "edge is long-short, not long-only" — chat-ideas.md), but
|
||||||
|
all canonical constructions are long-only (TopkDropout buys topk, drops, holds).
|
||||||
|
A market-neutral book (long topk, short bottom topk) should realize more of the
|
||||||
|
spread net of costs than the long-only book, IF short-side financing + doubled
|
||||||
|
turnover cost stays below the added spread capture.
|
||||||
|
|
||||||
|
## Change vs exp-26 reference (ONE variable)
|
||||||
|
- **Strategy**: long-only TopkDropout (topk 10, n_drop 1) → custom
|
||||||
|
`TopBottomDropoutStrategy` (long topk by rank, short bottom topk, equal
|
||||||
|
weight per side, same risk_degree), realized in a workflow with a cost model
|
||||||
|
that includes both sides (open/close cost symmetric).
|
||||||
|
- All signal/config unchanged.
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
- `net_IR > 0.21` AND `net_ann_return > +2.13%` AND `total_cost` within ~2× the
|
||||||
|
reference (doubled side count is the structural cost of this construction).
|
||||||
|
- Watch: benchmark neutrality (SPY beta ≈ 0) as a secondary sanity metric.
|
||||||
|
|
||||||
|
## Execution prerequisites
|
||||||
|
1. New contrib module `tac_qlib/contrib/strategy/top_bottom.py`
|
||||||
|
(`TopBottomDropoutStrategy` subclassing `BaseSignalStrategy`), copy to the
|
||||||
|
venv site-packages copy.
|
||||||
|
2. Workflow YAML wiring the strategy; PortAnaRecord benchmark SPY.
|
||||||
|
3. Trace + run + snapshot.
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
# QUEUE-10 — HMM regime overlay on the exp-26 book (overlay, not feature)
|
||||||
|
|
||||||
|
**Status:** QUEUED · **Priority:** P2 · **Effort:** custom strategy + feature compute + run
|
||||||
|
|
||||||
|
## Hypothesis (prove)
|
||||||
|
Regime flags failed as model **features** (exp 9 idea, exp 25 clean-lake
|
||||||
|
confirmation that model-specific families regress), but the surviving use is as
|
||||||
|
an **overlay**: a long-only/regime-gate that holds names only in the favourable
|
||||||
|
HMM state should cut drawdown / improve net IR on the same signal. Source:
|
||||||
|
`book/chapters/01` regime section + `chat-ideas.md`
|
||||||
|
(`TODO(evidence-needed: HMM regime gate as overlay on exp-26 book)`).
|
||||||
|
|
||||||
|
## Change vs exp-26 reference (ONE variable)
|
||||||
|
- **Strategy**: plain TopkDropout (topk 10, n_drop 1) → custom
|
||||||
|
`RegimeGateDropoutStrategy`: identical selection, but when the per-symbol
|
||||||
|
HMM posterior (`sp_hmm_p_regime1`) is below a calibrated threshold the name
|
||||||
|
is held in cash instead of bought (entry gate); no new features enter the
|
||||||
|
model — `sp_hmm_p_regime1` is computed for gating only, fit on the train
|
||||||
|
window (no lookahead), via `get_lake_sp` with `fit_end=<train end>`.
|
||||||
|
- All signal/config unchanged.
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
- `net_max_drawdown < 7.69%` (reference) AND `net_IR >= 0.21`. If the gate
|
||||||
|
never binds at a sensible threshold → the gate is a no-op on this signal
|
||||||
|
(exp 20 pattern) → recorded REFUTED/neutral, not a failure.
|
||||||
|
- Calibrate the threshold on the valid window only (avoid the exp 13/14
|
||||||
|
threshold-overfit trap).
|
||||||
|
|
||||||
|
## Execution prerequisites
|
||||||
|
1. Persist `sp_hmm_p_regime1` for the universe (get_lake_sp, fit_end =
|
||||||
|
2025-09-01) WITHOUT adding it to `feature_fields` of the model.
|
||||||
|
2. New contrib module `tac_qlib/contrib/strategy/regime_gate.py`, copy to the
|
||||||
|
venv site-packages copy.
|
||||||
|
3. Workflow YAML wiring the strategy.
|
||||||
|
4. Trace + run + snapshot.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# QUEUE-11 — Standalone 5-day reversal signal net of costs (unisolated)
|
||||||
|
|
||||||
|
**Status:** QUEUED · **Priority:** P2 · **Effort:** dataset study + backtest
|
||||||
|
|
||||||
|
## Hypothesis (prove)
|
||||||
|
5-day momentum strongly reverses on this panel (pooled regression:
|
||||||
|
`sp_trend_slope_5` β = −0.53, t = −24; VR < 1 at 5–20d for ~32/72 assets —
|
||||||
|
chat-derived, pre-clean-lake idea material). The reversal has never been tested
|
||||||
|
as a **standalone tradable strategy net of costs**. If it clears the 20bp
|
||||||
|
round-trip cost, it is an independent alpha source that can be blended with (or
|
||||||
|
replace) the model book.
|
||||||
|
Source: `book/ch01` "Timeline" + `chat-ideas.md`
|
||||||
|
(`TODO(evidence-needed: standalone 5d-reversal strategy net of costs)`).
|
||||||
|
|
||||||
|
## Change vs exp-26 reference
|
||||||
|
- This is NOT a model-construction variant — it isolates a SINGLE-FEATURE
|
||||||
|
signal: a model trained on `sp_trend_slope_5` (plus raw OHLCV) alone, or a
|
||||||
|
mechanical reversal book (rank by −`sp_trend_slope_5`, buy the most-reverted
|
||||||
|
topk), backtested net of costs over the exp-26 window.
|
||||||
|
- Control: exp-26 compact reference on the same window.
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
- Standalone reversal `net_ann_return > 0` (clears 20bp round-trip) — proves
|
||||||
|
the claim "reversal is tradable net of costs". Secondary: excess vs the
|
||||||
|
model book is the blend decision for a future round.
|
||||||
|
|
||||||
|
## Execution prerequisites
|
||||||
|
1. `rd_train`/workflow with `feature_fields = $open,$high,$low,$close,$vwap,$volume,sp_trend_slope_5`
|
||||||
|
(single feature) OR a mechanical rank backtest via `rd_backtest` on a
|
||||||
|
hand-built pred (pred = −rank(sp_trend_slope_5)).
|
||||||
|
2. Trace + run + record as a standalone study (dataset-study status, not
|
||||||
|
necessarily a traced model experiment).
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
# QUEUE-01 — M2 reproduction: risk-adjusted 22d Sharpe drift (sp_sharpe_22).
|
||||||
|
#
|
||||||
|
# Hypothesis (book ch.01/ch.07, EVIDENCE#018 -> exp 30): adding the
|
||||||
|
# risk-adjusted 22d Sharpe drift feature (sp_sharpe_22) to the compact
|
||||||
|
# stochastic reference IMPROVES net portfolio performance (exp 30: net +6.53%
|
||||||
|
# IR 0.62 vs reference +2.13% IR 0.21) while rank metrics dip (RankIC 0.0576 vs
|
||||||
|
# 0.0663). exp 30 is a SINGLE clean-lake run, unreproduced -> HYPOTHESIS.
|
||||||
|
#
|
||||||
|
# Change vs exp-26 reference (EVIDENCE#015, run 21afc6af...): ONE feature added,
|
||||||
|
# feature_fields = compact set + sp_sharpe_22. Everything else byte-identical.
|
||||||
|
#
|
||||||
|
# Acceptance: net_ann_return > +2.13% AND net_IR > 0.21 (else HYPOTHESIS -> REFUTED).
|
||||||
|
# Run: rd_run_workflow config_path=<repo>/experiments/queue/workflows/q01_m2_sharpe22_repro.yaml \
|
||||||
|
# experiment_name=tac-rd-q01-m2-sharpe22-repro
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
{%- set LAKE = TAC_LAKE_DIR %}
|
||||||
|
{%- set UNIVERSE = "SPY,QQQ,DIA,IWM,MDY,VTI,VOO,VEA,VWO,VT,EFA,EEM,TLT,IEF,SHY,AGG,BND,LQD,HYG,JNK,EMB,GLD,SLV,USO,UNG,DBA,DBC,XLK,XLF,XLE,XLV,XLI,XLY,XLP,XLU,XLB,XLRE,ARKK,SMH,SOXX,IBB,XBI,ITA,XAR,ICLN,TAN,FDN,IGV,ESPO,REM" %}
|
||||||
|
{%- set FEATURES = "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead,sp_sharpe_22" %}
|
||||||
|
|
||||||
|
qlib_init:
|
||||||
|
provider_uri: "{{ LAKE }}"
|
||||||
|
region: us
|
||||||
|
expression_cache: null
|
||||||
|
dataset_cache: null
|
||||||
|
|
||||||
|
calendar_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeCalendarProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
instrument_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeInstrumentProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
markets: {}
|
||||||
|
feature_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeFeatureProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
|
||||||
|
exp_manager:
|
||||||
|
class: MLflowExpManager
|
||||||
|
module_path: qlib.workflow.expm
|
||||||
|
kwargs:
|
||||||
|
uri: "sqlite:///{{ LAKE }}/mlruns.db"
|
||||||
|
default_exp_name: "tac-rd-q01-m2-sharpe22-repro"
|
||||||
|
|
||||||
|
task:
|
||||||
|
model:
|
||||||
|
class: RankICEnsembleLGBModel
|
||||||
|
module_path: tac_qlib.contrib.model.rank_ensemble
|
||||||
|
kwargs:
|
||||||
|
loss: mse
|
||||||
|
learning_rate: 0.02
|
||||||
|
num_leaves: 31
|
||||||
|
n_estimators: 3000
|
||||||
|
num_boost_round: 3000
|
||||||
|
early_stopping_rounds: 200
|
||||||
|
min_data_in_leaf: 20
|
||||||
|
lambda_l2: 0.5
|
||||||
|
colsample_bytree: 0.8
|
||||||
|
subsample: 0.8
|
||||||
|
subsample_freq: 1
|
||||||
|
reg_alpha: 0.1
|
||||||
|
reg_lambda: 1.0
|
||||||
|
seeds: "42,7,2026,99,123"
|
||||||
|
parallel: 5
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
class: DatasetH
|
||||||
|
module_path: qlib.data.dataset
|
||||||
|
kwargs:
|
||||||
|
handler:
|
||||||
|
class: TACHandler
|
||||||
|
module_path: tac_qlib.contrib.data.handler
|
||||||
|
kwargs:
|
||||||
|
instruments: "{{ UNIVERSE }}"
|
||||||
|
start_time: 2015-01-03
|
||||||
|
end_time: 2026-08-10
|
||||||
|
fit_start_time: 2016-01-04
|
||||||
|
fit_end_time: 2025-09-01
|
||||||
|
freq: day
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
label: "Ref($close,-6)/Ref($close,-1)-1"
|
||||||
|
feature_fields: "{{ FEATURES }}"
|
||||||
|
infer_processors:
|
||||||
|
- class: DropAllNaN
|
||||||
|
kwargs: {}
|
||||||
|
- class: ProcessInf
|
||||||
|
kwargs: {}
|
||||||
|
- class: CSRankNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: ZScoreNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: Fillna
|
||||||
|
kwargs: {}
|
||||||
|
segments:
|
||||||
|
train: [2016-01-04, 2025-09-01]
|
||||||
|
valid: [2025-09-03, 2026-01-03]
|
||||||
|
test: [2026-01-04, 2026-08-10]
|
||||||
|
|
||||||
|
record:
|
||||||
|
- class: SignalRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs: {}
|
||||||
|
- class: SigAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
ana_long_short: true
|
||||||
|
ann_scaler: 252
|
||||||
|
- class: PortAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
config:
|
||||||
|
strategy:
|
||||||
|
class: TopkDropoutStrategy
|
||||||
|
module_path: qlib.contrib.strategy
|
||||||
|
kwargs:
|
||||||
|
signal: "<PRED>"
|
||||||
|
topk: 10
|
||||||
|
n_drop: 1
|
||||||
|
only_tradable: true
|
||||||
|
risk_degree: 0.95
|
||||||
|
backtest:
|
||||||
|
start_time: 2026-01-04
|
||||||
|
end_time: 2026-08-10
|
||||||
|
account: 1000000
|
||||||
|
benchmark: SPY
|
||||||
|
exchange_kwargs:
|
||||||
|
codes: "{{ UNIVERSE }}"
|
||||||
|
deal_price: $close
|
||||||
|
freq: day
|
||||||
|
open_cost: 0.0005
|
||||||
|
close_cost: 0.0015
|
||||||
|
min_cost: 5.0
|
||||||
|
risk_analysis_freq: 1d
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
# QUEUE-02 — Seed-count 10 vs 5 on the compact reference.
|
||||||
|
#
|
||||||
|
# Hypothesis (book ch.05, EVIDENCE#016 -> exp 28): seed count is load-bearing
|
||||||
|
# (2 seeds lose to 5). Extending the same direction, does 10 seeds further
|
||||||
|
# raise ICIR and net performance? Tests whether averaging benefit saturates.
|
||||||
|
#
|
||||||
|
# Change vs exp-26 reference: ONE variable — seeds "42,7,2026,99,123" ->
|
||||||
|
# "42,7,2026,99,123,17,3,2020,88,55" (parallel: 10). Everything else identical.
|
||||||
|
#
|
||||||
|
# Acceptance: net_IR >= 0.21 AND ICIR/RankICIR >= reference (0.235 / 0.243);
|
||||||
|
# if seed count saturates, expect flat ICIR — that result also settles the
|
||||||
|
# mechanism question (variance reduction, not family diversification).
|
||||||
|
# Run: rd_run_workflow config_path=<repo>/experiments/queue/workflows/q02_seed10.yaml \
|
||||||
|
# experiment_name=tac-rd-q02-seed10
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
{%- set LAKE = TAC_LAKE_DIR %}
|
||||||
|
{%- set UNIVERSE = "SPY,QQQ,DIA,IWM,MDY,VTI,VOO,VEA,VWO,VT,EFA,EEM,TLT,IEF,SHY,AGG,BND,LQD,HYG,JNK,EMB,GLD,SLV,USO,UNG,DBA,DBC,XLK,XLF,XLE,XLV,XLI,XLY,XLP,XLU,XLB,XLRE,ARKK,SMH,SOXX,IBB,XBI,ITA,XAR,ICLN,TAN,FDN,IGV,ESPO,REM" %}
|
||||||
|
{%- set FEATURES = "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead" %}
|
||||||
|
|
||||||
|
qlib_init:
|
||||||
|
provider_uri: "{{ LAKE }}"
|
||||||
|
region: us
|
||||||
|
expression_cache: null
|
||||||
|
dataset_cache: null
|
||||||
|
|
||||||
|
calendar_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeCalendarProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
instrument_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeInstrumentProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
markets: {}
|
||||||
|
feature_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeFeatureProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
|
||||||
|
exp_manager:
|
||||||
|
class: MLflowExpManager
|
||||||
|
module_path: qlib.workflow.expm
|
||||||
|
kwargs:
|
||||||
|
uri: "sqlite:///{{ LAKE }}/mlruns.db"
|
||||||
|
default_exp_name: "tac-rd-q02-seed10"
|
||||||
|
|
||||||
|
task:
|
||||||
|
model:
|
||||||
|
class: RankICEnsembleLGBModel
|
||||||
|
module_path: tac_qlib.contrib.model.rank_ensemble
|
||||||
|
kwargs:
|
||||||
|
loss: mse
|
||||||
|
learning_rate: 0.02
|
||||||
|
num_leaves: 31
|
||||||
|
n_estimators: 3000
|
||||||
|
num_boost_round: 3000
|
||||||
|
early_stopping_rounds: 200
|
||||||
|
min_data_in_leaf: 20
|
||||||
|
lambda_l2: 0.5
|
||||||
|
colsample_bytree: 0.8
|
||||||
|
subsample: 0.8
|
||||||
|
subsample_freq: 1
|
||||||
|
reg_alpha: 0.1
|
||||||
|
reg_lambda: 1.0
|
||||||
|
seeds: "42,7,2026,99,123,17,3,2020,88,55"
|
||||||
|
parallel: 10
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
class: DatasetH
|
||||||
|
module_path: qlib.data.dataset
|
||||||
|
kwargs:
|
||||||
|
handler:
|
||||||
|
class: TACHandler
|
||||||
|
module_path: tac_qlib.contrib.data.handler
|
||||||
|
kwargs:
|
||||||
|
instruments: "{{ UNIVERSE }}"
|
||||||
|
start_time: 2015-01-03
|
||||||
|
end_time: 2026-08-10
|
||||||
|
fit_start_time: 2016-01-04
|
||||||
|
fit_end_time: 2025-09-01
|
||||||
|
freq: day
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
label: "Ref($close,-6)/Ref($close,-1)-1"
|
||||||
|
feature_fields: "{{ FEATURES }}"
|
||||||
|
infer_processors:
|
||||||
|
- class: DropAllNaN
|
||||||
|
kwargs: {}
|
||||||
|
- class: ProcessInf
|
||||||
|
kwargs: {}
|
||||||
|
- class: CSRankNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: ZScoreNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: Fillna
|
||||||
|
kwargs: {}
|
||||||
|
segments:
|
||||||
|
train: [2016-01-04, 2025-09-01]
|
||||||
|
valid: [2025-09-03, 2026-01-03]
|
||||||
|
test: [2026-01-04, 2026-08-10]
|
||||||
|
|
||||||
|
record:
|
||||||
|
- class: SignalRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs: {}
|
||||||
|
- class: SigAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
ana_long_short: true
|
||||||
|
ann_scaler: 252
|
||||||
|
- class: PortAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
config:
|
||||||
|
strategy:
|
||||||
|
class: TopkDropoutStrategy
|
||||||
|
module_path: qlib.contrib.strategy
|
||||||
|
kwargs:
|
||||||
|
signal: "<PRED>"
|
||||||
|
topk: 10
|
||||||
|
n_drop: 1
|
||||||
|
only_tradable: true
|
||||||
|
risk_degree: 0.95
|
||||||
|
backtest:
|
||||||
|
start_time: 2026-01-04
|
||||||
|
end_time: 2026-08-10
|
||||||
|
account: 1000000
|
||||||
|
benchmark: SPY
|
||||||
|
exchange_kwargs:
|
||||||
|
codes: "{{ UNIVERSE }}"
|
||||||
|
deal_price: $close
|
||||||
|
freq: day
|
||||||
|
open_cost: 0.0005
|
||||||
|
close_cost: 0.0015
|
||||||
|
min_cost: 5.0
|
||||||
|
risk_analysis_freq: 1d
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
# QUEUE-03 — topk 20 vs 10 diversification on the compact reference.
|
||||||
|
#
|
||||||
|
# Hypothesis (book ch.05/chat-ideas): the effective independent names in the
|
||||||
|
# 50-ETF book is small (~4, chat-derived eigenvalue analysis); raising topk
|
||||||
|
# diversifies the book and should cut drawdown / raise net IR without hurting
|
||||||
|
# the (weak) rank signal — cost relief by spreading the book wider.
|
||||||
|
#
|
||||||
|
# Change vs exp-26 reference: ONE variable — strategy topk 10 -> 20 (n_drop 1).
|
||||||
|
# Everything else identical.
|
||||||
|
#
|
||||||
|
# Acceptance: net_IR > 0.21 AND net_max_drawdown < 7.69% AND net_ann_return >=
|
||||||
|
# +2.13%; watch total_cost — more names held must not raise turnover/cost.
|
||||||
|
# Run: rd_run_workflow config_path=<repo>/experiments/queue/workflows/q03_topk20.yaml \
|
||||||
|
# experiment_name=tac-rd-q03-topk20
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
{%- set LAKE = TAC_LAKE_DIR %}
|
||||||
|
{%- set UNIVERSE = "SPY,QQQ,DIA,IWM,MDY,VTI,VOO,VEA,VWO,VT,EFA,EEM,TLT,IEF,SHY,AGG,BND,LQD,HYG,JNK,EMB,GLD,SLV,USO,UNG,DBA,DBC,XLK,XLF,XLE,XLV,XLI,XLY,XLP,XLU,XLB,XLRE,ARKK,SMH,SOXX,IBB,XBI,ITA,XAR,ICLN,TAN,FDN,IGV,ESPO,REM" %}
|
||||||
|
{%- set FEATURES = "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead" %}
|
||||||
|
|
||||||
|
qlib_init:
|
||||||
|
provider_uri: "{{ LAKE }}"
|
||||||
|
region: us
|
||||||
|
expression_cache: null
|
||||||
|
dataset_cache: null
|
||||||
|
|
||||||
|
calendar_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeCalendarProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
instrument_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeInstrumentProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
markets: {}
|
||||||
|
feature_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeFeatureProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
|
||||||
|
exp_manager:
|
||||||
|
class: MLflowExpManager
|
||||||
|
module_path: qlib.workflow.expm
|
||||||
|
kwargs:
|
||||||
|
uri: "sqlite:///{{ LAKE }}/mlruns.db"
|
||||||
|
default_exp_name: "tac-rd-q03-topk20"
|
||||||
|
|
||||||
|
task:
|
||||||
|
model:
|
||||||
|
class: RankICEnsembleLGBModel
|
||||||
|
module_path: tac_qlib.contrib.model.rank_ensemble
|
||||||
|
kwargs:
|
||||||
|
loss: mse
|
||||||
|
learning_rate: 0.02
|
||||||
|
num_leaves: 31
|
||||||
|
n_estimators: 3000
|
||||||
|
num_boost_round: 3000
|
||||||
|
early_stopping_rounds: 200
|
||||||
|
min_data_in_leaf: 20
|
||||||
|
lambda_l2: 0.5
|
||||||
|
colsample_bytree: 0.8
|
||||||
|
subsample: 0.8
|
||||||
|
subsample_freq: 1
|
||||||
|
reg_alpha: 0.1
|
||||||
|
reg_lambda: 1.0
|
||||||
|
seeds: "42,7,2026,99,123"
|
||||||
|
parallel: 5
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
class: DatasetH
|
||||||
|
module_path: qlib.data.dataset
|
||||||
|
kwargs:
|
||||||
|
handler:
|
||||||
|
class: TACHandler
|
||||||
|
module_path: tac_qlib.contrib.data.handler
|
||||||
|
kwargs:
|
||||||
|
instruments: "{{ UNIVERSE }}"
|
||||||
|
start_time: 2015-01-03
|
||||||
|
end_time: 2026-08-10
|
||||||
|
fit_start_time: 2016-01-04
|
||||||
|
fit_end_time: 2025-09-01
|
||||||
|
freq: day
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
label: "Ref($close,-6)/Ref($close,-1)-1"
|
||||||
|
feature_fields: "{{ FEATURES }}"
|
||||||
|
infer_processors:
|
||||||
|
- class: DropAllNaN
|
||||||
|
kwargs: {}
|
||||||
|
- class: ProcessInf
|
||||||
|
kwargs: {}
|
||||||
|
- class: CSRankNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: ZScoreNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: Fillna
|
||||||
|
kwargs: {}
|
||||||
|
segments:
|
||||||
|
train: [2016-01-04, 2025-09-01]
|
||||||
|
valid: [2025-09-03, 2026-01-03]
|
||||||
|
test: [2026-01-04, 2026-08-10]
|
||||||
|
|
||||||
|
record:
|
||||||
|
- class: SignalRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs: {}
|
||||||
|
- class: SigAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
ana_long_short: true
|
||||||
|
ann_scaler: 252
|
||||||
|
- class: PortAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
config:
|
||||||
|
strategy:
|
||||||
|
class: TopkDropoutStrategy
|
||||||
|
module_path: qlib.contrib.strategy
|
||||||
|
kwargs:
|
||||||
|
signal: "<PRED>"
|
||||||
|
topk: 20
|
||||||
|
n_drop: 1
|
||||||
|
only_tradable: true
|
||||||
|
risk_degree: 0.95
|
||||||
|
backtest:
|
||||||
|
start_time: 2026-01-04
|
||||||
|
end_time: 2026-08-10
|
||||||
|
account: 1000000
|
||||||
|
benchmark: SPY
|
||||||
|
exchange_kwargs:
|
||||||
|
codes: "{{ UNIVERSE }}"
|
||||||
|
deal_price: $close
|
||||||
|
freq: day
|
||||||
|
open_cost: 0.0005
|
||||||
|
close_cost: 0.0015
|
||||||
|
min_cost: 5.0
|
||||||
|
risk_analysis_freq: 1d
|
||||||
@@ -0,0 +1,141 @@
|
|||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
# QUEUE-04 — Non-overlapping 10-day label horizon.
|
||||||
|
#
|
||||||
|
# Hypothesis (book ch.01/chat-ideas): the 5d label is the campaign's best IC
|
||||||
|
# lever but sees short-horizon reversal only; a non-overlapping 10d label
|
||||||
|
# (`Ref($close,-11)/Ref($close,-1)-1`) tests whether a longer, cleaner horizon
|
||||||
|
# captures trend/reversal better and survives cost (lower effective turnover).
|
||||||
|
#
|
||||||
|
# Change vs exp-26 reference: ONE variable — label 5d -> 10d. Everything else
|
||||||
|
# identical (features, model, strategy).
|
||||||
|
#
|
||||||
|
# Acceptance: net_IR > 0.21 AND net_ann_return > +2.13%; secondary: ICIR and
|
||||||
|
# L/S Sharpe >= reference. A flat-but-not-worse result still settles the
|
||||||
|
# horizon-decomposition question (TODO: 5d can't see 1-12m drift).
|
||||||
|
# Run: rd_run_workflow config_path=<repo>/experiments/queue/workflows/q04_label10d.yaml \
|
||||||
|
# experiment_name=tac-rd-q04-label10d
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
{%- set LAKE = TAC_LAKE_DIR %}
|
||||||
|
{%- set UNIVERSE = "SPY,QQQ,DIA,IWM,MDY,VTI,VOO,VEA,VWO,VT,EFA,EEM,TLT,IEF,SHY,AGG,BND,LQD,HYG,JNK,EMB,GLD,SLV,USO,UNG,DBA,DBC,XLK,XLF,XLE,XLV,XLI,XLY,XLP,XLU,XLB,XLRE,ARKK,SMH,SOXX,IBB,XBI,ITA,XAR,ICLN,TAN,FDN,IGV,ESPO,REM" %}
|
||||||
|
{%- set FEATURES = "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead" %}
|
||||||
|
|
||||||
|
qlib_init:
|
||||||
|
provider_uri: "{{ LAKE }}"
|
||||||
|
region: us
|
||||||
|
expression_cache: null
|
||||||
|
dataset_cache: null
|
||||||
|
|
||||||
|
calendar_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeCalendarProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
instrument_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeInstrumentProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
markets: {}
|
||||||
|
feature_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeFeatureProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
|
||||||
|
exp_manager:
|
||||||
|
class: MLflowExpManager
|
||||||
|
module_path: qlib.workflow.expm
|
||||||
|
kwargs:
|
||||||
|
uri: "sqlite:///{{ LAKE }}/mlruns.db"
|
||||||
|
default_exp_name: "tac-rd-q04-label10d"
|
||||||
|
|
||||||
|
task:
|
||||||
|
model:
|
||||||
|
class: RankICEnsembleLGBModel
|
||||||
|
module_path: tac_qlib.contrib.model.rank_ensemble
|
||||||
|
kwargs:
|
||||||
|
loss: mse
|
||||||
|
learning_rate: 0.02
|
||||||
|
num_leaves: 31
|
||||||
|
n_estimators: 3000
|
||||||
|
num_boost_round: 3000
|
||||||
|
early_stopping_rounds: 200
|
||||||
|
min_data_in_leaf: 20
|
||||||
|
lambda_l2: 0.5
|
||||||
|
colsample_bytree: 0.8
|
||||||
|
subsample: 0.8
|
||||||
|
subsample_freq: 1
|
||||||
|
reg_alpha: 0.1
|
||||||
|
reg_lambda: 1.0
|
||||||
|
seeds: "42,7,2026,99,123"
|
||||||
|
parallel: 5
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
class: DatasetH
|
||||||
|
module_path: qlib.data.dataset
|
||||||
|
kwargs:
|
||||||
|
handler:
|
||||||
|
class: TACHandler
|
||||||
|
module_path: tac_qlib.contrib.data.handler
|
||||||
|
kwargs:
|
||||||
|
instruments: "{{ UNIVERSE }}"
|
||||||
|
start_time: 2015-01-03
|
||||||
|
end_time: 2026-08-10
|
||||||
|
fit_start_time: 2016-01-04
|
||||||
|
fit_end_time: 2025-09-01
|
||||||
|
freq: day
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
label: "Ref($close,-11)/Ref($close,-1)-1"
|
||||||
|
feature_fields: "{{ FEATURES }}"
|
||||||
|
infer_processors:
|
||||||
|
- class: DropAllNaN
|
||||||
|
kwargs: {}
|
||||||
|
- class: ProcessInf
|
||||||
|
kwargs: {}
|
||||||
|
- class: CSRankNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: ZScoreNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: Fillna
|
||||||
|
kwargs: {}
|
||||||
|
segments:
|
||||||
|
train: [2016-01-04, 2025-09-01]
|
||||||
|
valid: [2025-09-03, 2026-01-03]
|
||||||
|
test: [2026-01-04, 2026-08-10]
|
||||||
|
|
||||||
|
record:
|
||||||
|
- class: SignalRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs: {}
|
||||||
|
- class: SigAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
ana_long_short: true
|
||||||
|
ann_scaler: 252
|
||||||
|
- class: PortAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
config:
|
||||||
|
strategy:
|
||||||
|
class: TopkDropoutStrategy
|
||||||
|
module_path: qlib.contrib.strategy
|
||||||
|
kwargs:
|
||||||
|
signal: "<PRED>"
|
||||||
|
topk: 10
|
||||||
|
n_drop: 1
|
||||||
|
only_tradable: true
|
||||||
|
risk_degree: 0.95
|
||||||
|
backtest:
|
||||||
|
start_time: 2026-01-04
|
||||||
|
end_time: 2026-08-10
|
||||||
|
account: 1000000
|
||||||
|
benchmark: SPY
|
||||||
|
exchange_kwargs:
|
||||||
|
codes: "{{ UNIVERSE }}"
|
||||||
|
deal_price: $close
|
||||||
|
freq: day
|
||||||
|
open_cost: 0.0005
|
||||||
|
close_cost: 0.0015
|
||||||
|
min_cost: 5.0
|
||||||
|
risk_analysis_freq: 1d
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
# QUEUE-05 — Non-overlapping 22-day label horizon.
|
||||||
|
#
|
||||||
|
# Hypothesis (book ch.01/chat-ideas): a 22d (~monthly) non-overlapping label
|
||||||
|
# tests true trend-following — the 5d label can't distinguish a 1-12m drift
|
||||||
|
# (submartingale) from short-horizon reversal. Long-horizon labels also cut
|
||||||
|
# the rebalance-implied turnover, attacking the cost constraint directly.
|
||||||
|
#
|
||||||
|
# Change vs exp-26 reference: ONE variable — label 5d -> 22d
|
||||||
|
# (`Ref($close,-23)/Ref($close,-1)-1`). Everything else identical.
|
||||||
|
#
|
||||||
|
# Acceptance: net_IR > 0.21 AND net_ann_return > +2.13%; secondary: does the
|
||||||
|
# long-horizon signal survive cost with LOWER total_cost than the 5d book?
|
||||||
|
# Run: rd_run_workflow config_path=<repo>/experiments/queue/workflows/q05_label22d.yaml \
|
||||||
|
# experiment_name=tac-rd-q05-label22d
|
||||||
|
# -----------------------------------------------------------------------------
|
||||||
|
{%- set LAKE = TAC_LAKE_DIR %}
|
||||||
|
{%- set UNIVERSE = "SPY,QQQ,DIA,IWM,MDY,VTI,VOO,VEA,VWO,VT,EFA,EEM,TLT,IEF,SHY,AGG,BND,LQD,HYG,JNK,EMB,GLD,SLV,USO,UNG,DBA,DBC,XLK,XLF,XLE,XLV,XLI,XLY,XLP,XLU,XLB,XLRE,ARKK,SMH,SOXX,IBB,XBI,ITA,XAR,ICLN,TAN,FDN,IGV,ESPO,REM" %}
|
||||||
|
{%- set FEATURES = "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_jump_ratio,sp_jump_flag,sp_jump_tail,sp_max_move,sp_rv1,sp_rv5,sp_rv22,sp_vol_ratio_5_22,sp_vol_ratio_1_22,sp_trend_slope_5,sp_trend_slope_20,sp_trend_slope_60,sp_logp,sp_hurst_exponent,sp_sig_level1_lead,sp_sig_level1_lag,sp_sig_level2_lead_lag,sp_sig_level2_lag_lead" %}
|
||||||
|
|
||||||
|
qlib_init:
|
||||||
|
provider_uri: "{{ LAKE }}"
|
||||||
|
region: us
|
||||||
|
expression_cache: null
|
||||||
|
dataset_cache: null
|
||||||
|
|
||||||
|
calendar_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeCalendarProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
instrument_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeInstrumentProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
markets: {}
|
||||||
|
feature_provider:
|
||||||
|
class: tac_qlib.data.providers.LakeFeatureProvider
|
||||||
|
kwargs:
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
|
||||||
|
exp_manager:
|
||||||
|
class: MLflowExpManager
|
||||||
|
module_path: qlib.workflow.expm
|
||||||
|
kwargs:
|
||||||
|
uri: "sqlite:///{{ LAKE }}/mlruns.db"
|
||||||
|
default_exp_name: "tac-rd-q05-label22d"
|
||||||
|
|
||||||
|
task:
|
||||||
|
model:
|
||||||
|
class: RankICEnsembleLGBModel
|
||||||
|
module_path: tac_qlib.contrib.model.rank_ensemble
|
||||||
|
kwargs:
|
||||||
|
loss: mse
|
||||||
|
learning_rate: 0.02
|
||||||
|
num_leaves: 31
|
||||||
|
n_estimators: 3000
|
||||||
|
num_boost_round: 3000
|
||||||
|
early_stopping_rounds: 200
|
||||||
|
min_data_in_leaf: 20
|
||||||
|
lambda_l2: 0.5
|
||||||
|
colsample_bytree: 0.8
|
||||||
|
subsample: 0.8
|
||||||
|
subsample_freq: 1
|
||||||
|
reg_alpha: 0.1
|
||||||
|
reg_lambda: 1.0
|
||||||
|
seeds: "42,7,2026,99,123"
|
||||||
|
parallel: 5
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
class: DatasetH
|
||||||
|
module_path: qlib.data.dataset
|
||||||
|
kwargs:
|
||||||
|
handler:
|
||||||
|
class: TACHandler
|
||||||
|
module_path: tac_qlib.contrib.data.handler
|
||||||
|
kwargs:
|
||||||
|
instruments: "{{ UNIVERSE }}"
|
||||||
|
start_time: 2015-01-03
|
||||||
|
end_time: 2026-08-10
|
||||||
|
fit_start_time: 2016-01-04
|
||||||
|
fit_end_time: 2025-09-01
|
||||||
|
freq: day
|
||||||
|
lake_root: "{{ LAKE }}"
|
||||||
|
market: US
|
||||||
|
label: "Ref($close,-23)/Ref($close,-1)-1"
|
||||||
|
feature_fields: "{{ FEATURES }}"
|
||||||
|
infer_processors:
|
||||||
|
- class: DropAllNaN
|
||||||
|
kwargs: {}
|
||||||
|
- class: ProcessInf
|
||||||
|
kwargs: {}
|
||||||
|
- class: CSRankNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: ZScoreNorm
|
||||||
|
kwargs: {}
|
||||||
|
- class: Fillna
|
||||||
|
kwargs: {}
|
||||||
|
segments:
|
||||||
|
train: [2016-01-04, 2025-09-01]
|
||||||
|
valid: [2025-09-03, 2026-01-03]
|
||||||
|
test: [2026-01-04, 2026-08-10]
|
||||||
|
|
||||||
|
record:
|
||||||
|
- class: SignalRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs: {}
|
||||||
|
- class: SigAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
ana_long_short: true
|
||||||
|
ann_scaler: 252
|
||||||
|
- class: PortAnaRecord
|
||||||
|
module_path: qlib.workflow.record_temp
|
||||||
|
kwargs:
|
||||||
|
config:
|
||||||
|
strategy:
|
||||||
|
class: TopkDropoutStrategy
|
||||||
|
module_path: qlib.contrib.strategy
|
||||||
|
kwargs:
|
||||||
|
signal: "<PRED>"
|
||||||
|
topk: 10
|
||||||
|
n_drop: 1
|
||||||
|
only_tradable: true
|
||||||
|
risk_degree: 0.95
|
||||||
|
backtest:
|
||||||
|
start_time: 2026-01-04
|
||||||
|
end_time: 2026-08-10
|
||||||
|
account: 1000000
|
||||||
|
benchmark: SPY
|
||||||
|
exchange_kwargs:
|
||||||
|
codes: "{{ UNIVERSE }}"
|
||||||
|
deal_price: $close
|
||||||
|
freq: day
|
||||||
|
open_cost: 0.0005
|
||||||
|
close_cost: 0.0015
|
||||||
|
min_cost: 5.0
|
||||||
|
risk_analysis_freq: 1d
|
||||||
Reference in New Issue
Block a user