queue: pre-registered experiment backlog to prove better-trading-performance hypotheses

Mined from book/ on the 'book' branch (HEAD 436692a). 11 queued runs,
each = hypothesis + one-variable change vs the exp-26 reference + acceptance
metric, per the ch.02 isolation/falsification discipline.

- workflows/: 5 runnable config-only YAMLs (Q01 M2 repro, Q02 seed10, Q03 topk20,
  Q04 label10d, Q05 label22d) byte-derived from the exp-26 reference
- designs/: 6 design docs needing custom strategy modules or tool-only A/B
  (Q06 Kelly, Q07 weekly rebalance, Q08 risk-limit A/B, Q09 long-short,
   Q10 HMM overlay, Q11 standalone reversal)
This commit is contained in:
zhaoli
2026-08-19 15:30:46 +00:00
parent 9915a82d38
commit c0eb65efa7
12 changed files with 980 additions and 0 deletions
+33
View File
@@ -0,0 +1,33 @@
# QUEUE-06 — Fractional-Kelly sizing vs equal-weight top-k (re-run exp 15 on clean lake)
**Status:** QUEUED · **Priority:** P1 · **Effort:** custom strategy module + run
## Hypothesis (prove)
Fractional-Kelly sizing — sizing each name by the edge magnitude of its score
instead of equal-weight × risk_degree — is a sizing rule (not a strategy) that
throws away less edge and beats equal-weight top-k **net of costs** on the clean
lake. Source: `book/README.md` open questions (exp 15 run never finished),
`book/references/chat-ideas.md` ("Kelly sizing is a sizing rule, not a strategy").
## Change vs exp-26 reference (ONE variable)
- **Strategy**: equal-weight `TopkDropoutStrategy` (topk 10, n_drop 1) →
custom `FractionalKellyDropoutStrategy` (same topk/n_drop selection, sizing ∝
score magnitude, capped at a fraction f of the equal-weight notional; f as a
parameter, e.g. 0.5).
- All signal/config unchanged (compact stochastic features, 5-seed RankIC
ensemble, 5d label, train/valid/test, SPY benchmark, 5bp/15bp/$5 costs).
## Acceptance
- `net_IR > 0.21` AND `net_ann_return > +2.13%` (exp-26 reference), with
`total_cost` not higher than the reference book.
- If sizing flattens the book (over-concentration) and net degrades → REFUTED
(recorded negative; equal-weight stays canonical).
## Execution prerequisites
1. New contrib module `tac_qlib/contrib/strategy/kelly_dropout.py`
(`FractionalKellyDropoutStrategy` subclassing
`qlib.contrib.strategy.signal_strategy.TopkDropoutStrategy`), copy to the
venv site-packages copy (`/opt/venv/lib/python3.12/site-packages/tac_qlib/...`).
2. Workflow YAML with `strategy.class=FractionalKellyDropoutStrategy`,
`module_path=tac_qlib.contrib.strategy.kelly_dropout`.
3. Trace (rd_trace_start → run → rd_trace_finish), snapshot the new module.
+37
View File
@@ -0,0 +1,37 @@
# QUEUE-07 — Turnover relief: weekly rebalance vs daily (next cost lever after n_drop 1)
**Status:** QUEUED · **Priority:** P1 · **Effort:** custom strategy module + run
## Hypothesis (prove)
n_drop 2→1 proved the cost/turnover frontier is the binding constraint
(EVIDENCE#015, ch.03/ch.09: identical IC/RankIC, net flips −3.21% → +2.13%).
The next lever in the same direction: rebalance the TopkDropout book only
**weekly** (e.g. on Mondays) instead of daily — cutting forced churn further
should lift net performance at the same signal quality.
Source: `book/references/chat-ideas.md` ("weekly rebalance" among the turnover
reduction ideas), ch.09 claim inventory.
## Change vs exp-26 reference (ONE variable)
- **Strategy**: daily TopkDropout (topk 10, n_drop 1) → custom
`WeeklyRebalanceDropoutStrategy` that recomputes the target book once per
week and otherwise holds (no-trade buffer band for small deltas).
- All signal/config unchanged.
## Acceptance
- `total_cost`/turnover strictly below the reference AND `net_IR > 0.21` AND
`net_ann_return > +2.13%`.
- Reference numbers to beat: turnover ~0.74 (round-3 live), est. ~20% daily
book turnover at topk10/n_drop2 (pre-clean-lake estimate).
## Execution prerequisites
1. New contrib module `tac_qlib/contrib/strategy/weekly_rebalance.py`
(`WeeklyRebalanceDropoutStrategy` subclassing `TopkDropoutStrategy`, trade
only when the trade calendar day is the week's first trading day), copy to
the venv site-packages copy.
2. Workflow YAML wiring the strategy.
3. Trace + run + snapshot.
## Sibling (deferred)
No-trade buffer band and notional-vs-qty order sizing are variants of the same
cost lever; queue them only if Q07 reproduces positively.
+34
View File
@@ -0,0 +1,34 @@
# QUEUE-08 — Risk-limit A/B re-validation: $5M liquidity floor on the exp-26 reference
**Status:** QUEUED · **Priority:** P1 · **Effort:** tool-only (no new code)
## Hypothesis (prove)
The $5M liquidity floor improves net IR and cuts drawdown on the **post-reset**
reference signal (pre-reset exp 18, EVIDENCE#008: net IR 0.81→0.98, cumDD
7.93%→5.44%), while size/concentration caps hurt by cutting deployed capital.
Needs re-validation on the exp-26 lineage because exp 18 is pre-clean-lake and
not comparable (EVIDENCE#009/010). Source: `book/CLAIMS.md` open question +
`book/README.md` `TODO(evidence-needed: reconciliation of exp 18 risk-limit spec
on the post-reset reference signal)`.
## Change vs exp-26 reference (ONE variable)
- Reference: the saved exp-26 prediction (run `21afc6af…`, mlflow exp 25).
- A/B via `rd_risk_calibrate` (runs limit-vs-no-limit A/B + sensitivity grid
over size_cap_pct, concentration_cap_pct, liquidity_floor_adv) and/or
`rd_backtest` with `risk_limits` on the SAME saved `pred.pkl`:
- baseline: no limits (this must reproduce the exp-26 net +2.13% / IR 0.21);
- candidate: `{"liquidity_floor_adv": 5000000, "size_cap_pct": 0.12,
"concentration_cap_pct": 0.95, "drawdown_pause_pct": 0.10}` (round-3 spec).
- Pick the spec (B2 calibration) that keeps live ≈ backtest.
## Acceptance
- Candidate spec: `net_IR > 0.21` AND `net_max_drawdown < 7.69%` vs no-limit on
the same pred. Size/concentration caps expected to REDUCE deployed capital
(record the direction as confirmation of exp 18).
- If the floor is a no-op (gates don't bind at this signal) → report that gates
are no-ops when the signal is the bottleneck (exp 20 pattern) as a PROVEN
clean-lake result.
## Execution prerequisites
- None (uses saved pred + `rd_risk_calibrate`/`rd_backtest`). Trace the A/B as
an experiment; record the spec chosen for the next live round.
+30
View File
@@ -0,0 +1,30 @@
# QUEUE-09 — Long-short construction: capture the long-short edge net of costs
**Status:** QUEUED · **Priority:** P2 · **Effort:** custom strategy module + run
## Hypothesis (prove)
The compact stochastic signal's long-short spread is the real edge (L/S ann
Sharpe 4.54, exp 24; "edge is long-short, not long-only" — chat-ideas.md), but
all canonical constructions are long-only (TopkDropout buys topk, drops, holds).
A market-neutral book (long topk, short bottom topk) should realize more of the
spread net of costs than the long-only book, IF short-side financing + doubled
turnover cost stays below the added spread capture.
## Change vs exp-26 reference (ONE variable)
- **Strategy**: long-only TopkDropout (topk 10, n_drop 1) → custom
`TopBottomDropoutStrategy` (long topk by rank, short bottom topk, equal
weight per side, same risk_degree), realized in a workflow with a cost model
that includes both sides (open/close cost symmetric).
- All signal/config unchanged.
## Acceptance
- `net_IR > 0.21` AND `net_ann_return > +2.13%` AND `total_cost` within ~2× the
reference (doubled side count is the structural cost of this construction).
- Watch: benchmark neutrality (SPY beta ≈ 0) as a secondary sanity metric.
## Execution prerequisites
1. New contrib module `tac_qlib/contrib/strategy/top_bottom.py`
(`TopBottomDropoutStrategy` subclassing `BaseSignalStrategy`), copy to the
venv site-packages copy.
2. Workflow YAML wiring the strategy; PortAnaRecord benchmark SPY.
3. Trace + run + snapshot.
+35
View File
@@ -0,0 +1,35 @@
# QUEUE-10 — HMM regime overlay on the exp-26 book (overlay, not feature)
**Status:** QUEUED · **Priority:** P2 · **Effort:** custom strategy + feature compute + run
## Hypothesis (prove)
Regime flags failed as model **features** (exp 9 idea, exp 25 clean-lake
confirmation that model-specific families regress), but the surviving use is as
an **overlay**: a long-only/regime-gate that holds names only in the favourable
HMM state should cut drawdown / improve net IR on the same signal. Source:
`book/chapters/01` regime section + `chat-ideas.md`
(`TODO(evidence-needed: HMM regime gate as overlay on exp-26 book)`).
## Change vs exp-26 reference (ONE variable)
- **Strategy**: plain TopkDropout (topk 10, n_drop 1) → custom
`RegimeGateDropoutStrategy`: identical selection, but when the per-symbol
HMM posterior (`sp_hmm_p_regime1`) is below a calibrated threshold the name
is held in cash instead of bought (entry gate); no new features enter the
model — `sp_hmm_p_regime1` is computed for gating only, fit on the train
window (no lookahead), via `get_lake_sp` with `fit_end=<train end>`.
- All signal/config unchanged.
## Acceptance
- `net_max_drawdown < 7.69%` (reference) AND `net_IR >= 0.21`. If the gate
never binds at a sensible threshold → the gate is a no-op on this signal
(exp 20 pattern) → recorded REFUTED/neutral, not a failure.
- Calibrate the threshold on the valid window only (avoid the exp 13/14
threshold-overfit trap).
## Execution prerequisites
1. Persist `sp_hmm_p_regime1` for the universe (get_lake_sp, fit_end =
2025-09-01) WITHOUT adding it to `feature_fields` of the model.
2. New contrib module `tac_qlib/contrib/strategy/regime_gate.py`, copy to the
venv site-packages copy.
3. Workflow YAML wiring the strategy.
4. Trace + run + snapshot.
+32
View File
@@ -0,0 +1,32 @@
# QUEUE-11 — Standalone 5-day reversal signal net of costs (unisolated)
**Status:** QUEUED · **Priority:** P2 · **Effort:** dataset study + backtest
## Hypothesis (prove)
5-day momentum strongly reverses on this panel (pooled regression:
`sp_trend_slope_5` β = −0.53, t = −24; VR < 1 at 5–20d for ~32/72 assets —
chat-derived, pre-clean-lake idea material). The reversal has never been tested
as a **standalone tradable strategy net of costs**. If it clears the 20bp
round-trip cost, it is an independent alpha source that can be blended with (or
replace) the model book.
Source: `book/ch01` "Timeline" + `chat-ideas.md`
(`TODO(evidence-needed: standalone 5d-reversal strategy net of costs)`).
## Change vs exp-26 reference
- This is NOT a model-construction variant — it isolates a SINGLE-FEATURE
signal: a model trained on `sp_trend_slope_5` (plus raw OHLCV) alone, or a
mechanical reversal book (rank by −`sp_trend_slope_5`, buy the most-reverted
topk), backtested net of costs over the exp-26 window.
- Control: exp-26 compact reference on the same window.
## Acceptance
- Standalone reversal `net_ann_return > 0` (clears 20bp round-trip) — proves
the claim "reversal is tradable net of costs". Secondary: excess vs the
model book is the blend decision for a future round.
## Execution prerequisites
1. `rd_train`/workflow with `feature_fields = $open,$high,$low,$close,$vwap,$volume,sp_trend_slope_5`
(single feature) OR a mechanical rank backtest via `rd_backtest` on a
hand-built pred (pred = −rank(sp_trend_slope_5)).
2. Trace + run + record as a standalone study (dataset-study status, not
necessarily a traced model experiment).