queue: pre-register Series 2 (Q12-Q20) targeting unproven book hypotheses

Series 1 (Q01-Q11, exp 33-43) executed and folded into book/CLAIMS.md.
Series 2 covers the remaining HYPOTHESIS rows and open questions:
  Q12 22d label + weekly recompute (untested combo)
  Q13 weekly rebalance reproduction on a 2nd OOS window
  Q14 out-of-universe validation (single-stock panel, needs lake backfill)
  Q15 5-seed vs single-seed clean A/B
  Q16 hmm family as features
  Q17 realized-moments family as features
  Q18 OptimalStopControl clean re-test
  Q19/Q20 martingale-VR + effective-names scripted studies

Each workflow pins one-variable change vs exp-26 reference and acceptance.
This commit is contained in:
zhaoli
2026-08-20 03:51:58 +00:00
parent c0eb65efa7
commit 7124ef5d8e
17 changed files with 502 additions and 627 deletions
-33
View File
@@ -1,33 +0,0 @@
# QUEUE-06 — Fractional-Kelly sizing vs equal-weight top-k (re-run exp 15 on clean lake)
**Status:** QUEUED · **Priority:** P1 · **Effort:** custom strategy module + run
## Hypothesis (prove)
Fractional-Kelly sizing — sizing each name by the edge magnitude of its score
instead of equal-weight × risk_degree — is a sizing rule (not a strategy) that
throws away less edge and beats equal-weight top-k **net of costs** on the clean
lake. Source: `book/README.md` open questions (exp 15 run never finished),
`book/references/chat-ideas.md` ("Kelly sizing is a sizing rule, not a strategy").
## Change vs exp-26 reference (ONE variable)
- **Strategy**: equal-weight `TopkDropoutStrategy` (topk 10, n_drop 1) →
custom `FractionalKellyDropoutStrategy` (same topk/n_drop selection, sizing ∝
score magnitude, capped at a fraction f of the equal-weight notional; f as a
parameter, e.g. 0.5).
- All signal/config unchanged (compact stochastic features, 5-seed RankIC
ensemble, 5d label, train/valid/test, SPY benchmark, 5bp/15bp/$5 costs).
## Acceptance
- `net_IR > 0.21` AND `net_ann_return > +2.13%` (exp-26 reference), with
`total_cost` not higher than the reference book.
- If sizing flattens the book (over-concentration) and net degrades → REFUTED
(recorded negative; equal-weight stays canonical).
## Execution prerequisites
1. New contrib module `tac_qlib/contrib/strategy/kelly_dropout.py`
(`FractionalKellyDropoutStrategy` subclassing
`qlib.contrib.strategy.signal_strategy.TopkDropoutStrategy`), copy to the
venv site-packages copy (`/opt/venv/lib/python3.12/site-packages/tac_qlib/...`).
2. Workflow YAML with `strategy.class=FractionalKellyDropoutStrategy`,
`module_path=tac_qlib.contrib.strategy.kelly_dropout`.
3. Trace (rd_trace_start → run → rd_trace_finish), snapshot the new module.
-37
View File
@@ -1,37 +0,0 @@
# QUEUE-07 — Turnover relief: weekly rebalance vs daily (next cost lever after n_drop 1)
**Status:** QUEUED · **Priority:** P1 · **Effort:** custom strategy module + run
## Hypothesis (prove)
n_drop 2→1 proved the cost/turnover frontier is the binding constraint
(EVIDENCE#015, ch.03/ch.09: identical IC/RankIC, net flips −3.21% → +2.13%).
The next lever in the same direction: rebalance the TopkDropout book only
**weekly** (e.g. on Mondays) instead of daily — cutting forced churn further
should lift net performance at the same signal quality.
Source: `book/references/chat-ideas.md` ("weekly rebalance" among the turnover
reduction ideas), ch.09 claim inventory.
## Change vs exp-26 reference (ONE variable)
- **Strategy**: daily TopkDropout (topk 10, n_drop 1) → custom
`WeeklyRebalanceDropoutStrategy` that recomputes the target book once per
week and otherwise holds (no-trade buffer band for small deltas).
- All signal/config unchanged.
## Acceptance
- `total_cost`/turnover strictly below the reference AND `net_IR > 0.21` AND
`net_ann_return > +2.13%`.
- Reference numbers to beat: turnover ~0.74 (round-3 live), est. ~20% daily
book turnover at topk10/n_drop2 (pre-clean-lake estimate).
## Execution prerequisites
1. New contrib module `tac_qlib/contrib/strategy/weekly_rebalance.py`
(`WeeklyRebalanceDropoutStrategy` subclassing `TopkDropoutStrategy`, trade
only when the trade calendar day is the week's first trading day), copy to
the venv site-packages copy.
2. Workflow YAML wiring the strategy.
3. Trace + run + snapshot.
## Sibling (deferred)
No-trade buffer band and notional-vs-qty order sizing are variants of the same
cost lever; queue them only if Q07 reproduces positively.
-34
View File
@@ -1,34 +0,0 @@
# QUEUE-08 — Risk-limit A/B re-validation: $5M liquidity floor on the exp-26 reference
**Status:** QUEUED · **Priority:** P1 · **Effort:** tool-only (no new code)
## Hypothesis (prove)
The $5M liquidity floor improves net IR and cuts drawdown on the **post-reset**
reference signal (pre-reset exp 18, EVIDENCE#008: net IR 0.81→0.98, cumDD
7.93%→5.44%), while size/concentration caps hurt by cutting deployed capital.
Needs re-validation on the exp-26 lineage because exp 18 is pre-clean-lake and
not comparable (EVIDENCE#009/010). Source: `book/CLAIMS.md` open question +
`book/README.md` `TODO(evidence-needed: reconciliation of exp 18 risk-limit spec
on the post-reset reference signal)`.
## Change vs exp-26 reference (ONE variable)
- Reference: the saved exp-26 prediction (run `21afc6af…`, mlflow exp 25).
- A/B via `rd_risk_calibrate` (runs limit-vs-no-limit A/B + sensitivity grid
over size_cap_pct, concentration_cap_pct, liquidity_floor_adv) and/or
`rd_backtest` with `risk_limits` on the SAME saved `pred.pkl`:
- baseline: no limits (this must reproduce the exp-26 net +2.13% / IR 0.21);
- candidate: `{"liquidity_floor_adv": 5000000, "size_cap_pct": 0.12,
"concentration_cap_pct": 0.95, "drawdown_pause_pct": 0.10}` (round-3 spec).
- Pick the spec (B2 calibration) that keeps live ≈ backtest.
## Acceptance
- Candidate spec: `net_IR > 0.21` AND `net_max_drawdown < 7.69%` vs no-limit on
the same pred. Size/concentration caps expected to REDUCE deployed capital
(record the direction as confirmation of exp 18).
- If the floor is a no-op (gates don't bind at this signal) → report that gates
are no-ops when the signal is the bottleneck (exp 20 pattern) as a PROVEN
clean-lake result.
## Execution prerequisites
- None (uses saved pred + `rd_risk_calibrate`/`rd_backtest`). Trace the A/B as
an experiment; record the spec chosen for the next live round.
-30
View File
@@ -1,30 +0,0 @@
# QUEUE-09 — Long-short construction: capture the long-short edge net of costs
**Status:** QUEUED · **Priority:** P2 · **Effort:** custom strategy module + run
## Hypothesis (prove)
The compact stochastic signal's long-short spread is the real edge (L/S ann
Sharpe 4.54, exp 24; "edge is long-short, not long-only" — chat-ideas.md), but
all canonical constructions are long-only (TopkDropout buys topk, drops, holds).
A market-neutral book (long topk, short bottom topk) should realize more of the
spread net of costs than the long-only book, IF short-side financing + doubled
turnover cost stays below the added spread capture.
## Change vs exp-26 reference (ONE variable)
- **Strategy**: long-only TopkDropout (topk 10, n_drop 1) → custom
`TopBottomDropoutStrategy` (long topk by rank, short bottom topk, equal
weight per side, same risk_degree), realized in a workflow with a cost model
that includes both sides (open/close cost symmetric).
- All signal/config unchanged.
## Acceptance
- `net_IR > 0.21` AND `net_ann_return > +2.13%` AND `total_cost` within ~2× the
reference (doubled side count is the structural cost of this construction).
- Watch: benchmark neutrality (SPY beta ≈ 0) as a secondary sanity metric.
## Execution prerequisites
1. New contrib module `tac_qlib/contrib/strategy/top_bottom.py`
(`TopBottomDropoutStrategy` subclassing `BaseSignalStrategy`), copy to the
venv site-packages copy.
2. Workflow YAML wiring the strategy; PortAnaRecord benchmark SPY.
3. Trace + run + snapshot.
-35
View File
@@ -1,35 +0,0 @@
# QUEUE-10 — HMM regime overlay on the exp-26 book (overlay, not feature)
**Status:** QUEUED · **Priority:** P2 · **Effort:** custom strategy + feature compute + run
## Hypothesis (prove)
Regime flags failed as model **features** (exp 9 idea, exp 25 clean-lake
confirmation that model-specific families regress), but the surviving use is as
an **overlay**: a long-only/regime-gate that holds names only in the favourable
HMM state should cut drawdown / improve net IR on the same signal. Source:
`book/chapters/01` regime section + `chat-ideas.md`
(`TODO(evidence-needed: HMM regime gate as overlay on exp-26 book)`).
## Change vs exp-26 reference (ONE variable)
- **Strategy**: plain TopkDropout (topk 10, n_drop 1) → custom
`RegimeGateDropoutStrategy`: identical selection, but when the per-symbol
HMM posterior (`sp_hmm_p_regime1`) is below a calibrated threshold the name
is held in cash instead of bought (entry gate); no new features enter the
model — `sp_hmm_p_regime1` is computed for gating only, fit on the train
window (no lookahead), via `get_lake_sp` with `fit_end=<train end>`.
- All signal/config unchanged.
## Acceptance
- `net_max_drawdown < 7.69%` (reference) AND `net_IR >= 0.21`. If the gate
never binds at a sensible threshold → the gate is a no-op on this signal
(exp 20 pattern) → recorded REFUTED/neutral, not a failure.
- Calibrate the threshold on the valid window only (avoid the exp 13/14
threshold-overfit trap).
## Execution prerequisites
1. Persist `sp_hmm_p_regime1` for the universe (get_lake_sp, fit_end =
2025-09-01) WITHOUT adding it to `feature_fields` of the model.
2. New contrib module `tac_qlib/contrib/strategy/regime_gate.py`, copy to the
venv site-packages copy.
3. Workflow YAML wiring the strategy.
4. Trace + run + snapshot.
-32
View File
@@ -1,32 +0,0 @@
# QUEUE-11 — Standalone 5-day reversal signal net of costs (unisolated)
**Status:** QUEUED · **Priority:** P2 · **Effort:** dataset study + backtest
## Hypothesis (prove)
5-day momentum strongly reverses on this panel (pooled regression:
`sp_trend_slope_5` β = −0.53, t = −24; VR < 1 at 5–20d for ~32/72 assets —
chat-derived, pre-clean-lake idea material). The reversal has never been tested
as a **standalone tradable strategy net of costs**. If it clears the 20bp
round-trip cost, it is an independent alpha source that can be blended with (or
replace) the model book.
Source: `book/ch01` "Timeline" + `chat-ideas.md`
(`TODO(evidence-needed: standalone 5d-reversal strategy net of costs)`).
## Change vs exp-26 reference
- This is NOT a model-construction variant — it isolates a SINGLE-FEATURE
signal: a model trained on `sp_trend_slope_5` (plus raw OHLCV) alone, or a
mechanical reversal book (rank by −`sp_trend_slope_5`, buy the most-reverted
topk), backtested net of costs over the exp-26 window.
- Control: exp-26 compact reference on the same window.
## Acceptance
- Standalone reversal `net_ann_return > 0` (clears 20bp round-trip) — proves
the claim "reversal is tradable net of costs". Secondary: excess vs the
model book is the blend decision for a future round.
## Execution prerequisites
1. `rd_train`/workflow with `feature_fields = $open,$high,$low,$close,$vwap,$volume,sp_trend_slope_5`
(single feature) OR a mechanical rank backtest via `rd_backtest` on a
hand-built pred (pred = −rank(sp_trend_slope_5)).
2. Trace + run + record as a standalone study (dataset-study status, not
necessarily a traced model experiment).
+26
View File
@@ -0,0 +1,26 @@
# QUEUE-19 — Martingale / variance-ratio study close-out (no qrun)
**Status:** QUEUED · **Priority:** P2 · **Effort:** ad-hoc script under `book/data/`
## Hypothesis (settle)
Assets are submartingales long-horizon / mean-reverting short-horizon
(`VR < 1` at 5–20d). CLAIMS.md marks this HYPOTHESIS (chat-derived martingale
study; exp 19 was opened but never closed). It is a market-structure claim, not a
trading claim — settle it with a clean-lake script, then close exp 19 or open a
scripted EVIDENCE entry.
## Method (persist everything under `book/data/evidence/q19-vr/`)
1. Load the 50-ETF panel 1d bars from the lake for 2015-01-01..2026-08-19.
2. Compute the Lo–MacKinlay variance ratio at horizons 5 / 10 / 20d per symbol,
with heteroskedasticity-robust z-stats.
3. Report: per-horizon VR distribution, fraction of symbols with VR < 1 and the
z-significance, pooled drift vs daily variance (submartingale check).
4. Cross-check the pooled `sp_trend_slope_5` regression beta claim (β ≈ −0.53,
t ≈ −24) on the clean lake.
5. Write `VR_stats.csv` + a one-page summary into the evidence dir.
## Acceptance
- VR < 1 at 5–20d for a material fraction of the panel with |z| > 2 → supports
the mean-reversion HYPOTHESIS; else mark REFUTED or REFERENCED.
- The result updates CLAIMS.md's "Assets are submartingales…" row and closes the
exp-19 open thread.
+22
View File
@@ -0,0 +1,22 @@
# QUEUE-20 — Effective independent names in the 50-ETF book (no qrun)
**Status:** QUEUED · **Priority:** P2 · **Effort:** ad-hoc script under `book/data/`
## Hypothesis (settle)
The 50-ETF book has only ~4 effective independent names (CLAIMS.md HYPOTHESIS,
chat-derived eigenvalue analysis, pre-reset). This is a concentration/diversification
claim with direct sizing relevance; verify it on the clean lake.
## Method (persist everything under `book/data/evidence/q20-effective-names/`)
1. Load the 50-ETF panel 1d returns from the lake for the test window 2026-01-04..2026-08-10.
2. Standardize returns; compute the correlation matrix and its eigendecomposition.
3. Count eigenvalues above the Marchenko–Pastur bound (N=50, T≈150) and report the
cumulative-variance share of the top k components.
4. Effective-rank measures: participation ratio `(Σλ)² / Σλ²` and cumulative 80%
variance count.
5. Write `eigenanalysis.csv` + a one-page summary.
## Acceptance
- If effective rank ≈ 4 (top-4 explain ~80%+ variance), the concentration claim is
PROVEN and feeds chapter 08 sizing guidance (why topk 10→20 adds no breadth).
- If effective rank is much larger, mark the claim REFUTED.