Files
book-tac/tac-qlib/skills/tac-qlib-custom/SKILL.md
T

530 lines
30 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: tac-qlib-custom
description: "Guide agents to customize and extend Qlib on the TradeAC R&D stack — how to configure workflow YAMLs (qlib_init, model, dataset/handler, processors, records, PortAnaRecord strategies), how to extend Qlib classes wired into those workflows (custom Model, BaseStrategy, DataHandler, Record), and the empirically-tested knobs from this repo (RankIC early-stopping, stochastic-control strategies, stochastic-process features, catch22/GARCH/Hurst/signature). Also encodes the experiment traceability loop: every backtest runs as a workflow-with-recorder, is recorded in the Postgres experiments table (rationale/details/evaluation/metrics with pgvector embeddings, evolution chain) and on a per-experiment git branch that is committed + pushed. Companion to tradeac-rd (MCP run tools) and tradeac-lake (parquet lake)."
---
# tac-qlib-custom
Customizing and extending Qlib on the TradeAC stack. This skill encodes what was
learned from actual experiments in this repo: how a workflow YAML maps to Qlib
classes, how to write a custom class that the YAML can load, and which training /
strategy / feature knobs measurably moved IC, RankIC and the backtest.
Read `tac-qlib/skills/tradeac-rd/SKILL.md` for the MCP run/inspect tools and
`tac-qlib/README.md` for the package layout. The venv is `/app/.venv`
(qlib 0.1.dev2066); `tac_qlib` is installed into the venv's `site-packages`
(editable copy under `/opt/venv/.../tac_qlib/`), so **any new module must be
copied to `/opt/venv/lib/python3.12/site-packages/tac_qlib/...` too** (or use an
editable install) before `rd_run_workflow` can import it.
## MCP-first policy
- **Drive every backtest and run through the `tac-qlib-rd` MCP tools** (`rd_run_workflow`,
`rd_train`, `rd_predict`, `rd_exp_*`) and the tac-engine lake tools for data prep. Do not
reimplement them with ad-hoc scripts (custom qlib glue, own mlruns readers, direct
JSON-RPC/stdio clients).
- **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server` /
`tac-engine`, bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and
ask the user to confirm first**.
- The traceability bookkeeping (Postgres `rd_experiments` row + pgvector embeddings +
branch-per-experiment git) is exposed as the **`rd_trace_*` MCP tools** on the tac-qlib-rd
server — use those, not bash scripts. Data prep, training, evaluation and backtests also go
through MCP tools.
- If the venv is missing a runtime dep (`duckdb`, `pyarrow`, feature libs), lazy-install it
(`uv pip install --python $VIRTUAL_ENV/bin/python <pkg>`) instead of switching tools.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or
credential-bearing URLs (`DATABASE_URL`, `GIT_PASS`, `EMBEDDING_API_KEY`) in
workflow YAMLs, scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on
`.env`). That pulls secrets into this session and leaks them to any agent
sharing it.
- When a tool or command needs an env var, ASK the user to set it in the
environment (shell/container env, or the user-owned `.env`) and reference it
by name (`$VAR`), never by value. If it's missing, report which variable is
required instead of reading it yourself.
- Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows —
`rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else
the lake sqlite. Never hardcode a `postgres://user:pass@…` URI.
- If you find a committed secret, flag it, remove it, and replace it with a
placeholder. (The `rd_trace_*` MCP tools' commit guard blocks adding
credential-shaped lines.)
## How a workflow YAML maps to Qlib classes
A workflow YAML (`tac-qlib/workflows/*.yaml`) is rendered by Jinja (vars like
`{{ LAKE }}` from `TAC_LAKE_DIR`) then executed by `qrun` / `rd_run_workflow`.
Every block is a Qlib class reference resolved by `module_path` + `class`:
```yaml
{%- set LAKE = TAC_LAKE_DIR %}
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
calendar_provider: # custom tac-qlib providers read the parquet lake
class: LakeCalendarProvider
module_path: tac_qlib.data.providers
instrument_provider: # ... (markets: {} => lake universe)
feature_provider: # LakeFeatureProvider: routes $open..$volume from bars,
class: LakeFeatureProvider # $<ta-lib/sp_*> from features parquet, $amount derived
exp_manager:
class: MLflowExpManager
module_path: qlib.workflow.expm
kwargs: { uri: "sqlite:///{{ LAKE }}/mlruns.db", default_exp_name: "my-exp" }
task:
model: # <MODEL BLOCK> — custom model → new module_path
class: RankICLGBModel
module_path: tac_qlib.contrib.model.rank_gbdt
kwargs: { loss: mse, learning_rate: 0.02, num_leaves: 31, ... }
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler: # <HANDLER BLOCK> — feature selection + processors live here
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: "SPY,QQQ,..."
start_time: 2015-01-03
end_time: 2026-08-10
fit_start_time: 2015-01-03 # processors fit on this window
fit_end_time: 2025-09-01
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-6)/Ref($close,-1)-1" # 5d forward return
feature_fields: "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_ou_zscore,..."
infer_processors: # feature-time transforms, fit on fit_*
- { class: DropAllNaN, kwargs: {} }
- { class: ProcessInf, kwargs: {} }
- { class: CSRankNorm, kwargs: {} } # per-day cross-sectional rank
- { class: ZScoreNorm, kwargs: {} }
- { class: Fillna, kwargs: {} }
segments:
train: [2015-01-03, 2025-09-01]
valid: [2025-09-03, 2026-01-03]
test: [2026-01-04, 2026-08-10]
record: # each entry records one artifact type to the run
- { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} }
- { class: SigAnaRecord, module_path: qlib.workflow.record_temp,
kwargs: { ana_long_short: true, ann_scaler: 252 } }
- { class: PortAnaRecord, module_path: qlib.workflow.record_temp,
kwargs: { config: { strategy: <STRATEGY BLOCK>, backtest: {...} }, risk_analysis_freq: 1d } }
```
`rd_run_workflow config_path=<yaml> experiment_name=<exp>` runs it; the MCP call
may time out for long runs (RankIC tuning, heavy feature sets) — the run keeps
executing; poll via `rd_exp_list` / `rd_exp_get_run` on the returned experiment.
## Experiment traceability (DB + git + embeddings)
Every backtest you run as an agent MUST be tracked: it runs as a workflow with the
`record` block (SignalRecord/SigAnaRecord/PortAnaRecord → MLflow artifacts on disk
under `<lake>/mlruns/<exp_id>/<run_id>`), and a row is written to the Postgres
`experiments` table plus a git branch per experiment. The `tac-app` UI owns the
schema (Drizzle migrations in `tac-app/drizzle/`); this skill's `lib/` scripts are
the executor the agent drives.
**Trigger the lineage as part of the run — automatically, not on prompt.** Any
time you execute a qlib workflow (`rd_run_workflow`) or a train/predict pipeline
on this stack, the traceability bookkeeping is part of that run, not a separate
step the user must ask for: open the traced experiment with `rd_trace_start`
before running, commit intermediates with `rd_trace_commit`, and close it with
`rd_trace_finish` after — without waiting to be prompted (see "The
per-experiment procedure" below).
### Env vars
| Var | Purpose |
|-----|---------|
| `DATABASE_URL` | Postgres URL for the `rd_experiments` table AND the MLflow tracking store (set in repo `.env`) |
| `EMBEDDING_API_BASE_URL` | embedding POST endpoint (e.g. `https://embd.h.lizhao.net/embeddings`) |
| `EMBEDDING_API_KEY` | basic-auth credential (`user:pass` form is supported) |
| `GIT_USER` / `GIT_PASS` | git remote credentials for push/fetch |
| `GIT_REPO_URL` | experiment git repo tracked by the `experiments` submodule (branches are pushed here) |
| `TAC_LAKE_DIR` | lake root (mlruns artifact files live under it) |
The experiment repo is the **`experiments` git submodule** at the workspace root
(`<repo-root>/experiments`), always tracking `$GIT_REPO_URL`. `rd_trace_init`
creates/validates it; it errors if `experiments/` exists but points at a
different URL. There is no `TAC_EXP_GIT_DIR` — the submodule path IS the
experiment repo, and ALL experiment/backtest changes (workflow YAMLs, notes,
outputs) must live inside it, never in the parent tradeac repo.
### The `rd_experiments` table
Owned by tac-app's Drizzle schema (`tac-app/src/db/schema.ts`); `rd_trace_init`
can `init` it idempotently. The table is named **`rd_experiments`** (NOT
`experiments`) because MLflow's Postgres tracking store creates its own
`experiments` table in the same database. Key columns: `id` (PK), `rational` +
`rational_embedding` (pgvector `vector(384)`), `details` + `details_embedding`,
`evaluation`, `metrics` (jsonb), `evolved_from` (FK → rd_experiments.id),
`start_ts`/`end_ts`, `git_branch`, `experiment_ref_id`, `mlruns_dir`, `status`.
`experiment_ref_id` holds the **mlflow run id** returned by `rd_run_workflow` and
is an FK to MLflow's `runs(run_uuid)` (added by `rd_trace_init` after the
mlflow store tables exist — MLflow creates `runs` lazily).
Tracking store: **Postgres `$DATABASE_URL`** (MLflow's own tables) when set,
falling back to the unified lake sqlite `sqlite:///<lake>/mlruns.db`. Artifact
files always stay on disk under `<lake>/mlruns/<exp_id>/<run_id>/artifacts`.
Embedding model: `michaelfeil/bge-small-en-v1.5` (384-dim, **512-token context**).
Rational/details are written paper-summary style (≤512 tokens) and embedded verbatim —
NEVER truncate; if a text is longer, summarize it first (the embed helper rejects
over-limit input).
### Git repo + branch-per-experiment
The experiment repo is the `experiments` submodule at the workspace root
(`<repo-root>/experiments`, tracking `$GIT_REPO_URL`). The `rd_trace_*` MCP
tools handle it, and every git operation is scoped to that submodule —
experiments NEVER stage or push parent-repo (tradeac) files.
- `rd_trace_init` creates/validates the submodule and the base branch. If
`experiments/` does not exist it runs `git clone $GIT_REPO_URL experiments`;
if it exists but tracks a different URL, init errors out.
- Base branch: `main` (or `master`). If the submodule is empty, a seed commit is
made and pushed so there are commits to fork from.
- Every experiment runs on its own branch `exp/<id>-<slug>`.
- `evolved_from` resolution (in order):
1. If the wizard prompt explicitly says `evolved_from=<id>` (run wizard click on an
existing experiment) — use that id directly.
2. Otherwise `--evolved-from auto`: the user prompt / rational is embedded and
cosine-searched over the `experiments.rational_embedding` column; the top hit
above the similarity threshold (0.5) becomes `evolved_from`.
3. Otherwise (first experiment, or a new chat with no predecessor) — no evolved_from;
fork from `main`'s latest commits.
- The new branch is forked from the **evolved-from experiment's branch** (its latest
commits), or from `main` when there is no predecessor — so experiment lineages form
a git branch chain.
- On every finish, and for intermediate steps, changes are committed + pushed.
### Custom code is part of the lineage (code snapshot)
Custom contrib modules (`tac_qlib/contrib/model/`, `tac_qlib/contrib/strategy/`,
`tac_qlib/contrib/data/`, `tac_qlib/data/providers.py`) live in the **parent**
tradeac repo, not in the `experiments/` submodule — so they are normally invisible
to the experiment branch and a descendant forking from it would reinvent them.
The lineage tooling fixes this: **every experiment branch carries a `code/`
snapshot of exactly the qlib extension code that run depended on**, so descendants
reuse it instead of re-authoring it.
- `rd_trace_start` and `rd_trace_finish` automatically snapshot the default paths
(`tac-qlib/tac_qlib/contrib`, `tac-qlib/tac_qlib/data`) into
`<experiments>/code/<parent-relative-path>` on the experiment branch.
- `rd_trace_snapshot` snapshots mid-run (e.g. after writing a
new custom model) without waiting for finish.
- The snapshot also writes `code/MANIFEST.txt` recording the **parent-repo HEAD
commit** and the per-file blob hashes it was taken from — so a run can be traced
back to the exact parent commit that produced its custom code.
- Descendants: the custom modules your run needs are under `code/tac_qlib/...` on the
evolved-from branch. Reuse them (copy/`git show`) instead of writing new ones; check
`code/MANIFEST.txt` to see which parent commit they came from and port fixes back.
- Guardrail exception: parent-repo changes under `tac_qlib/tac_qlib/contrib` and
`tac_qlib/tac_qlib/data` are **expected** (they are the snapshotted code);
`parent_changes` reports them as a note, not a violation. Any OTHER parent change
is still a guardrail violation.
Guardrail — experiments must NOT introduce side effects to the parent repo:
- Write workflow YAMLs, notes and experiment outputs ONLY inside
`<repo-root>/experiments/` (they are committed on the experiment branch).
- Never `git add`/commit/stage anything in the parent tradeac repo.
- Run `rd_trace_guard` to list any parent
changes outside the submodule pointer; `rd_trace_finish` also surfaces them.
Revert any accidental parent edits before finishing.
- If an experiment reveals a PRODUCT change (workflow template, skill, tac-app),
propose it separately for the tradeac repo — do not mix it into the experiment
branch.
The `rd_trace_*` MCP tools perform git operations with the mandated credential
helper (from `GIT_USER` / `GIT_PASS`), so you do not need to construct it by hand.
### The per-experiment procedure
**Use the `rd_trace_*` MCP tools (tac-qlib-rd)** — they replace the old
`trace.sh`/`trace_db.py` scripts. The server is long-lived (psycopg imported
once, DB connection reused per call) and every tool returns one JSON object, so
no output parsing is needed:
```text
# 0. ensure ready (rd_experiments table + experiments git repo + base main)
rd_trace_init
# 1. start — inserts the row, resolves evolved_from, forks+pushes the branch.
# Returns {experiment_id, branch, evolved_from, base_branch} as JSON.
rd_trace_start rational="5-day forward label, RankIC early stop, 50-ETF universe" \
details="LGBModel mse lr=0.02 num_leaves=15 num_boost_round=3000; TopkDropout topk=2; benchmark QQQ" \
experiment_name="tac-rd-expN" \
evolved_from="auto" \
session_id="<this chat's opencode session id, if started from a chat>"
# -> {"experiment_id": N, "branch": "exp/N-...", "evolved_from": ..., "base_branch": ...}
# 2. write the workflow YAML INSIDE the experiments submodule
# (e.g. <repo-root>/experiments/workflows/<exp>/workflow.yaml), then commit it:
rd_trace_commit experiment_id=<N> message="add workflow yaml"
# 2b. if the workflow uses a NEW custom module, snapshot it onto the branch
# (start/finish auto-snapshot contrib+data; do this to capture mid-run):
rd_trace_snapshot experiment_id=<N> # default contrib+data
# or: rd_trace_snapshot experiment_id=<N> paths="tac-qlib/tac_qlib/contrib/model/rank_gbdt.py"
# 3. run the backtest through the WORKFLOW with the recorder (MUST write mlruns):
rd_run_workflow config_path=<repo-root>/experiments/workflows/<exp>/workflow.yaml experiment_name=tac-rd-expN
# -> returns run_id (= experiment_ref_id) + metrics
# 4. inspect with rd_exp_result / rd_exp_blotter, then finish — updates the row
# (re-embeds rational/details, sets metrics/eval/end_ts), snapshots the custom
# code, and commits+pushes. finish also surfaces parent-repo side effects.
rd_trace_finish experiment_id=<N> \
ref_id=<mlflow-run-id> \
evaluation="IC 0.0645, RankIC 0.075; net excess +0.85% ann" \
metrics='{"IC":0.0645,"RankIC":0.075,"ann_excess":0.85}' \
mlruns_dir=<lake>/mlruns/<exp_id>/<run_id>
```
Helpers (MCP tools): `rd_trace_search` (semantic), `rd_trace_get` (one row),
`rd_trace_list`, `rd_trace_mlruns_dir` (resolves the mlruns dir for an
experiment name), `rd_trace_guard` (parent-repo side-effect check).
Rules:
- **Always** run backtests as workflows with the `record` block (req 2) — never a bare
`rd_backtest` for a traced experiment.
- **Always** open the lineage (`rd_trace_start`) BEFORE the run and **Always**
`rd_trace_finish` + push after it completes (req 5) — this happens as part of the run,
do not wait for the user to ask; intermediate `rd_trace_commit` is encouraged (req 5).
- **Always** snapshot the custom qlib code (`rd_trace_snapshot`, or rely on the
auto-snapshot at start/finish) so the experiment branch carries the exact contrib/data
modules the run used — descendants fork and reuse `code/` instead of reinventing it.
- Keep rational/details ≤ 512 tokens (paper-summary style) so embeddings are exact —
no truncation.
- **Confine experiments to the `experiments/` submodule** — never write to, stage, or
commit parent tradeac repo files; run `rd_trace_guard` to check for side effects.
(Custom code edits under `tac-qlib/tac_qlib/contrib` and `.../data` are the sanctioned
exception — they are the snapshotted modules; see "Custom code is part of the lineage".)
- **Follow the Secrets policy above** — no secrets in files, no reading `.env*`, ask the
user to set env vars; use `uri: "sqlite:///mlruns.db"` for the tracking store.
- Workflow YAMLs are jinja-rendered with `os.environ` as the context, so env-var
placeholders work (`{%- set LAKE = TAC_LAKE_DIR %}` then `{{ LAKE }}`). Use them for
paths/config — never for secrets that get committed.
## Extending Qlib — the 4 class families you can override
### 1. Custom Model (train-time) — `tac_qlib/contrib/model/`
Subclass `qlib.contrib.model.gbdt.LGBModel` (or `qlib.model.base.BaseModel`) and
implement `fit(dataset, ...)` + `predict(dataset)`. `LGBModel.fit` calls
`self._prepare_data(dataset)` → `lgb.Dataset`s, then `lgb.train` with
`early_stopping` on the valid set. Override points that matter:
- `_prepare_data` → build the `lgb.Dataset` with `group=` (per-day query groups)
when you need ranking metrics per trading day.
- `fit` → change what early-stops training (the biggest IC/backtest lever, see §Knobs).
- `predict` → return the Series keyed (datetime, instrument).
Reference: `tac_qlib/tac_qlib/contrib/model/rank_gbdt.py` — `RankICLGBModel`
subclasses `LGBModel`, adds per-day `group` in `_prepare_data`, injects
`feval=rankic_feval` (mean per-day Spearman) into `lgb.train`, and forces
`metric='None'` + `first_metric_only=True` so early-stopping tracks RankIC only.
### 2. Custom Strategy (backtest-time) — `tac_qlib/contrib/strategy/`
Subclass `qlib.contrib.strategy.signal_strategy.BaseSignalStrategy` (which wraps
`qlib.strategy.base.BaseStrategy`) and implement:
```python
def generate_trade_decision(self, execute_result=None):
# trade_step, trade_start/end = self.trade_calendar.get_step_time(trade_step)
# pred = self.signal.get_signal(start_time=pred_shift, end_time=pred_shift) # shift=-1 => signal known at t-1
# self.trade_position / self.trade_exchange / self.trade_calendar injected by the executor
# build qlib.backtest.Order(stock_id, amount, start_time, end_time, direction=Order.BUY/SELL)
# return TradeDecisionWO(orders, self)
```
Wire it into the YAML under `PortAnaRecord.config.strategy`:
```yaml
strategy:
class: OptimalStopControl
module_path: tac_qlib.contrib.strategy.optimal_stop
kwargs:
signal: "<PRED>" # placeholder replaced with the recorded pred
topk: 10
entry_pct: 0.85
exit_pct: 0.7
max_hold_days: 10
min_hold_days: 2
sl: -0.08
risk_degree: 0.95
```
Reference: `tac_qlib/tac_qlib/contrib/strategy/optimal_stop.py`
(`OptimalStopControl` — entry gated by cross-sectional signal percentile, exits
by percentile/time/stop-loss, equal-weight control sizing).
### 3. Custom DataHandler / processors — `tac_qlib/contrib/data/handler.py`
`TACHandler(DataHandlerLP)` already wraps the lake via `QlibDataLoader` +
`LakeFeatureProvider`. Key config surface (all usable from YAML without new code):
- `feature_fields` — explicit list; the handler prefixes `$` and de-dups. Anything
the provider can route is usable: bar fields, `$amount` (v*vw), and any column
present in the lake `features/.../symbol=*.parquet` files.
- `infer_processors` / `learn_processors` — add `CSRankNorm`, `CSZScoreNorm`
(label), `ZScoreNorm`, `DropnaLabel`, `Fillna`, etc. `DropAllNaN` is a
tac-qlib processor (drops all-NaN columns on the fit window).
- `label` — any qlib expression, e.g. `Ref($close,-6)/Ref($close,-1)-1`.
To add a *new feature family*: compute it once (see `examples/sp_features.py` +
`examples/persist_sp_features.py`), persist extra columns into
`features/market=US/timeframe=1d/symbol=*.parquet` (drop stale `sp_*` columns
first on re-runs), then reference them in `feature_fields`.
**The Rust engine already ships the SP feature pipeline as a lake MCP tool**:
`get_lake_sp` (tac-engine, stochastic-rs) computes `sp_ou_*`, `sp_hmm_*`,
`sp_jump_*`, `sp_rv*`/`sp_vol_ratio_*` (+ `sp_rv_ac1`, `sp_rv_cv_22`),
`sp_max_up`/`sp_max_down`, `sp_trend_slope_*`, `sp_logp`,
`sp_hurst_exponent`, `sp_sig_*` (levels 1/2 at lag 1 and 5),
`sp_rskew_*`/`sp_rkurt_*`/`sp_dsv_*` (realized moments via stochastic-rs
`realized`) + `sp_ret` from lake bars and persists them into
the feature parquets (replacing stale `sp_*`), all in one call:
```json
{"symbol": "AAPL", "timeframe": "1d", "start": "2015-01-03", "end": "2026-08-10", "fit_end": "2025-09-01"}
```
`fit_end` pins the Gaussian-HMM fit to the train window (no lookahead), matching
the `FIT_END` convention. **Deferred families** (`garch`, `entropy`, `catch22`)
are still computed with the Python `sp_features.py` path until their ports land.
Note two deliberate differences vs the Python reference: the Rust HMM uses the
causal *forward filter* (`filtered_state_probs`) rather than hmmlearn's smoothed
`predict_proba`, and `hurst` is estimated on the returns series directly
(`take_differences=false`) rather than the reference's double-differenced
`kind="random_walk"` — regime *state* assignments agree, probability levels are
comparable but not identical.
### 4. Custom Record (artifact writers)
Subclass `qlib.workflow.record_temp.SignalRecord` / a `Record` and log metrics +
artifacts into the MLflow run. There is no shipped example Record in `contrib/`
yet — write one against the pattern in `qlib.workflow.record_temp` when a
workflow needs a bespoke simulator (e.g. beta-neutral 3L/3S) that
`PortAnaRecord` doesn't cover.
## Empirical knobs that moved the numbers (measured on the 50-ETF lake)
All experiments used: 50-ETF universe, train 2015-01-03..2025-09-01 / valid
2025-09-03..2026-01-03 / test 2026-01-04..2026-08-10, benchmark SPY, TopkDropout
or OptimalStopControl, costs open 0.0005 / close 0.0015 / min 5.
> **Rank-dimension reminder**: when the goal is to improve the *ranking* quality of
> a signal (RankIC, long-short spread, top-decile precision), do NOT reinvent the
> stack — use the contrib modules already shipped and verified in this repo:
> `tac_qlib.contrib.model.rank_gbdt.RankICLGBModel` (early-stops training on
> per-day cross-sectional RankIC, `metric='None'` + `first_metric_only`) and
> `tac_qlib.contrib.strategy.optimal_stop.OptimalStopControl` (entry/exit gated by
> signal percentile instead of raw levels). Both are loadable from a workflow YAML
> via `module_path` — see the canonical `tac-qlib/workflows/workflow_lgb_sp5d_rankic.yaml`
> (rank dimension: model) and `workflow_lgb_sp5d_optstop.yaml` (rank dimension:
> portfolio construction). Verified end-to-end on 2026-01-04..2026-08-10:
> RankIC 0.071 / net-of-cost excess +20.7% ann (IR 0.70) vs SPY. Only write a new
> custom Model/Strategy when these proven paths are insufficient.
### Label
- **5-day forward return `Ref($close,-6)/Ref($close,-1)-1` ≫ 2-day.** IC nearly
tripled (0.0207 → 0.0645 standalone; the biggest single lever found). The 2-day
target is too noisy.
### Features
- **Stochastic-process features beat hand-rolled TA.** 55-feature set: OU
(`sp_ou_*`), 2-state HMM (`sp_hmm_*`), jump intensity (`sp_jump_*`, incl.
`sp_max_up`/`sp_max_down`), HARRV vol (`sp_rv*` + `sp_rv_ac1`/`sp_rv_cv_22`),
trend (`sp_trend_slope_*`, `sp_logp`), GARCH (`sp_garch_*`), Hurst
(`sp_hurst_exponent`), path signatures (`sp_sig_*`, lag 1 & 5), entropy
(`sp_ent_*`), realized moments (`sp_rskew_*`/`sp_rkurt_*`/`sp_dsv_*`),
catch22 (`sp_c22_*`). IC 0.036 → 0.047 vs the 19-feature v1.
- **Do NOT add ta-lib indicators on top** (SP+TA, 74 feats): IC dropped
0.047 → 0.031, RankIC 0.047 → 0.020. They're redundant with rv22/hmm/garch/catch22
and dilute CSRankNorm + LGBM.
- **CSRankNorm** (per-day cross-sectional rank) is important for the rank signal.
- Warm-up rows persist as all-NaN feature rows — expected; DropAllNaN/DropnaLabel
handle them.
### Model / training loop
- **LambdaRank / rank_xendcg objectives FAIL here** (RankIC → ~0): with only ~50
"documents" per query the rank gradient is noise.
- **Early-stopping metric beats objective.** MSE objective + early-stop on a
**RankIC feval** (mean per-day Spearman) lifted RankIC 0.047 → 0.075 (standalone).
- **The workflow gap was qlib's training loop**: `lgb.train` default
`first_metric_only=False` + `metric=l2` keeps training while l2 improves after
RankIC peaks. `RankICLGBModel` sets `metric='None'` + `first_metric_only=True`
so early-stopping tracks RankIC only.
- **RankIC-only early stop + bigger/smaller budget is the win**: `num_boost_round
3000`, `learning_rate 0.02`, `early_stopping_rounds 200`, `min_data_in_leaf 20`,
`lambda_l2 0.5` → test excess **+9.1% ann w/o cost (IR 1.03, maxDD −3.8%)** and
**+0.85% ann after costs** — the only config that beat SPY net. Note IC/RankIC
themselves were slightly lower (0.042) than the 500-tree run (0.051); the tuned
budget selects the iteration maximizing *valid* RankIC, converting to realized
excess return.
### Strategy / portfolio construction
- **Long-only construction leaves the edge on the table.** The SP-5d signal has
long-short **+31.6% ann (Sharpe 2.51)**, but TopkDropout long-only ≈ flat vs SPY,
and OptimalStopControl underperformed (valid-window threshold overfit: valid
+7.5% → test −17.7% on one calibration).
- **Costs eat most of the gross edge** (+9.1% → +0.85% net). Reduce turnover or go
long-short to widen the net edge.
- OptimalStopControl thresholds must be calibrated on the *valid* window and are
sensitive to overfit — prefer robust defaults or penalize turnover in selection.
## Gotchas
- **Installed package copy**: `tac_qlib` in the venv is a copy under
`/opt/venv/lib/python3.12/site-packages/tac_qlib/`. After editing any
`tac_qlib/contrib/**` module, `cp` it there or the workflow imports the stale
version. New subpackages need `mkdir -p` first.
- `qlib.backtest` exports `Order` but not `OrderDir`/`Position` at top level —
import `Order` from `qlib.backtest`, `OrderDir`/`TradeDecisionWO` from
`qlib.backtest.decision`, `Position` from `qlib.backtest.position`.
- `qlib.backtest.high_performance_ds` may not export `Order` in this build — don't
import from it.
- HMM / GARCH / catch22 features must not see test data at fit time: fit the HMM
on the train window only (`fit_end=FIT_END`), and compute rolling windows ending
at each day. GARCH/entropy use a stride + forward-fill for speed (~5x).
- `pycatch22`, `arch`, `hurst`, `antropy`, `hmmlearn` are required for the full
feature set; install with `uv pip install --python /app/.venv/bin/python <pkg>`
(a C compiler is needed for `pycatch22`). `duckdb` and `pyarrow` are declared in
`tac-qlib/pyproject.toml`; if a workflow import fails on either, lazy-install with
`uv pip install --python /app/.venv/bin/python duckdb pyarrow`.
- `rd_run_workflow` defaults to `wait=false`: it returns immediately with
`status: started` and the workflow runs in a background thread — poll
`rd_exp_get_run` / `rd_exp_list` for the newest run of the experiment
(status `RUNNING` until it finishes), then reuse its `run_id`. Pass
`wait=true` only for small windows that finish within the MCP call timeout.
- After fixing a YAML model/handler change, remember both `/app/tac-qlib/...` and
the `/opt/venv` copy stay in sync.
## Files this skill is based on
Minimal, runnable examples live next to this skill in `examples/` — they are the
canonical reference for every artifact the skill describes:
- Workflows (full `record` block → MLflow on disk):
- `examples/workflow_minimal.yaml` — the canonical backtest template (req: every
traced backtest runs through a workflow like this via `rd_run_workflow`)
- `examples/workflow_rankic.yaml` — RankIC-early-stop model wired in
- Repo workflows for reference: `tac-qlib/workflows/workflow_lgb_taclake.yaml`,
`tune_run1_wider_5d.yaml`, `tune_run2_regularized.yaml`, `tune_run3_label5d_clean_universe.yaml`,
`tune_run4_fix_universe_longtrain.yaml`, `tune_run5_longtest.yaml`
- Models: `examples/model_rank_gbdt.py` (`RankICLGBModel`: per-day groups +
`feval=rankic` + `metric='None'`). Repo: `tac_qlib/contrib/model/rank_gbdt.py`
- Strategies: `examples/strategy_optimal_stop.py` (`OptimalStopControl`),
`examples/strategy_beta_neutral.py` (doc-only 3L/3S stub — pattern for a
custom strategy + Record; not wired into the package)
- Handler: `examples/handler.py` (how to subclass `TACHandler`); repo:
`tac_qlib/contrib/data/handler.py`; providers: `tac_qlib/data/providers.py`
- Feature engineering: `examples/sp_features.py` (OU + Hurst) and
`examples/persist_sp_features.py` (persist `sp_*` into the lake features parquet)
- Ranking experiments: `examples/run_rank_objectives.py` (mse vs lambdarank vs
rank_xendcg ablation on the lake)
- Optstop calibration: `examples/run_optstop_compare.py` (valid-window grid +
overfit warning)
- Traceability tooling: the `rd_trace_*` MCP tools (tac-qlib-rd,
`tac_qlib/trace.py`) — see the traceability section above