30 KiB
name, description
| name | description |
|---|---|
| tac-qlib-custom | Guide agents to customize and extend Qlib on the TradeAC R&D stack — how to configure workflow YAMLs (qlib_init, model, dataset/handler, processors, records, PortAnaRecord strategies), how to extend Qlib classes wired into those workflows (custom Model, BaseStrategy, DataHandler, Record), and the empirically-tested knobs from this repo (RankIC early-stopping, stochastic-control strategies, stochastic-process features, catch22/GARCH/Hurst/signature). Also encodes the experiment traceability loop: every backtest runs as a workflow-with-recorder, is recorded in the Postgres experiments table (rationale/details/evaluation/metrics with pgvector embeddings, evolution chain) and on a per-experiment git branch that is committed + pushed. Companion to tradeac-rd (MCP run tools) and tradeac-lake (parquet lake). |
tac-qlib-custom
Customizing and extending Qlib on the TradeAC stack. This skill encodes what was learned from actual experiments in this repo: how a workflow YAML maps to Qlib classes, how to write a custom class that the YAML can load, and which training / strategy / feature knobs measurably moved IC, RankIC and the backtest.
Read tac-qlib/skills/tradeac-rd/SKILL.md for the MCP run/inspect tools and
tac-qlib/README.md for the package layout. The venv is /app/.venv
(qlib 0.1.dev2066); tac_qlib is installed into the venv's site-packages
(editable copy under /opt/venv/.../tac_qlib/), so any new module must be
copied to /opt/venv/lib/python3.12/site-packages/tac_qlib/... too (or use an
editable install) before rd_run_workflow can import it.
MCP-first policy
- Drive every backtest and run through the
tac-qlib-rdMCP tools (rd_run_workflow,rd_train,rd_predict,rd_exp_*) and the tac-engine lake tools for data prep. Do not reimplement them with ad-hoc scripts (custom qlib glue, own mlruns readers, direct JSON-RPC/stdio clients). - NEVER script directly against the MCP server (spawning
tac_qlib.rd_server/tac-engine, bash/curl/stdio) unless a tool genuinely can't do the job — then stop and ask the user to confirm first. - The traceability bookkeeping (Postgres
rd_experimentsrow + pgvector embeddings + branch-per-experiment git) is exposed as therd_trace_*MCP tools on the tac-qlib-rd server — use those, not bash scripts. Data prep, training, evaluation and backtests also go through MCP tools. - If the venv is missing a runtime dep (
duckdb,pyarrow, feature libs), lazy-install it (uv pip install --python $VIRTUAL_ENV/bin/python <pkg>) instead of switching tools.
Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or
credential-bearing URLs (
DATABASE_URL,GIT_PASS,EMBEDDING_API_KEY) in workflow YAMLs, scripts, configs, notes or committed code. - NEVER read
*.env/.env.*directly (cat/tail/grep/sed/headon.env). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the
environment (shell/container env, or the user-owned
.env) and reference it by name ($VAR), never by value. If it's missing, report which variable is required instead of reading it yourself. - Tracking store: use
uri: "sqlite:///mlruns.db"(relative) in workflows —rd_run_workflownormalizes it to Postgres when$DATABASE_URLis set, else the lake sqlite. Never hardcode apostgres://user:pass@…URI. - If you find a committed secret, flag it, remove it, and replace it with a
placeholder. (The
rd_trace_*MCP tools' commit guard blocks adding credential-shaped lines.)
How a workflow YAML maps to Qlib classes
A workflow YAML (tac-qlib/workflows/*.yaml) is rendered by Jinja (vars like
{{ LAKE }} from TAC_LAKE_DIR) then executed by qrun / rd_run_workflow.
Every block is a Qlib class reference resolved by module_path + class:
{%- set LAKE = TAC_LAKE_DIR %}
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
calendar_provider: # custom tac-qlib providers read the parquet lake
class: LakeCalendarProvider
module_path: tac_qlib.data.providers
instrument_provider: # ... (markets: {} => lake universe)
feature_provider: # LakeFeatureProvider: routes $open..$volume from bars,
class: LakeFeatureProvider # $<ta-lib/sp_*> from features parquet, $amount derived
exp_manager:
class: MLflowExpManager
module_path: qlib.workflow.expm
kwargs: { uri: "sqlite:///{{ LAKE }}/mlruns.db", default_exp_name: "my-exp" }
task:
model: # <MODEL BLOCK> — custom model → new module_path
class: RankICLGBModel
module_path: tac_qlib.contrib.model.rank_gbdt
kwargs: { loss: mse, learning_rate: 0.02, num_leaves: 31, ... }
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler: # <HANDLER BLOCK> — feature selection + processors live here
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: "SPY,QQQ,..."
start_time: 2015-01-03
end_time: 2026-08-10
fit_start_time: 2015-01-03 # processors fit on this window
fit_end_time: 2025-09-01
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-6)/Ref($close,-1)-1" # 5d forward return
feature_fields: "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_ou_zscore,..."
infer_processors: # feature-time transforms, fit on fit_*
- { class: DropAllNaN, kwargs: {} }
- { class: ProcessInf, kwargs: {} }
- { class: CSRankNorm, kwargs: {} } # per-day cross-sectional rank
- { class: ZScoreNorm, kwargs: {} }
- { class: Fillna, kwargs: {} }
segments:
train: [2015-01-03, 2025-09-01]
valid: [2025-09-03, 2026-01-03]
test: [2026-01-04, 2026-08-10]
record: # each entry records one artifact type to the run
- { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} }
- { class: SigAnaRecord, module_path: qlib.workflow.record_temp,
kwargs: { ana_long_short: true, ann_scaler: 252 } }
- { class: PortAnaRecord, module_path: qlib.workflow.record_temp,
kwargs: { config: { strategy: <STRATEGY BLOCK>, backtest: {...} }, risk_analysis_freq: 1d } }
rd_run_workflow config_path=<yaml> experiment_name=<exp> runs it; the MCP call
may time out for long runs (RankIC tuning, heavy feature sets) — the run keeps
executing; poll via rd_exp_list / rd_exp_get_run on the returned experiment.
Experiment traceability (DB + git + embeddings)
Every backtest you run as an agent MUST be tracked: it runs as a workflow with the
record block (SignalRecord/SigAnaRecord/PortAnaRecord → MLflow artifacts on disk
under <lake>/mlruns/<exp_id>/<run_id>), and a row is written to the Postgres
experiments table plus a git branch per experiment. The tac-app UI owns the
schema (Drizzle migrations in tac-app/drizzle/); this skill's lib/ scripts are
the executor the agent drives.
Trigger the lineage as part of the run — automatically, not on prompt. Any
time you execute a qlib workflow (rd_run_workflow) or a train/predict pipeline
on this stack, the traceability bookkeeping is part of that run, not a separate
step the user must ask for: open the traced experiment with rd_trace_start
before running, commit intermediates with rd_trace_commit, and close it with
rd_trace_finish after — without waiting to be prompted (see "The
per-experiment procedure" below).
Env vars
| Var | Purpose |
|---|---|
DATABASE_URL |
Postgres URL for the rd_experiments table AND the MLflow tracking store (set in repo .env) |
EMBEDDING_API_BASE_URL |
embedding POST endpoint (e.g. https://embd.h.lizhao.net/embeddings) |
EMBEDDING_API_KEY |
basic-auth credential (user:pass form is supported) |
GIT_USER / GIT_PASS |
git remote credentials for push/fetch |
GIT_REPO_URL |
experiment git repo tracked by the experiments submodule (branches are pushed here) |
TAC_LAKE_DIR |
lake root (mlruns artifact files live under it) |
The experiment repo is the experiments git submodule at the workspace root
(<repo-root>/experiments), always tracking $GIT_REPO_URL. rd_trace_init
creates/validates it; it errors if experiments/ exists but points at a
different URL. There is no TAC_EXP_GIT_DIR — the submodule path IS the
experiment repo, and ALL experiment/backtest changes (workflow YAMLs, notes,
outputs) must live inside it, never in the parent tradeac repo.
The rd_experiments table
Owned by tac-app's Drizzle schema (tac-app/src/db/schema.ts); rd_trace_init
can init it idempotently. The table is named rd_experiments (NOT
experiments) because MLflow's Postgres tracking store creates its own
experiments table in the same database. Key columns: id (PK), rational +
rational_embedding (pgvector vector(384)), details + details_embedding,
evaluation, metrics (jsonb), evolved_from (FK → rd_experiments.id),
start_ts/end_ts, git_branch, experiment_ref_id, mlruns_dir, status.
experiment_ref_id holds the mlflow run id returned by rd_run_workflow and
is an FK to MLflow's runs(run_uuid) (added by rd_trace_init after the
mlflow store tables exist — MLflow creates runs lazily).
Tracking store: Postgres $DATABASE_URL (MLflow's own tables) when set,
falling back to the unified lake sqlite sqlite:///<lake>/mlruns.db. Artifact
files always stay on disk under <lake>/mlruns/<exp_id>/<run_id>/artifacts.
Embedding model: michaelfeil/bge-small-en-v1.5 (384-dim, 512-token context).
Rational/details are written paper-summary style (≤512 tokens) and embedded verbatim —
NEVER truncate; if a text is longer, summarize it first (the embed helper rejects
over-limit input).
Git repo + branch-per-experiment
The experiment repo is the experiments submodule at the workspace root
(<repo-root>/experiments, tracking $GIT_REPO_URL). The rd_trace_* MCP
tools handle it, and every git operation is scoped to that submodule —
experiments NEVER stage or push parent-repo (tradeac) files.
rd_trace_initcreates/validates the submodule and the base branch. Ifexperiments/does not exist it runsgit clone $GIT_REPO_URL experiments; if it exists but tracks a different URL, init errors out.- Base branch:
main(ormaster). If the submodule is empty, a seed commit is made and pushed so there are commits to fork from. - Every experiment runs on its own branch
exp/<id>-<slug>. evolved_fromresolution (in order):- If the wizard prompt explicitly says
evolved_from=<id>(run wizard click on an existing experiment) — use that id directly. - Otherwise
--evolved-from auto: the user prompt / rational is embedded and cosine-searched over theexperiments.rational_embeddingcolumn; the top hit above the similarity threshold (0.5) becomesevolved_from. - Otherwise (first experiment, or a new chat with no predecessor) — no evolved_from;
fork from
main's latest commits.
- If the wizard prompt explicitly says
- The new branch is forked from the evolved-from experiment's branch (its latest
commits), or from
mainwhen there is no predecessor — so experiment lineages form a git branch chain. - On every finish, and for intermediate steps, changes are committed + pushed.
Custom code is part of the lineage (code snapshot)
Custom contrib modules (tac_qlib/contrib/model/, tac_qlib/contrib/strategy/,
tac_qlib/contrib/data/, tac_qlib/data/providers.py) live in the parent
tradeac repo, not in the experiments/ submodule — so they are normally invisible
to the experiment branch and a descendant forking from it would reinvent them.
The lineage tooling fixes this: every experiment branch carries a code/
snapshot of exactly the qlib extension code that run depended on, so descendants
reuse it instead of re-authoring it.
rd_trace_startandrd_trace_finishautomatically snapshot the default paths (tac-qlib/tac_qlib/contrib,tac-qlib/tac_qlib/data) into<experiments>/code/<parent-relative-path>on the experiment branch.rd_trace_snapshotsnapshots mid-run (e.g. after writing a new custom model) without waiting for finish.- The snapshot also writes
code/MANIFEST.txtrecording the parent-repo HEAD commit and the per-file blob hashes it was taken from — so a run can be traced back to the exact parent commit that produced its custom code. - Descendants: the custom modules your run needs are under
code/tac_qlib/...on the evolved-from branch. Reuse them (copy/git show) instead of writing new ones; checkcode/MANIFEST.txtto see which parent commit they came from and port fixes back. - Guardrail exception: parent-repo changes under
tac_qlib/tac_qlib/contribandtac_qlib/tac_qlib/dataare expected (they are the snapshotted code);parent_changesreports them as a note, not a violation. Any OTHER parent change is still a guardrail violation.
Guardrail — experiments must NOT introduce side effects to the parent repo:
- Write workflow YAMLs, notes and experiment outputs ONLY inside
<repo-root>/experiments/(they are committed on the experiment branch). - Never
git add/commit/stage anything in the parent tradeac repo. - Run
rd_trace_guardto list any parent changes outside the submodule pointer;rd_trace_finishalso surfaces them. Revert any accidental parent edits before finishing. - If an experiment reveals a PRODUCT change (workflow template, skill, tac-app), propose it separately for the tradeac repo — do not mix it into the experiment branch.
The rd_trace_* MCP tools perform git operations with the mandated credential
helper (from GIT_USER / GIT_PASS), so you do not need to construct it by hand.
The per-experiment procedure
Use the rd_trace_* MCP tools (tac-qlib-rd) — they replace the old
trace.sh/trace_db.py scripts. The server is long-lived (psycopg imported
once, DB connection reused per call) and every tool returns one JSON object, so
no output parsing is needed:
# 0. ensure ready (rd_experiments table + experiments git repo + base main)
rd_trace_init
# 1. start — inserts the row, resolves evolved_from, forks+pushes the branch.
# Returns {experiment_id, branch, evolved_from, base_branch} as JSON.
rd_trace_start rational="5-day forward label, RankIC early stop, 50-ETF universe" \
details="LGBModel mse lr=0.02 num_leaves=15 num_boost_round=3000; TopkDropout topk=2; benchmark QQQ" \
experiment_name="tac-rd-expN" \
evolved_from="auto" \
session_id="<this chat's opencode session id, if started from a chat>"
# -> {"experiment_id": N, "branch": "exp/N-...", "evolved_from": ..., "base_branch": ...}
# 2. write the workflow YAML INSIDE the experiments submodule
# (e.g. <repo-root>/experiments/workflows/<exp>/workflow.yaml), then commit it:
rd_trace_commit experiment_id=<N> message="add workflow yaml"
# 2b. if the workflow uses a NEW custom module, snapshot it onto the branch
# (start/finish auto-snapshot contrib+data; do this to capture mid-run):
rd_trace_snapshot experiment_id=<N> # default contrib+data
# or: rd_trace_snapshot experiment_id=<N> paths="tac-qlib/tac_qlib/contrib/model/rank_gbdt.py"
# 3. run the backtest through the WORKFLOW with the recorder (MUST write mlruns):
rd_run_workflow config_path=<repo-root>/experiments/workflows/<exp>/workflow.yaml experiment_name=tac-rd-expN
# -> returns run_id (= experiment_ref_id) + metrics
# 4. inspect with rd_exp_result / rd_exp_blotter, then finish — updates the row
# (re-embeds rational/details, sets metrics/eval/end_ts), snapshots the custom
# code, and commits+pushes. finish also surfaces parent-repo side effects.
rd_trace_finish experiment_id=<N> \
ref_id=<mlflow-run-id> \
evaluation="IC 0.0645, RankIC 0.075; net excess +0.85% ann" \
metrics='{"IC":0.0645,"RankIC":0.075,"ann_excess":0.85}' \
mlruns_dir=<lake>/mlruns/<exp_id>/<run_id>
Helpers (MCP tools): rd_trace_search (semantic), rd_trace_get (one row),
rd_trace_list, rd_trace_mlruns_dir (resolves the mlruns dir for an
experiment name), rd_trace_guard (parent-repo side-effect check).
Rules:
- Always run backtests as workflows with the
recordblock (req 2) — never a barerd_backtestfor a traced experiment. - Always open the lineage (
rd_trace_start) BEFORE the run and Alwaysrd_trace_finish+ push after it completes (req 5) — this happens as part of the run, do not wait for the user to ask; intermediaterd_trace_commitis encouraged (req 5). - Always snapshot the custom qlib code (
rd_trace_snapshot, or rely on the auto-snapshot at start/finish) so the experiment branch carries the exact contrib/data modules the run used — descendants fork and reusecode/instead of reinventing it. - Keep rational/details ≤ 512 tokens (paper-summary style) so embeddings are exact — no truncation.
- Confine experiments to the
experiments/submodule — never write to, stage, or commit parent tradeac repo files; runrd_trace_guardto check for side effects. (Custom code edits undertac-qlib/tac_qlib/contriband.../dataare the sanctioned exception — they are the snapshotted modules; see "Custom code is part of the lineage".) - Follow the Secrets policy above — no secrets in files, no reading
.env*, ask the user to set env vars; useuri: "sqlite:///mlruns.db"for the tracking store. - Workflow YAMLs are jinja-rendered with
os.environas the context, so env-var placeholders work ({%- set LAKE = TAC_LAKE_DIR %}then{{ LAKE }}). Use them for paths/config — never for secrets that get committed.
Extending Qlib — the 4 class families you can override
1. Custom Model (train-time) — tac_qlib/contrib/model/
Subclass qlib.contrib.model.gbdt.LGBModel (or qlib.model.base.BaseModel) and
implement fit(dataset, ...) + predict(dataset). LGBModel.fit calls
self._prepare_data(dataset) → lgb.Datasets, then lgb.train with
early_stopping on the valid set. Override points that matter:
_prepare_data→ build thelgb.Datasetwithgroup=(per-day query groups) when you need ranking metrics per trading day.fit→ change what early-stops training (the biggest IC/backtest lever, see §Knobs).predict→ return the Series keyed (datetime, instrument).
Reference: tac_qlib/tac_qlib/contrib/model/rank_gbdt.py — RankICLGBModel
subclasses LGBModel, adds per-day group in _prepare_data, injects
feval=rankic_feval (mean per-day Spearman) into lgb.train, and forces
metric='None' + first_metric_only=True so early-stopping tracks RankIC only.
2. Custom Strategy (backtest-time) — tac_qlib/contrib/strategy/
Subclass qlib.contrib.strategy.signal_strategy.BaseSignalStrategy (which wraps
qlib.strategy.base.BaseStrategy) and implement:
def generate_trade_decision(self, execute_result=None):
# trade_step, trade_start/end = self.trade_calendar.get_step_time(trade_step)
# pred = self.signal.get_signal(start_time=pred_shift, end_time=pred_shift) # shift=-1 => signal known at t-1
# self.trade_position / self.trade_exchange / self.trade_calendar injected by the executor
# build qlib.backtest.Order(stock_id, amount, start_time, end_time, direction=Order.BUY/SELL)
# return TradeDecisionWO(orders, self)
Wire it into the YAML under PortAnaRecord.config.strategy:
strategy:
class: OptimalStopControl
module_path: tac_qlib.contrib.strategy.optimal_stop
kwargs:
signal: "<PRED>" # placeholder replaced with the recorded pred
topk: 10
entry_pct: 0.85
exit_pct: 0.7
max_hold_days: 10
min_hold_days: 2
sl: -0.08
risk_degree: 0.95
Reference: tac_qlib/tac_qlib/contrib/strategy/optimal_stop.py
(OptimalStopControl — entry gated by cross-sectional signal percentile, exits
by percentile/time/stop-loss, equal-weight control sizing).
3. Custom DataHandler / processors — tac_qlib/contrib/data/handler.py
TACHandler(DataHandlerLP) already wraps the lake via QlibDataLoader +
LakeFeatureProvider. Key config surface (all usable from YAML without new code):
feature_fields— explicit list; the handler prefixes$and de-dups. Anything the provider can route is usable: bar fields,$amount(v*vw), and any column present in the lakefeatures/.../symbol=*.parquetfiles.infer_processors/learn_processors— addCSRankNorm,CSZScoreNorm(label),ZScoreNorm,DropnaLabel,Fillna, etc.DropAllNaNis a tac-qlib processor (drops all-NaN columns on the fit window).label— any qlib expression, e.g.Ref($close,-6)/Ref($close,-1)-1.
To add a new feature family: compute it once (see examples/sp_features.py +
examples/persist_sp_features.py), persist extra columns into
features/market=US/timeframe=1d/symbol=*.parquet (drop stale sp_* columns
first on re-runs), then reference them in feature_fields.
The Rust engine already ships the SP feature pipeline as a lake MCP tool:
get_lake_sp (tac-engine, stochastic-rs) computes sp_ou_*, sp_hmm_*,
sp_jump_*, sp_rv*/sp_vol_ratio_* (+ sp_rv_ac1, sp_rv_cv_22),
sp_max_up/sp_max_down, sp_trend_slope_*, sp_logp,
sp_hurst_exponent, sp_sig_* (levels 1/2 at lag 1 and 5),
sp_rskew_*/sp_rkurt_*/sp_dsv_* (realized moments via stochastic-rs
realized) + sp_ret from lake bars and persists them into
the feature parquets (replacing stale sp_*), all in one call:
{"symbol": "AAPL", "timeframe": "1d", "start": "2015-01-03", "end": "2026-08-10", "fit_end": "2025-09-01"}
fit_end pins the Gaussian-HMM fit to the train window (no lookahead), matching
the FIT_END convention. Deferred families (garch, entropy, catch22)
are still computed with the Python sp_features.py path until their ports land.
Note two deliberate differences vs the Python reference: the Rust HMM uses the
causal forward filter (filtered_state_probs) rather than hmmlearn's smoothed
predict_proba, and hurst is estimated on the returns series directly
(take_differences=false) rather than the reference's double-differenced
kind="random_walk" — regime state assignments agree, probability levels are
comparable but not identical.
4. Custom Record (artifact writers)
Subclass qlib.workflow.record_temp.SignalRecord / a Record and log metrics +
artifacts into the MLflow run. There is no shipped example Record in contrib/
yet — write one against the pattern in qlib.workflow.record_temp when a
workflow needs a bespoke simulator (e.g. beta-neutral 3L/3S) that
PortAnaRecord doesn't cover.
Empirical knobs that moved the numbers (measured on the 50-ETF lake)
All experiments used: 50-ETF universe, train 2015-01-03..2025-09-01 / valid 2025-09-03..2026-01-03 / test 2026-01-04..2026-08-10, benchmark SPY, TopkDropout or OptimalStopControl, costs open 0.0005 / close 0.0015 / min 5.
Rank-dimension reminder: when the goal is to improve the ranking quality of a signal (RankIC, long-short spread, top-decile precision), do NOT reinvent the stack — use the contrib modules already shipped and verified in this repo:
tac_qlib.contrib.model.rank_gbdt.RankICLGBModel(early-stops training on per-day cross-sectional RankIC,metric='None'+first_metric_only) andtac_qlib.contrib.strategy.optimal_stop.OptimalStopControl(entry/exit gated by signal percentile instead of raw levels). Both are loadable from a workflow YAML viamodule_path— see the canonicaltac-qlib/workflows/workflow_lgb_sp5d_rankic.yaml(rank dimension: model) andworkflow_lgb_sp5d_optstop.yaml(rank dimension: portfolio construction). Verified end-to-end on 2026-01-04..2026-08-10: RankIC 0.071 / net-of-cost excess +20.7% ann (IR 0.70) vs SPY. Only write a new custom Model/Strategy when these proven paths are insufficient.
Label
- 5-day forward return
Ref($close,-6)/Ref($close,-1)-1≫ 2-day. IC nearly tripled (0.0207 → 0.0645 standalone; the biggest single lever found). The 2-day target is too noisy.
Features
- Stochastic-process features beat hand-rolled TA. 55-feature set: OU
(
sp_ou_*), 2-state HMM (sp_hmm_*), jump intensity (sp_jump_*, incl.sp_max_up/sp_max_down), HARRV vol (sp_rv*+sp_rv_ac1/sp_rv_cv_22), trend (sp_trend_slope_*,sp_logp), GARCH (sp_garch_*), Hurst (sp_hurst_exponent), path signatures (sp_sig_*, lag 1 & 5), entropy (sp_ent_*), realized moments (sp_rskew_*/sp_rkurt_*/sp_dsv_*), catch22 (sp_c22_*). IC 0.036 → 0.047 vs the 19-feature v1. - Do NOT add ta-lib indicators on top (SP+TA, 74 feats): IC dropped 0.047 → 0.031, RankIC 0.047 → 0.020. They're redundant with rv22/hmm/garch/catch22 and dilute CSRankNorm + LGBM.
- CSRankNorm (per-day cross-sectional rank) is important for the rank signal.
- Warm-up rows persist as all-NaN feature rows — expected; DropAllNaN/DropnaLabel handle them.
Model / training loop
- LambdaRank / rank_xendcg objectives FAIL here (RankIC → ~0): with only ~50 "documents" per query the rank gradient is noise.
- Early-stopping metric beats objective. MSE objective + early-stop on a RankIC feval (mean per-day Spearman) lifted RankIC 0.047 → 0.075 (standalone).
- The workflow gap was qlib's training loop:
lgb.traindefaultfirst_metric_only=False+metric=l2keeps training while l2 improves after RankIC peaks.RankICLGBModelsetsmetric='None'+first_metric_only=Trueso early-stopping tracks RankIC only. - RankIC-only early stop + bigger/smaller budget is the win:
num_boost_round 3000,learning_rate 0.02,early_stopping_rounds 200,min_data_in_leaf 20,lambda_l2 0.5→ test excess +9.1% ann w/o cost (IR 1.03, maxDD −3.8%) and +0.85% ann after costs — the only config that beat SPY net. Note IC/RankIC themselves were slightly lower (0.042) than the 500-tree run (0.051); the tuned budget selects the iteration maximizing valid RankIC, converting to realized excess return.
Strategy / portfolio construction
- Long-only construction leaves the edge on the table. The SP-5d signal has long-short +31.6% ann (Sharpe 2.51), but TopkDropout long-only ≈ flat vs SPY, and OptimalStopControl underperformed (valid-window threshold overfit: valid +7.5% → test −17.7% on one calibration).
- Costs eat most of the gross edge (+9.1% → +0.85% net). Reduce turnover or go long-short to widen the net edge.
- OptimalStopControl thresholds must be calibrated on the valid window and are sensitive to overfit — prefer robust defaults or penalize turnover in selection.
Gotchas
- Installed package copy:
tac_qlibin the venv is a copy under/opt/venv/lib/python3.12/site-packages/tac_qlib/. After editing anytac_qlib/contrib/**module,cpit there or the workflow imports the stale version. New subpackages needmkdir -pfirst. qlib.backtestexportsOrderbut notOrderDir/Positionat top level — importOrderfromqlib.backtest,OrderDir/TradeDecisionWOfromqlib.backtest.decision,Positionfromqlib.backtest.position.qlib.backtest.high_performance_dsmay not exportOrderin this build — don't import from it.- HMM / GARCH / catch22 features must not see test data at fit time: fit the HMM
on the train window only (
fit_end=FIT_END), and compute rolling windows ending at each day. GARCH/entropy use a stride + forward-fill for speed (~5x). pycatch22,arch,hurst,antropy,hmmlearnare required for the full feature set; install withuv pip install --python /app/.venv/bin/python <pkg>(a C compiler is needed forpycatch22).duckdbandpyarroware declared intac-qlib/pyproject.toml; if a workflow import fails on either, lazy-install withuv pip install --python /app/.venv/bin/python duckdb pyarrow.rd_run_workflowdefaults towait=false: it returns immediately withstatus: startedand the workflow runs in a background thread — pollrd_exp_get_run/rd_exp_listfor the newest run of the experiment (statusRUNNINGuntil it finishes), then reuse itsrun_id. Passwait=trueonly for small windows that finish within the MCP call timeout.- After fixing a YAML model/handler change, remember both
/app/tac-qlib/...and the/opt/venvcopy stay in sync.
Files this skill is based on
Minimal, runnable examples live next to this skill in examples/ — they are the
canonical reference for every artifact the skill describes:
- Workflows (full
recordblock → MLflow on disk):examples/workflow_minimal.yaml— the canonical backtest template (req: every traced backtest runs through a workflow like this viard_run_workflow)examples/workflow_rankic.yaml— RankIC-early-stop model wired in- Repo workflows for reference:
tac-qlib/workflows/workflow_lgb_taclake.yaml,tune_run1_wider_5d.yaml,tune_run2_regularized.yaml,tune_run3_label5d_clean_universe.yaml,tune_run4_fix_universe_longtrain.yaml,tune_run5_longtest.yaml
- Models:
examples/model_rank_gbdt.py(RankICLGBModel: per-day groups +feval=rankic+metric='None'). Repo:tac_qlib/contrib/model/rank_gbdt.py - Strategies:
examples/strategy_optimal_stop.py(OptimalStopControl),examples/strategy_beta_neutral.py(doc-only 3L/3S stub — pattern for a custom strategy + Record; not wired into the package) - Handler:
examples/handler.py(how to subclassTACHandler); repo:tac_qlib/contrib/data/handler.py; providers:tac_qlib/data/providers.py - Feature engineering:
examples/sp_features.py(OU + Hurst) andexamples/persist_sp_features.py(persistsp_*into the lake features parquet) - Ranking experiments:
examples/run_rank_objectives.py(mse vs lambdarank vs rank_xendcg ablation on the lake) - Optstop calibration:
examples/run_optstop_compare.py(valid-window grid + overfit warning) - Traceability tooling: the
rd_trace_*MCP tools (tac-qlib-rd,tac_qlib/trace.py) — see the traceability section above