Files
tac-exp-dev/tac-qlib/skills/tac-qlib-custom/SKILL.md
T

30 KiB
Raw Blame History

name, description
name description
tac-qlib-custom Guide agents to customize and extend Qlib on the TradeAC R&D stack — how to configure workflow YAMLs (qlib_init, model, dataset/handler, processors, records, PortAnaRecord strategies), how to extend Qlib classes wired into those workflows (custom Model, BaseStrategy, DataHandler, Record), and the empirically-tested knobs from this repo (RankIC early-stopping, stochastic-control strategies, stochastic-process features, catch22/GARCH/Hurst/signature). Also encodes the experiment traceability loop: every backtest runs as a workflow-with-recorder, is recorded in the Postgres experiments table (rationale/details/evaluation/metrics with pgvector embeddings, evolution chain) and on a per-experiment git branch that is committed + pushed. Companion to tradeac-rd (MCP run tools) and tradeac-lake (parquet lake).

tac-qlib-custom

Customizing and extending Qlib on the TradeAC stack. This skill encodes what was learned from actual experiments in this repo: how a workflow YAML maps to Qlib classes, how to write a custom class that the YAML can load, and which training / strategy / feature knobs measurably moved IC, RankIC and the backtest.

Read tac-qlib/skills/tradeac-rd/SKILL.md for the MCP run/inspect tools and tac-qlib/README.md for the package layout. The venv is /app/.venv (qlib 0.1.dev2066); tac_qlib is installed into the venv's site-packages (editable copy under /opt/venv/.../tac_qlib/), so any new module must be copied to /opt/venv/lib/python3.12/site-packages/tac_qlib/... too (or use an editable install) before rd_run_workflow can import it.

MCP-first policy

  • Drive every backtest and run through the tac-qlib-rd MCP tools (rd_run_workflow, rd_train, rd_predict, rd_exp_*) and the tac-engine lake tools for data prep. Do not reimplement them with ad-hoc scripts (custom qlib glue, own mlruns readers, direct JSON-RPC/stdio clients).
  • NEVER script directly against the MCP server (spawning tac_qlib.rd_server / tac-engine, bash/curl/stdio) unless a tool genuinely can't do the job — then stop and ask the user to confirm first.
  • The traceability bookkeeping (Postgres rd_experiments row + pgvector embeddings + branch-per-experiment git) is exposed as the rd_trace_* MCP tools on the tac-qlib-rd server — use those, not bash scripts. Data prep, training, evaluation and backtests also go through MCP tools.
  • If the venv is missing a runtime dep (duckdb, pyarrow, feature libs), lazy-install it (uv pip install --python $VIRTUAL_ENV/bin/python <pkg>) instead of switching tools.

Secrets policy

  • NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (DATABASE_URL, GIT_PASS, EMBEDDING_API_KEY) in workflow YAMLs, scripts, configs, notes or committed code.
  • NEVER read *.env / .env.* directly (cat/tail/grep/sed/head on .env). That pulls secrets into this session and leaks them to any agent sharing it.
  • When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned .env) and reference it by name ($VAR), never by value. If it's missing, report which variable is required instead of reading it yourself.
  • Tracking store: use uri: "sqlite:///mlruns.db" (relative) in workflows — rd_run_workflow normalizes it to Postgres when $DATABASE_URL is set, else the lake sqlite. Never hardcode a postgres://user:pass@… URI.
  • If you find a committed secret, flag it, remove it, and replace it with a placeholder. (The rd_trace_* MCP tools' commit guard blocks adding credential-shaped lines.)

How a workflow YAML maps to Qlib classes

A workflow YAML (tac-qlib/workflows/*.yaml) is rendered by Jinja (vars like {{ LAKE }} from TAC_LAKE_DIR) then executed by qrun / rd_run_workflow. Every block is a Qlib class reference resolved by module_path + class:

{%- set LAKE = TAC_LAKE_DIR %}
qlib_init:
    provider_uri: "{{ LAKE }}"
    region: us
    calendar_provider:        # custom tac-qlib providers read the parquet lake
        class: LakeCalendarProvider
        module_path: tac_qlib.data.providers
    instrument_provider:      # ... (markets: {} => lake universe)
    feature_provider:         # LakeFeatureProvider: routes $open..$volume from bars,
        class: LakeFeatureProvider   # $<ta-lib/sp_*> from features parquet, $amount derived
    exp_manager:
        class: MLflowExpManager
        module_path: qlib.workflow.expm
        kwargs: { uri: "sqlite:///{{ LAKE }}/mlruns.db", default_exp_name: "my-exp" }

task:
    model:                    # <MODEL BLOCK> — custom model → new module_path
        class: RankICLGBModel
        module_path: tac_qlib.contrib.model.rank_gbdt
        kwargs: { loss: mse, learning_rate: 0.02, num_leaves: 31, ... }
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:          # <HANDLER BLOCK> — feature selection + processors live here
                class: TACHandler
                module_path: tac_qlib.contrib.data.handler
                kwargs:
                    instruments: "SPY,QQQ,..."
                    start_time: 2015-01-03
                    end_time: 2026-08-10
                    fit_start_time: 2015-01-03     # processors fit on this window
                    fit_end_time: 2025-09-01
                    freq: day
                    lake_root: "{{ LAKE }}"
                    market: US
                    label: "Ref($close,-6)/Ref($close,-1)-1"   # 5d forward return
                    feature_fields: "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_ou_zscore,..."
                    infer_processors:              # feature-time transforms, fit on fit_*
                        - { class: DropAllNaN, kwargs: {} }
                        - { class: ProcessInf, kwargs: {} }
                        - { class: CSRankNorm, kwargs: {} }   # per-day cross-sectional rank
                        - { class: ZScoreNorm, kwargs: {} }
                        - { class: Fillna, kwargs: {} }
            segments:
                train: [2015-01-03, 2025-09-01]
                valid: [2025-09-03, 2026-01-03]
                test:  [2026-01-04, 2026-08-10]
    record:                   # each entry records one artifact type to the run
        - { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} }
        - { class: SigAnaRecord, module_path: qlib.workflow.record_temp,
            kwargs: { ana_long_short: true, ann_scaler: 252 } }
        - { class: PortAnaRecord, module_path: qlib.workflow.record_temp,
            kwargs: { config: { strategy: <STRATEGY BLOCK>, backtest: {...} }, risk_analysis_freq: 1d } }

rd_run_workflow config_path=<yaml> experiment_name=<exp> runs it; the MCP call may time out for long runs (RankIC tuning, heavy feature sets) — the run keeps executing; poll via rd_exp_list / rd_exp_get_run on the returned experiment.

Experiment traceability (DB + git + embeddings)

Every backtest you run as an agent MUST be tracked: it runs as a workflow with the record block (SignalRecord/SigAnaRecord/PortAnaRecord → MLflow artifacts on disk under <lake>/mlruns/<exp_id>/<run_id>), and a row is written to the Postgres experiments table plus a git branch per experiment. The tac-app UI owns the schema (Drizzle migrations in tac-app/drizzle/); this skill's lib/ scripts are the executor the agent drives.

Trigger the lineage as part of the run — automatically, not on prompt. Any time you execute a qlib workflow (rd_run_workflow) or a train/predict pipeline on this stack, the traceability bookkeeping is part of that run, not a separate step the user must ask for: open the traced experiment with rd_trace_start before running, commit intermediates with rd_trace_commit, and close it with rd_trace_finish after — without waiting to be prompted (see "The per-experiment procedure" below).

Env vars

Var Purpose
DATABASE_URL Postgres URL for the rd_experiments table AND the MLflow tracking store (set in repo .env)
EMBEDDING_API_BASE_URL embedding POST endpoint (e.g. https://embd.h.lizhao.net/embeddings)
EMBEDDING_API_KEY basic-auth credential (user:pass form is supported)
GIT_USER / GIT_PASS git remote credentials for push/fetch
GIT_REPO_URL experiment git repo tracked by the experiments submodule (branches are pushed here)
TAC_LAKE_DIR lake root (mlruns artifact files live under it)

The experiment repo is the experiments git submodule at the workspace root (<repo-root>/experiments), always tracking $GIT_REPO_URL. rd_trace_init creates/validates it; it errors if experiments/ exists but points at a different URL. There is no TAC_EXP_GIT_DIR — the submodule path IS the experiment repo, and ALL experiment/backtest changes (workflow YAMLs, notes, outputs) must live inside it, never in the parent tradeac repo.

The rd_experiments table

Owned by tac-app's Drizzle schema (tac-app/src/db/schema.ts); rd_trace_init can init it idempotently. The table is named rd_experiments (NOT experiments) because MLflow's Postgres tracking store creates its own experiments table in the same database. Key columns: id (PK), rational + rational_embedding (pgvector vector(384)), details + details_embedding, evaluation, metrics (jsonb), evolved_from (FK → rd_experiments.id), start_ts/end_ts, git_branch, experiment_ref_id, mlruns_dir, status.

experiment_ref_id holds the mlflow run id returned by rd_run_workflow and is an FK to MLflow's runs(run_uuid) (added by rd_trace_init after the mlflow store tables exist — MLflow creates runs lazily).

Tracking store: Postgres $DATABASE_URL (MLflow's own tables) when set, falling back to the unified lake sqlite sqlite:///<lake>/mlruns.db. Artifact files always stay on disk under <lake>/mlruns/<exp_id>/<run_id>/artifacts.

Embedding model: michaelfeil/bge-small-en-v1.5 (384-dim, 512-token context). Rational/details are written paper-summary style (≤512 tokens) and embedded verbatim — NEVER truncate; if a text is longer, summarize it first (the embed helper rejects over-limit input).

Git repo + branch-per-experiment

The experiment repo is the experiments submodule at the workspace root (<repo-root>/experiments, tracking $GIT_REPO_URL). The rd_trace_* MCP tools handle it, and every git operation is scoped to that submodule — experiments NEVER stage or push parent-repo (tradeac) files.

  • rd_trace_init creates/validates the submodule and the base branch. If experiments/ does not exist it runs git clone $GIT_REPO_URL experiments; if it exists but tracks a different URL, init errors out.
  • Base branch: main (or master). If the submodule is empty, a seed commit is made and pushed so there are commits to fork from.
  • Every experiment runs on its own branch exp/<id>-<slug>.
  • evolved_from resolution (in order):
    1. If the wizard prompt explicitly says evolved_from=<id> (run wizard click on an existing experiment) — use that id directly.
    2. Otherwise --evolved-from auto: the user prompt / rational is embedded and cosine-searched over the experiments.rational_embedding column; the top hit above the similarity threshold (0.5) becomes evolved_from.
    3. Otherwise (first experiment, or a new chat with no predecessor) — no evolved_from; fork from main's latest commits.
  • The new branch is forked from the evolved-from experiment's branch (its latest commits), or from main when there is no predecessor — so experiment lineages form a git branch chain.
  • On every finish, and for intermediate steps, changes are committed + pushed.

Custom code is part of the lineage (code snapshot)

Custom contrib modules (tac_qlib/contrib/model/, tac_qlib/contrib/strategy/, tac_qlib/contrib/data/, tac_qlib/data/providers.py) live in the parent tradeac repo, not in the experiments/ submodule — so they are normally invisible to the experiment branch and a descendant forking from it would reinvent them. The lineage tooling fixes this: every experiment branch carries a code/ snapshot of exactly the qlib extension code that run depended on, so descendants reuse it instead of re-authoring it.

  • rd_trace_start and rd_trace_finish automatically snapshot the default paths (tac-qlib/tac_qlib/contrib, tac-qlib/tac_qlib/data) into <experiments>/code/<parent-relative-path> on the experiment branch.
  • rd_trace_snapshot snapshots mid-run (e.g. after writing a new custom model) without waiting for finish.
  • The snapshot also writes code/MANIFEST.txt recording the parent-repo HEAD commit and the per-file blob hashes it was taken from — so a run can be traced back to the exact parent commit that produced its custom code.
  • Descendants: the custom modules your run needs are under code/tac_qlib/... on the evolved-from branch. Reuse them (copy/git show) instead of writing new ones; check code/MANIFEST.txt to see which parent commit they came from and port fixes back.
  • Guardrail exception: parent-repo changes under tac_qlib/tac_qlib/contrib and tac_qlib/tac_qlib/data are expected (they are the snapshotted code); parent_changes reports them as a note, not a violation. Any OTHER parent change is still a guardrail violation.

Guardrail — experiments must NOT introduce side effects to the parent repo:

  • Write workflow YAMLs, notes and experiment outputs ONLY inside <repo-root>/experiments/ (they are committed on the experiment branch).
  • Never git add/commit/stage anything in the parent tradeac repo.
  • Run rd_trace_guard to list any parent changes outside the submodule pointer; rd_trace_finish also surfaces them. Revert any accidental parent edits before finishing.
  • If an experiment reveals a PRODUCT change (workflow template, skill, tac-app), propose it separately for the tradeac repo — do not mix it into the experiment branch.

The rd_trace_* MCP tools perform git operations with the mandated credential helper (from GIT_USER / GIT_PASS), so you do not need to construct it by hand.

The per-experiment procedure

Use the rd_trace_* MCP tools (tac-qlib-rd) — they replace the old trace.sh/trace_db.py scripts. The server is long-lived (psycopg imported once, DB connection reused per call) and every tool returns one JSON object, so no output parsing is needed:

# 0. ensure ready (rd_experiments table + experiments git repo + base main)
rd_trace_init

# 1. start — inserts the row, resolves evolved_from, forks+pushes the branch.
#    Returns {experiment_id, branch, evolved_from, base_branch} as JSON.
rd_trace_start  rational="5-day forward label, RankIC early stop, 50-ETF universe" \
                details="LGBModel mse lr=0.02 num_leaves=15 num_boost_round=3000; TopkDropout topk=2; benchmark QQQ" \
                experiment_name="tac-rd-expN" \
                evolved_from="auto" \
                session_id="<this chat's opencode session id, if started from a chat>"
# -> {"experiment_id": N, "branch": "exp/N-...", "evolved_from": ..., "base_branch": ...}

# 2. write the workflow YAML INSIDE the experiments submodule
#    (e.g. <repo-root>/experiments/workflows/<exp>/workflow.yaml), then commit it:
rd_trace_commit  experiment_id=<N>  message="add workflow yaml"

# 2b. if the workflow uses a NEW custom module, snapshot it onto the branch
#     (start/finish auto-snapshot contrib+data; do this to capture mid-run):
rd_trace_snapshot  experiment_id=<N>                    # default contrib+data
#    or: rd_trace_snapshot experiment_id=<N> paths="tac-qlib/tac_qlib/contrib/model/rank_gbdt.py"

# 3. run the backtest through the WORKFLOW with the recorder (MUST write mlruns):
rd_run_workflow config_path=<repo-root>/experiments/workflows/<exp>/workflow.yaml experiment_name=tac-rd-expN
# -> returns run_id (= experiment_ref_id) + metrics

# 4. inspect with rd_exp_result / rd_exp_blotter, then finish — updates the row
#    (re-embeds rational/details, sets metrics/eval/end_ts), snapshots the custom
#    code, and commits+pushes. finish also surfaces parent-repo side effects.
rd_trace_finish  experiment_id=<N> \
                 ref_id=<mlflow-run-id> \
                 evaluation="IC 0.0645, RankIC 0.075; net excess +0.85% ann" \
                 metrics='{"IC":0.0645,"RankIC":0.075,"ann_excess":0.85}' \
                 mlruns_dir=<lake>/mlruns/<exp_id>/<run_id>

Helpers (MCP tools): rd_trace_search (semantic), rd_trace_get (one row), rd_trace_list, rd_trace_mlruns_dir (resolves the mlruns dir for an experiment name), rd_trace_guard (parent-repo side-effect check).

Rules:

  • Always run backtests as workflows with the record block (req 2) — never a bare rd_backtest for a traced experiment.
  • Always open the lineage (rd_trace_start) BEFORE the run and Always rd_trace_finish + push after it completes (req 5) — this happens as part of the run, do not wait for the user to ask; intermediate rd_trace_commit is encouraged (req 5).
  • Always snapshot the custom qlib code (rd_trace_snapshot, or rely on the auto-snapshot at start/finish) so the experiment branch carries the exact contrib/data modules the run used — descendants fork and reuse code/ instead of reinventing it.
  • Keep rational/details ≤ 512 tokens (paper-summary style) so embeddings are exact — no truncation.
  • Confine experiments to the experiments/ submodule — never write to, stage, or commit parent tradeac repo files; run rd_trace_guard to check for side effects. (Custom code edits under tac-qlib/tac_qlib/contrib and .../data are the sanctioned exception — they are the snapshotted modules; see "Custom code is part of the lineage".)
  • Follow the Secrets policy above — no secrets in files, no reading .env*, ask the user to set env vars; use uri: "sqlite:///mlruns.db" for the tracking store.
  • Workflow YAMLs are jinja-rendered with os.environ as the context, so env-var placeholders work ({%- set LAKE = TAC_LAKE_DIR %} then {{ LAKE }}). Use them for paths/config — never for secrets that get committed.

Extending Qlib — the 4 class families you can override

1. Custom Model (train-time) — tac_qlib/contrib/model/

Subclass qlib.contrib.model.gbdt.LGBModel (or qlib.model.base.BaseModel) and implement fit(dataset, ...) + predict(dataset). LGBModel.fit calls self._prepare_data(dataset) → lgb.Datasets, then lgb.train with early_stopping on the valid set. Override points that matter:

  • _prepare_data → build the lgb.Dataset with group= (per-day query groups) when you need ranking metrics per trading day.
  • fit → change what early-stops training (the biggest IC/backtest lever, see §Knobs).
  • predict → return the Series keyed (datetime, instrument).

Reference: tac_qlib/tac_qlib/contrib/model/rank_gbdt.py — RankICLGBModel subclasses LGBModel, adds per-day group in _prepare_data, injects feval=rankic_feval (mean per-day Spearman) into lgb.train, and forces metric='None' + first_metric_only=True so early-stopping tracks RankIC only.

2. Custom Strategy (backtest-time) — tac_qlib/contrib/strategy/

Subclass qlib.contrib.strategy.signal_strategy.BaseSignalStrategy (which wraps qlib.strategy.base.BaseStrategy) and implement:

def generate_trade_decision(self, execute_result=None):
    # trade_step, trade_start/end = self.trade_calendar.get_step_time(trade_step)
    # pred = self.signal.get_signal(start_time=pred_shift, end_time=pred_shift)  # shift=-1 => signal known at t-1
    # self.trade_position / self.trade_exchange / self.trade_calendar injected by the executor
    # build qlib.backtest.Order(stock_id, amount, start_time, end_time, direction=Order.BUY/SELL)
    # return TradeDecisionWO(orders, self)

Wire it into the YAML under PortAnaRecord.config.strategy:

strategy:
    class: OptimalStopControl
    module_path: tac_qlib.contrib.strategy.optimal_stop
    kwargs:
        signal: "<PRED>"     # placeholder replaced with the recorded pred
        topk: 10
        entry_pct: 0.85
        exit_pct: 0.7
        max_hold_days: 10
        min_hold_days: 2
        sl: -0.08
        risk_degree: 0.95

Reference: tac_qlib/tac_qlib/contrib/strategy/optimal_stop.py (OptimalStopControl — entry gated by cross-sectional signal percentile, exits by percentile/time/stop-loss, equal-weight control sizing).

3. Custom DataHandler / processors — tac_qlib/contrib/data/handler.py

TACHandler(DataHandlerLP) already wraps the lake via QlibDataLoader + LakeFeatureProvider. Key config surface (all usable from YAML without new code):

  • feature_fields — explicit list; the handler prefixes $ and de-dups. Anything the provider can route is usable: bar fields, $amount (v*vw), and any column present in the lake features/.../symbol=*.parquet files.
  • infer_processors / learn_processors — add CSRankNorm, CSZScoreNorm (label), ZScoreNorm, DropnaLabel, Fillna, etc. DropAllNaN is a tac-qlib processor (drops all-NaN columns on the fit window).
  • label — any qlib expression, e.g. Ref($close,-6)/Ref($close,-1)-1.

To add a new feature family: compute it once (see examples/sp_features.py + examples/persist_sp_features.py), persist extra columns into features/market=US/timeframe=1d/symbol=*.parquet (drop stale sp_* columns first on re-runs), then reference them in feature_fields.

The Rust engine already ships the SP feature pipeline as a lake MCP tool: get_lake_sp (tac-engine, stochastic-rs) computes sp_ou_*, sp_hmm_*, sp_jump_*, sp_rv*/sp_vol_ratio_* (+ sp_rv_ac1, sp_rv_cv_22), sp_max_up/sp_max_down, sp_trend_slope_*, sp_logp, sp_hurst_exponent, sp_sig_* (levels 1/2 at lag 1 and 5), sp_rskew_*/sp_rkurt_*/sp_dsv_* (realized moments via stochastic-rs realized) + sp_ret from lake bars and persists them into the feature parquets (replacing stale sp_*), all in one call:

{"symbol": "AAPL", "timeframe": "1d", "start": "2015-01-03", "end": "2026-08-10", "fit_end": "2025-09-01"}

fit_end pins the Gaussian-HMM fit to the train window (no lookahead), matching the FIT_END convention. Deferred families (garch, entropy, catch22) are still computed with the Python sp_features.py path until their ports land. Note two deliberate differences vs the Python reference: the Rust HMM uses the causal forward filter (filtered_state_probs) rather than hmmlearn's smoothed predict_proba, and hurst is estimated on the returns series directly (take_differences=false) rather than the reference's double-differenced kind="random_walk" — regime state assignments agree, probability levels are comparable but not identical.

4. Custom Record (artifact writers)

Subclass qlib.workflow.record_temp.SignalRecord / a Record and log metrics + artifacts into the MLflow run. There is no shipped example Record in contrib/ yet — write one against the pattern in qlib.workflow.record_temp when a workflow needs a bespoke simulator (e.g. beta-neutral 3L/3S) that PortAnaRecord doesn't cover.

Empirical knobs that moved the numbers (measured on the 50-ETF lake)

All experiments used: 50-ETF universe, train 2015-01-03..2025-09-01 / valid 2025-09-03..2026-01-03 / test 2026-01-04..2026-08-10, benchmark SPY, TopkDropout or OptimalStopControl, costs open 0.0005 / close 0.0015 / min 5.

Rank-dimension reminder: when the goal is to improve the ranking quality of a signal (RankIC, long-short spread, top-decile precision), do NOT reinvent the stack — use the contrib modules already shipped and verified in this repo: tac_qlib.contrib.model.rank_gbdt.RankICLGBModel (early-stops training on per-day cross-sectional RankIC, metric='None' + first_metric_only) and tac_qlib.contrib.strategy.optimal_stop.OptimalStopControl (entry/exit gated by signal percentile instead of raw levels). Both are loadable from a workflow YAML via module_path — see the canonical tac-qlib/workflows/workflow_lgb_sp5d_rankic.yaml (rank dimension: model) and workflow_lgb_sp5d_optstop.yaml (rank dimension: portfolio construction). Verified end-to-end on 2026-01-04..2026-08-10: RankIC 0.071 / net-of-cost excess +20.7% ann (IR 0.70) vs SPY. Only write a new custom Model/Strategy when these proven paths are insufficient.

Label

  • 5-day forward return Ref($close,-6)/Ref($close,-1)-1 ≫ 2-day. IC nearly tripled (0.0207 → 0.0645 standalone; the biggest single lever found). The 2-day target is too noisy.

Features

  • Stochastic-process features beat hand-rolled TA. 55-feature set: OU (sp_ou_*), 2-state HMM (sp_hmm_*), jump intensity (sp_jump_*, incl. sp_max_up/sp_max_down), HARRV vol (sp_rv* + sp_rv_ac1/sp_rv_cv_22), trend (sp_trend_slope_*, sp_logp), GARCH (sp_garch_*), Hurst (sp_hurst_exponent), path signatures (sp_sig_*, lag 1 & 5), entropy (sp_ent_*), realized moments (sp_rskew_*/sp_rkurt_*/sp_dsv_*), catch22 (sp_c22_*). IC 0.036 → 0.047 vs the 19-feature v1.
  • Do NOT add ta-lib indicators on top (SP+TA, 74 feats): IC dropped 0.047 → 0.031, RankIC 0.047 → 0.020. They're redundant with rv22/hmm/garch/catch22 and dilute CSRankNorm + LGBM.
  • CSRankNorm (per-day cross-sectional rank) is important for the rank signal.
  • Warm-up rows persist as all-NaN feature rows — expected; DropAllNaN/DropnaLabel handle them.

Model / training loop

  • LambdaRank / rank_xendcg objectives FAIL here (RankIC → ~0): with only ~50 "documents" per query the rank gradient is noise.
  • Early-stopping metric beats objective. MSE objective + early-stop on a RankIC feval (mean per-day Spearman) lifted RankIC 0.047 → 0.075 (standalone).
  • The workflow gap was qlib's training loop: lgb.train default first_metric_only=False + metric=l2 keeps training while l2 improves after RankIC peaks. RankICLGBModel sets metric='None' + first_metric_only=True so early-stopping tracks RankIC only.
  • RankIC-only early stop + bigger/smaller budget is the win: num_boost_round 3000, learning_rate 0.02, early_stopping_rounds 200, min_data_in_leaf 20, lambda_l2 0.5 → test excess +9.1% ann w/o cost (IR 1.03, maxDD −3.8%) and +0.85% ann after costs — the only config that beat SPY net. Note IC/RankIC themselves were slightly lower (0.042) than the 500-tree run (0.051); the tuned budget selects the iteration maximizing valid RankIC, converting to realized excess return.

Strategy / portfolio construction

  • Long-only construction leaves the edge on the table. The SP-5d signal has long-short +31.6% ann (Sharpe 2.51), but TopkDropout long-only ≈ flat vs SPY, and OptimalStopControl underperformed (valid-window threshold overfit: valid +7.5% → test −17.7% on one calibration).
  • Costs eat most of the gross edge (+9.1% → +0.85% net). Reduce turnover or go long-short to widen the net edge.
  • OptimalStopControl thresholds must be calibrated on the valid window and are sensitive to overfit — prefer robust defaults or penalize turnover in selection.

Gotchas

  • Installed package copy: tac_qlib in the venv is a copy under /opt/venv/lib/python3.12/site-packages/tac_qlib/. After editing any tac_qlib/contrib/** module, cp it there or the workflow imports the stale version. New subpackages need mkdir -p first.
  • qlib.backtest exports Order but not OrderDir/Position at top level — import Order from qlib.backtest, OrderDir/TradeDecisionWO from qlib.backtest.decision, Position from qlib.backtest.position.
  • qlib.backtest.high_performance_ds may not export Order in this build — don't import from it.
  • HMM / GARCH / catch22 features must not see test data at fit time: fit the HMM on the train window only (fit_end=FIT_END), and compute rolling windows ending at each day. GARCH/entropy use a stride + forward-fill for speed (~5x).
  • pycatch22, arch, hurst, antropy, hmmlearn are required for the full feature set; install with uv pip install --python /app/.venv/bin/python <pkg> (a C compiler is needed for pycatch22). duckdb and pyarrow are declared in tac-qlib/pyproject.toml; if a workflow import fails on either, lazy-install with uv pip install --python /app/.venv/bin/python duckdb pyarrow.
  • rd_run_workflow defaults to wait=false: it returns immediately with status: started and the workflow runs in a background thread — poll rd_exp_get_run / rd_exp_list for the newest run of the experiment (status RUNNING until it finishes), then reuse its run_id. Pass wait=true only for small windows that finish within the MCP call timeout.
  • After fixing a YAML model/handler change, remember both /app/tac-qlib/... and the /opt/venv copy stay in sync.

Files this skill is based on

Minimal, runnable examples live next to this skill in examples/ — they are the canonical reference for every artifact the skill describes:

  • Workflows (full record block → MLflow on disk):
    • examples/workflow_minimal.yaml — the canonical backtest template (req: every traced backtest runs through a workflow like this via rd_run_workflow)
    • examples/workflow_rankic.yaml — RankIC-early-stop model wired in
    • Repo workflows for reference: tac-qlib/workflows/workflow_lgb_taclake.yaml, tune_run1_wider_5d.yaml, tune_run2_regularized.yaml, tune_run3_label5d_clean_universe.yaml, tune_run4_fix_universe_longtrain.yaml, tune_run5_longtest.yaml
  • Models: examples/model_rank_gbdt.py (RankICLGBModel: per-day groups + feval=rankic + metric='None'). Repo: tac_qlib/contrib/model/rank_gbdt.py
  • Strategies: examples/strategy_optimal_stop.py (OptimalStopControl), examples/strategy_beta_neutral.py (doc-only 3L/3S stub — pattern for a custom strategy + Record; not wired into the package)
  • Handler: examples/handler.py (how to subclass TACHandler); repo: tac_qlib/contrib/data/handler.py; providers: tac_qlib/data/providers.py
  • Feature engineering: examples/sp_features.py (OU + Hurst) and examples/persist_sp_features.py (persist sp_* into the lake features parquet)
  • Ranking experiments: examples/run_rank_objectives.py (mse vs lambdarank vs rank_xendcg ablation on the lake)
  • Optstop calibration: examples/run_optstop_compare.py (valid-window grid + overfit warning)
  • Traceability tooling: the rd_trace_* MCP tools (tac-qlib-rd, tac_qlib/trace.py) — see the traceability section above