--- name: tac-qlib-custom description: "Guide agents to customize and extend Qlib on the TradeAC R&D stack — how to configure workflow YAMLs (qlib_init, model, dataset/handler, processors, records, PortAnaRecord strategies), how to extend Qlib classes wired into those workflows (custom Model, BaseStrategy, DataHandler, Record), and the empirically-tested knobs from this repo (RankIC early-stopping, stochastic-control strategies, stochastic-process features, catch22/GARCH/Hurst/signature). Also encodes the experiment traceability loop: every backtest runs as a workflow-with-recorder, is recorded in the Postgres experiments table (rationale/details/evaluation/metrics with pgvector embeddings, evolution chain) and on a per-experiment git branch that is committed + pushed. Companion to tradeac-rd (MCP run tools) and tradeac-lake (parquet lake)." --- # tac-qlib-custom Customizing and extending Qlib on the TradeAC stack. This skill encodes what was learned from actual experiments in this repo: how a workflow YAML maps to Qlib classes, how to write a custom class that the YAML can load, and which training / strategy / feature knobs measurably moved IC, RankIC and the backtest. Read `tac-qlib/skills/tradeac-rd/SKILL.md` for the MCP run/inspect tools and `tac-qlib/README.md` for the package layout. The venv is `/app/.venv` (qlib 0.1.dev2066); `tac_qlib` is installed into the venv's `site-packages` (editable copy under `/opt/venv/.../tac_qlib/`), so **any new module must be copied to `/opt/venv/lib/python3.12/site-packages/tac_qlib/...` too** (or use an editable install) before `rd_run_workflow` can import it. ## MCP-first policy - **Drive every backtest and run through the `tac-qlib-rd` MCP tools** (`rd_run_workflow`, `rd_train`, `rd_predict`, `rd_exp_*`) and the tac-engine lake tools for data prep. Do not reimplement them with ad-hoc scripts (custom qlib glue, own mlruns readers, direct JSON-RPC/stdio clients). - **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server` / `tac-engine`, bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**. - The traceability bookkeeping (Postgres `rd_experiments` row + pgvector embeddings + branch-per-experiment git) is exposed as the **`rd_trace_*` MCP tools** on the tac-qlib-rd server — use those, not bash scripts. Data prep, training, evaluation and backtests also go through MCP tools. - If the venv is missing a runtime dep (`duckdb`, `pyarrow`, feature libs), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python `) instead of switching tools. ## Secrets policy - NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `GIT_PASS`, `EMBEDDING_API_KEY`) in workflow YAMLs, scripts, configs, notes or committed code. - NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself. - Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows — `rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else the lake sqlite. Never hardcode a `postgres://user:pass@…` URI. - If you find a committed secret, flag it, remove it, and replace it with a placeholder. (The `rd_trace_*` MCP tools' commit guard blocks adding credential-shaped lines.) ## How a workflow YAML maps to Qlib classes A workflow YAML (`tac-qlib/workflows/*.yaml`) is rendered by Jinja (vars like `{{ LAKE }}` from `TAC_LAKE_DIR`) then executed by `qrun` / `rd_run_workflow`. Every block is a Qlib class reference resolved by `module_path` + `class`: ```yaml {%- set LAKE = TAC_LAKE_DIR %} qlib_init: provider_uri: "{{ LAKE }}" region: us calendar_provider: # custom tac-qlib providers read the parquet lake class: LakeCalendarProvider module_path: tac_qlib.data.providers instrument_provider: # ... (markets: {} => lake universe) feature_provider: # LakeFeatureProvider: routes $open..$volume from bars, class: LakeFeatureProvider # $ from features parquet, $amount derived exp_manager: class: MLflowExpManager module_path: qlib.workflow.expm kwargs: { uri: "sqlite:///{{ LAKE }}/mlruns.db", default_exp_name: "my-exp" } task: model: # — custom model → new module_path class: RankICLGBModel module_path: tac_qlib.contrib.model.rank_gbdt kwargs: { loss: mse, learning_rate: 0.02, num_leaves: 31, ... } dataset: class: DatasetH module_path: qlib.data.dataset kwargs: handler: # — feature selection + processors live here class: TACHandler module_path: tac_qlib.contrib.data.handler kwargs: instruments: "SPY,QQQ,..." start_time: 2015-01-03 end_time: 2026-08-10 fit_start_time: 2015-01-03 # processors fit on this window fit_end_time: 2025-09-01 freq: day lake_root: "{{ LAKE }}" market: US label: "Ref($close,-6)/Ref($close,-1)-1" # 5d forward return feature_fields: "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_ou_zscore,..." infer_processors: # feature-time transforms, fit on fit_* - { class: DropAllNaN, kwargs: {} } - { class: ProcessInf, kwargs: {} } - { class: CSRankNorm, kwargs: {} } # per-day cross-sectional rank - { class: ZScoreNorm, kwargs: {} } - { class: Fillna, kwargs: {} } segments: train: [2015-01-03, 2025-09-01] valid: [2025-09-03, 2026-01-03] test: [2026-01-04, 2026-08-10] record: # each entry records one artifact type to the run - { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} } - { class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: { ana_long_short: true, ann_scaler: 252 } } - { class: PortAnaRecord, module_path: qlib.workflow.record_temp, kwargs: { config: { strategy: , backtest: {...} }, risk_analysis_freq: 1d } } ``` `rd_run_workflow config_path= experiment_name=` runs it; the MCP call may time out for long runs (RankIC tuning, heavy feature sets) — the run keeps executing; poll via `rd_exp_list` / `rd_exp_get_run` on the returned experiment. ## Experiment traceability (DB + git + embeddings) Every backtest you run as an agent MUST be tracked: it runs as a workflow with the `record` block (SignalRecord/SigAnaRecord/PortAnaRecord → MLflow artifacts on disk under `/mlruns//`), and a row is written to the Postgres `experiments` table plus a git branch per experiment. The `tac-app` UI owns the schema (Drizzle migrations in `tac-app/drizzle/`); this skill's `lib/` scripts are the executor the agent drives. **Trigger the lineage as part of the run — automatically, not on prompt.** Any time you execute a qlib workflow (`rd_run_workflow`) or a train/predict pipeline on this stack, the traceability bookkeeping is part of that run, not a separate step the user must ask for: open the traced experiment with `rd_trace_start` before running, commit intermediates with `rd_trace_commit`, and close it with `rd_trace_finish` after — without waiting to be prompted (see "The per-experiment procedure" below). ### Env vars | Var | Purpose | |-----|---------| | `DATABASE_URL` | Postgres URL for the `rd_experiments` table AND the MLflow tracking store (set in repo `.env`) | | `EMBEDDING_API_BASE_URL` | embedding POST endpoint (e.g. `https://embd.h.lizhao.net/embeddings`) | | `EMBEDDING_API_KEY` | basic-auth credential (`user:pass` form is supported) | | `GIT_USER` / `GIT_PASS` | git remote credentials for push/fetch | | `GIT_REPO_URL` | experiment git repo tracked by the `experiments` submodule (branches are pushed here) | | `TAC_LAKE_DIR` | lake root (mlruns artifact files live under it) | The experiment repo is the **`experiments` git submodule** at the workspace root (`/experiments`), always tracking `$GIT_REPO_URL`. `rd_trace_init` creates/validates it; it errors if `experiments/` exists but points at a different URL. There is no `TAC_EXP_GIT_DIR` — the submodule path IS the experiment repo, and ALL experiment/backtest changes (workflow YAMLs, notes, outputs) must live inside it, never in the parent tradeac repo. ### The `rd_experiments` table Owned by tac-app's Drizzle schema (`tac-app/src/db/schema.ts`); `rd_trace_init` can `init` it idempotently. The table is named **`rd_experiments`** (NOT `experiments`) because MLflow's Postgres tracking store creates its own `experiments` table in the same database. Key columns: `id` (PK), `rational` + `rational_embedding` (pgvector `vector(384)`), `details` + `details_embedding`, `evaluation`, `metrics` (jsonb), `evolved_from` (FK → rd_experiments.id), `start_ts`/`end_ts`, `git_branch`, `experiment_ref_id`, `mlruns_dir`, `status`. `experiment_ref_id` holds the **mlflow run id** returned by `rd_run_workflow` and is an FK to MLflow's `runs(run_uuid)` (added by `rd_trace_init` after the mlflow store tables exist — MLflow creates `runs` lazily). Tracking store: **Postgres `$DATABASE_URL`** (MLflow's own tables) when set, falling back to the unified lake sqlite `sqlite:////mlruns.db`. Artifact files always stay on disk under `/mlruns///artifacts`. Embedding model: `michaelfeil/bge-small-en-v1.5` (384-dim, **512-token context**). Rational/details are written paper-summary style (≤512 tokens) and embedded verbatim — NEVER truncate; if a text is longer, summarize it first (the embed helper rejects over-limit input). ### Git repo + branch-per-experiment The experiment repo is the `experiments` submodule at the workspace root (`/experiments`, tracking `$GIT_REPO_URL`). The `rd_trace_*` MCP tools handle it, and every git operation is scoped to that submodule — experiments NEVER stage or push parent-repo (tradeac) files. - `rd_trace_init` creates/validates the submodule and the base branch. If `experiments/` does not exist it runs `git clone $GIT_REPO_URL experiments`; if it exists but tracks a different URL, init errors out. - Base branch: `main` (or `master`). If the submodule is empty, a seed commit is made and pushed so there are commits to fork from. - Every experiment runs on its own branch `exp/-`. - `evolved_from` resolution (in order): 1. If the wizard prompt explicitly says `evolved_from=` (run wizard click on an existing experiment) — use that id directly. 2. Otherwise `--evolved-from auto`: the user prompt / rational is embedded and cosine-searched over the `experiments.rational_embedding` column; the top hit above the similarity threshold (0.5) becomes `evolved_from`. 3. Otherwise (first experiment, or a new chat with no predecessor) — no evolved_from; fork from `main`'s latest commits. - The new branch is forked from the **evolved-from experiment's branch** (its latest commits), or from `main` when there is no predecessor — so experiment lineages form a git branch chain. - On every finish, and for intermediate steps, changes are committed + pushed. ### Custom code is part of the lineage (code snapshot) Custom contrib modules (`tac_qlib/contrib/model/`, `tac_qlib/contrib/strategy/`, `tac_qlib/contrib/data/`, `tac_qlib/data/providers.py`) live in the **parent** tradeac repo, not in the `experiments/` submodule — so they are normally invisible to the experiment branch and a descendant forking from it would reinvent them. The lineage tooling fixes this: **every experiment branch carries a `code/` snapshot of exactly the qlib extension code that run depended on**, so descendants reuse it instead of re-authoring it. - `rd_trace_start` and `rd_trace_finish` automatically snapshot the default paths (`tac-qlib/tac_qlib/contrib`, `tac-qlib/tac_qlib/data`) into `/code/` on the experiment branch. - `rd_trace_snapshot` snapshots mid-run (e.g. after writing a new custom model) without waiting for finish. - The snapshot also writes `code/MANIFEST.txt` recording the **parent-repo HEAD commit** and the per-file blob hashes it was taken from — so a run can be traced back to the exact parent commit that produced its custom code. - Descendants: the custom modules your run needs are under `code/tac_qlib/...` on the evolved-from branch. Reuse them (copy/`git show`) instead of writing new ones; check `code/MANIFEST.txt` to see which parent commit they came from and port fixes back. - Guardrail exception: parent-repo changes under `tac_qlib/tac_qlib/contrib` and `tac_qlib/tac_qlib/data` are **expected** (they are the snapshotted code); `parent_changes` reports them as a note, not a violation. Any OTHER parent change is still a guardrail violation. Guardrail — experiments must NOT introduce side effects to the parent repo: - Write workflow YAMLs, notes and experiment outputs ONLY inside `/experiments/` (they are committed on the experiment branch). - Never `git add`/commit/stage anything in the parent tradeac repo. - Run `rd_trace_guard` to list any parent changes outside the submodule pointer; `rd_trace_finish` also surfaces them. Revert any accidental parent edits before finishing. - If an experiment reveals a PRODUCT change (workflow template, skill, tac-app), propose it separately for the tradeac repo — do not mix it into the experiment branch. The `rd_trace_*` MCP tools perform git operations with the mandated credential helper (from `GIT_USER` / `GIT_PASS`), so you do not need to construct it by hand. ### The per-experiment procedure **Use the `rd_trace_*` MCP tools (tac-qlib-rd)** — they replace the old `trace.sh`/`trace_db.py` scripts. The server is long-lived (psycopg imported once, DB connection reused per call) and every tool returns one JSON object, so no output parsing is needed: ```text # 0. ensure ready (rd_experiments table + experiments git repo + base main) rd_trace_init # 1. start — inserts the row, resolves evolved_from, forks+pushes the branch. # Returns {experiment_id, branch, evolved_from, base_branch} as JSON. rd_trace_start rational="5-day forward label, RankIC early stop, 50-ETF universe" \ details="LGBModel mse lr=0.02 num_leaves=15 num_boost_round=3000; TopkDropout topk=2; benchmark QQQ" \ experiment_name="tac-rd-expN" \ evolved_from="auto" \ session_id="" # -> {"experiment_id": N, "branch": "exp/N-...", "evolved_from": ..., "base_branch": ...} # 2. write the workflow YAML INSIDE the experiments submodule # (e.g. /experiments/workflows//workflow.yaml), then commit it: rd_trace_commit experiment_id= message="add workflow yaml" # 2b. if the workflow uses a NEW custom module, snapshot it onto the branch # (start/finish auto-snapshot contrib+data; do this to capture mid-run): rd_trace_snapshot experiment_id= # default contrib+data # or: rd_trace_snapshot experiment_id= paths="tac-qlib/tac_qlib/contrib/model/rank_gbdt.py" # 3. run the backtest through the WORKFLOW with the recorder (MUST write mlruns): rd_run_workflow config_path=/experiments/workflows//workflow.yaml experiment_name=tac-rd-expN # -> returns run_id (= experiment_ref_id) + metrics # 4. inspect with rd_exp_result / rd_exp_blotter, then finish — updates the row # (re-embeds rational/details, sets metrics/eval/end_ts), snapshots the custom # code, and commits+pushes. finish also surfaces parent-repo side effects. rd_trace_finish experiment_id= \ ref_id= \ evaluation="IC 0.0645, RankIC 0.075; net excess +0.85% ann" \ metrics='{"IC":0.0645,"RankIC":0.075,"ann_excess":0.85}' \ mlruns_dir=/mlruns// ``` Helpers (MCP tools): `rd_trace_search` (semantic), `rd_trace_get` (one row), `rd_trace_list`, `rd_trace_mlruns_dir` (resolves the mlruns dir for an experiment name), `rd_trace_guard` (parent-repo side-effect check). Rules: - **Always** run backtests as workflows with the `record` block (req 2) — never a bare `rd_backtest` for a traced experiment. - **Always** open the lineage (`rd_trace_start`) BEFORE the run and **Always** `rd_trace_finish` + push after it completes (req 5) — this happens as part of the run, do not wait for the user to ask; intermediate `rd_trace_commit` is encouraged (req 5). - **Always** snapshot the custom qlib code (`rd_trace_snapshot`, or rely on the auto-snapshot at start/finish) so the experiment branch carries the exact contrib/data modules the run used — descendants fork and reuse `code/` instead of reinventing it. - Keep rational/details ≤ 512 tokens (paper-summary style) so embeddings are exact — no truncation. - **Confine experiments to the `experiments/` submodule** — never write to, stage, or commit parent tradeac repo files; run `rd_trace_guard` to check for side effects. (Custom code edits under `tac-qlib/tac_qlib/contrib` and `.../data` are the sanctioned exception — they are the snapshotted modules; see "Custom code is part of the lineage".) - **Follow the Secrets policy above** — no secrets in files, no reading `.env*`, ask the user to set env vars; use `uri: "sqlite:///mlruns.db"` for the tracking store. - Workflow YAMLs are jinja-rendered with `os.environ` as the context, so env-var placeholders work (`{%- set LAKE = TAC_LAKE_DIR %}` then `{{ LAKE }}`). Use them for paths/config — never for secrets that get committed. ## Extending Qlib — the 4 class families you can override ### 1. Custom Model (train-time) — `tac_qlib/contrib/model/` Subclass `qlib.contrib.model.gbdt.LGBModel` (or `qlib.model.base.BaseModel`) and implement `fit(dataset, ...)` + `predict(dataset)`. `LGBModel.fit` calls `self._prepare_data(dataset)` → `lgb.Dataset`s, then `lgb.train` with `early_stopping` on the valid set. Override points that matter: - `_prepare_data` → build the `lgb.Dataset` with `group=` (per-day query groups) when you need ranking metrics per trading day. - `fit` → change what early-stops training (the biggest IC/backtest lever, see §Knobs). - `predict` → return the Series keyed (datetime, instrument). Reference: `tac_qlib/tac_qlib/contrib/model/rank_gbdt.py` — `RankICLGBModel` subclasses `LGBModel`, adds per-day `group` in `_prepare_data`, injects `feval=rankic_feval` (mean per-day Spearman) into `lgb.train`, and forces `metric='None'` + `first_metric_only=True` so early-stopping tracks RankIC only. ### 2. Custom Strategy (backtest-time) — `tac_qlib/contrib/strategy/` Subclass `qlib.contrib.strategy.signal_strategy.BaseSignalStrategy` (which wraps `qlib.strategy.base.BaseStrategy`) and implement: ```python def generate_trade_decision(self, execute_result=None): # trade_step, trade_start/end = self.trade_calendar.get_step_time(trade_step) # pred = self.signal.get_signal(start_time=pred_shift, end_time=pred_shift) # shift=-1 => signal known at t-1 # self.trade_position / self.trade_exchange / self.trade_calendar injected by the executor # build qlib.backtest.Order(stock_id, amount, start_time, end_time, direction=Order.BUY/SELL) # return TradeDecisionWO(orders, self) ``` Wire it into the YAML under `PortAnaRecord.config.strategy`: ```yaml strategy: class: OptimalStopControl module_path: tac_qlib.contrib.strategy.optimal_stop kwargs: signal: "" # placeholder replaced with the recorded pred topk: 10 entry_pct: 0.85 exit_pct: 0.7 max_hold_days: 10 min_hold_days: 2 sl: -0.08 risk_degree: 0.95 ``` Reference: `tac_qlib/tac_qlib/contrib/strategy/optimal_stop.py` (`OptimalStopControl` — entry gated by cross-sectional signal percentile, exits by percentile/time/stop-loss, equal-weight control sizing). ### 3. Custom DataHandler / processors — `tac_qlib/contrib/data/handler.py` `TACHandler(DataHandlerLP)` already wraps the lake via `QlibDataLoader` + `LakeFeatureProvider`. Key config surface (all usable from YAML without new code): - `feature_fields` — explicit list; the handler prefixes `$` and de-dups. Anything the provider can route is usable: bar fields, `$amount` (v*vw), and any column present in the lake `features/.../symbol=*.parquet` files. - `infer_processors` / `learn_processors` — add `CSRankNorm`, `CSZScoreNorm` (label), `ZScoreNorm`, `DropnaLabel`, `Fillna`, etc. `DropAllNaN` is a tac-qlib processor (drops all-NaN columns on the fit window). - `label` — any qlib expression, e.g. `Ref($close,-6)/Ref($close,-1)-1`. To add a *new feature family*: compute it once (see `examples/sp_features.py` + `examples/persist_sp_features.py`), persist extra columns into `features/market=US/timeframe=1d/symbol=*.parquet` (drop stale `sp_*` columns first on re-runs), then reference them in `feature_fields`. **The Rust engine already ships the SP feature pipeline as a lake MCP tool**: `get_lake_sp` (tac-engine, stochastic-rs) computes `sp_ou_*`, `sp_hmm_*`, `sp_jump_*`, `sp_rv*`/`sp_vol_ratio_*` (+ `sp_rv_ac1`, `sp_rv_cv_22`), `sp_max_up`/`sp_max_down`, `sp_trend_slope_*`, `sp_logp`, `sp_hurst_exponent`, `sp_sig_*` (levels 1/2 at lag 1 and 5), `sp_rskew_*`/`sp_rkurt_*`/`sp_dsv_*` (realized moments via stochastic-rs `realized`) + `sp_ret` from lake bars and persists them into the feature parquets (replacing stale `sp_*`), all in one call: ```json {"symbol": "AAPL", "timeframe": "1d", "start": "2015-01-03", "end": "2026-08-10", "fit_end": "2025-09-01"} ``` `fit_end` pins the Gaussian-HMM fit to the train window (no lookahead), matching the `FIT_END` convention. **Deferred families** (`garch`, `entropy`, `catch22`) are still computed with the Python `sp_features.py` path until their ports land. Note two deliberate differences vs the Python reference: the Rust HMM uses the causal *forward filter* (`filtered_state_probs`) rather than hmmlearn's smoothed `predict_proba`, and `hurst` is estimated on the returns series directly (`take_differences=false`) rather than the reference's double-differenced `kind="random_walk"` — regime *state* assignments agree, probability levels are comparable but not identical. ### 4. Custom Record (artifact writers) Subclass `qlib.workflow.record_temp.SignalRecord` / a `Record` and log metrics + artifacts into the MLflow run. There is no shipped example Record in `contrib/` yet — write one against the pattern in `qlib.workflow.record_temp` when a workflow needs a bespoke simulator (e.g. beta-neutral 3L/3S) that `PortAnaRecord` doesn't cover. ## Empirical knobs that moved the numbers (measured on the 50-ETF lake) All experiments used: 50-ETF universe, train 2015-01-03..2025-09-01 / valid 2025-09-03..2026-01-03 / test 2026-01-04..2026-08-10, benchmark SPY, TopkDropout or OptimalStopControl, costs open 0.0005 / close 0.0015 / min 5. > **Rank-dimension reminder**: when the goal is to improve the *ranking* quality of > a signal (RankIC, long-short spread, top-decile precision), do NOT reinvent the > stack — use the contrib modules already shipped and verified in this repo: > `tac_qlib.contrib.model.rank_gbdt.RankICLGBModel` (early-stops training on > per-day cross-sectional RankIC, `metric='None'` + `first_metric_only`) and > `tac_qlib.contrib.strategy.optimal_stop.OptimalStopControl` (entry/exit gated by > signal percentile instead of raw levels). Both are loadable from a workflow YAML > via `module_path` — see the canonical `tac-qlib/workflows/workflow_lgb_sp5d_rankic.yaml` > (rank dimension: model) and `workflow_lgb_sp5d_optstop.yaml` (rank dimension: > portfolio construction). Verified end-to-end on 2026-01-04..2026-08-10: > RankIC 0.071 / net-of-cost excess +20.7% ann (IR 0.70) vs SPY. Only write a new > custom Model/Strategy when these proven paths are insufficient. ### Label - **5-day forward return `Ref($close,-6)/Ref($close,-1)-1` ≫ 2-day.** IC nearly tripled (0.0207 → 0.0645 standalone; the biggest single lever found). The 2-day target is too noisy. ### Features - **Stochastic-process features beat hand-rolled TA.** 55-feature set: OU (`sp_ou_*`), 2-state HMM (`sp_hmm_*`), jump intensity (`sp_jump_*`, incl. `sp_max_up`/`sp_max_down`), HARRV vol (`sp_rv*` + `sp_rv_ac1`/`sp_rv_cv_22`), trend (`sp_trend_slope_*`, `sp_logp`), GARCH (`sp_garch_*`), Hurst (`sp_hurst_exponent`), path signatures (`sp_sig_*`, lag 1 & 5), entropy (`sp_ent_*`), realized moments (`sp_rskew_*`/`sp_rkurt_*`/`sp_dsv_*`), catch22 (`sp_c22_*`). IC 0.036 → 0.047 vs the 19-feature v1. - **Do NOT add ta-lib indicators on top** (SP+TA, 74 feats): IC dropped 0.047 → 0.031, RankIC 0.047 → 0.020. They're redundant with rv22/hmm/garch/catch22 and dilute CSRankNorm + LGBM. - **CSRankNorm** (per-day cross-sectional rank) is important for the rank signal. - Warm-up rows persist as all-NaN feature rows — expected; DropAllNaN/DropnaLabel handle them. ### Model / training loop - **LambdaRank / rank_xendcg objectives FAIL here** (RankIC → ~0): with only ~50 "documents" per query the rank gradient is noise. - **Early-stopping metric beats objective.** MSE objective + early-stop on a **RankIC feval** (mean per-day Spearman) lifted RankIC 0.047 → 0.075 (standalone). - **The workflow gap was qlib's training loop**: `lgb.train` default `first_metric_only=False` + `metric=l2` keeps training while l2 improves after RankIC peaks. `RankICLGBModel` sets `metric='None'` + `first_metric_only=True` so early-stopping tracks RankIC only. - **RankIC-only early stop + bigger/smaller budget is the win**: `num_boost_round 3000`, `learning_rate 0.02`, `early_stopping_rounds 200`, `min_data_in_leaf 20`, `lambda_l2 0.5` → test excess **+9.1% ann w/o cost (IR 1.03, maxDD −3.8%)** and **+0.85% ann after costs** — the only config that beat SPY net. Note IC/RankIC themselves were slightly lower (0.042) than the 500-tree run (0.051); the tuned budget selects the iteration maximizing *valid* RankIC, converting to realized excess return. ### Strategy / portfolio construction - **Long-only construction leaves the edge on the table.** The SP-5d signal has long-short **+31.6% ann (Sharpe 2.51)**, but TopkDropout long-only ≈ flat vs SPY, and OptimalStopControl underperformed (valid-window threshold overfit: valid +7.5% → test −17.7% on one calibration). - **Costs eat most of the gross edge** (+9.1% → +0.85% net). Reduce turnover or go long-short to widen the net edge. - OptimalStopControl thresholds must be calibrated on the *valid* window and are sensitive to overfit — prefer robust defaults or penalize turnover in selection. ## Gotchas - **Installed package copy**: `tac_qlib` in the venv is a copy under `/opt/venv/lib/python3.12/site-packages/tac_qlib/`. After editing any `tac_qlib/contrib/**` module, `cp` it there or the workflow imports the stale version. New subpackages need `mkdir -p` first. - `qlib.backtest` exports `Order` but not `OrderDir`/`Position` at top level — import `Order` from `qlib.backtest`, `OrderDir`/`TradeDecisionWO` from `qlib.backtest.decision`, `Position` from `qlib.backtest.position`. - `qlib.backtest.high_performance_ds` may not export `Order` in this build — don't import from it. - HMM / GARCH / catch22 features must not see test data at fit time: fit the HMM on the train window only (`fit_end=FIT_END`), and compute rolling windows ending at each day. GARCH/entropy use a stride + forward-fill for speed (~5x). - `pycatch22`, `arch`, `hurst`, `antropy`, `hmmlearn` are required for the full feature set; install with `uv pip install --python /app/.venv/bin/python ` (a C compiler is needed for `pycatch22`). `duckdb` and `pyarrow` are declared in `tac-qlib/pyproject.toml`; if a workflow import fails on either, lazy-install with `uv pip install --python /app/.venv/bin/python duckdb pyarrow`. - `rd_run_workflow` defaults to `wait=false`: it returns immediately with `status: started` and the workflow runs in a background thread — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of the experiment (status `RUNNING` until it finishes), then reuse its `run_id`. Pass `wait=true` only for small windows that finish within the MCP call timeout. - After fixing a YAML model/handler change, remember both `/app/tac-qlib/...` and the `/opt/venv` copy stay in sync. ## Files this skill is based on Minimal, runnable examples live next to this skill in `examples/` — they are the canonical reference for every artifact the skill describes: - Workflows (full `record` block → MLflow on disk): - `examples/workflow_minimal.yaml` — the canonical backtest template (req: every traced backtest runs through a workflow like this via `rd_run_workflow`) - `examples/workflow_rankic.yaml` — RankIC-early-stop model wired in - Repo workflows for reference: `tac-qlib/workflows/workflow_lgb_taclake.yaml`, `tune_run1_wider_5d.yaml`, `tune_run2_regularized.yaml`, `tune_run3_label5d_clean_universe.yaml`, `tune_run4_fix_universe_longtrain.yaml`, `tune_run5_longtest.yaml` - Models: `examples/model_rank_gbdt.py` (`RankICLGBModel`: per-day groups + `feval=rankic` + `metric='None'`). Repo: `tac_qlib/contrib/model/rank_gbdt.py` - Strategies: `examples/strategy_optimal_stop.py` (`OptimalStopControl`), `examples/strategy_beta_neutral.py` (doc-only 3L/3S stub — pattern for a custom strategy + Record; not wired into the package) - Handler: `examples/handler.py` (how to subclass `TACHandler`); repo: `tac_qlib/contrib/data/handler.py`; providers: `tac_qlib/data/providers.py` - Feature engineering: `examples/sp_features.py` (OU + Hurst) and `examples/persist_sp_features.py` (persist `sp_*` into the lake features parquet) - Ranking experiments: `examples/run_rank_objectives.py` (mse vs lambdarank vs rank_xendcg ablation on the lake) - Optstop calibration: `examples/run_optstop_compare.py` (valid-window grid + overfit warning) - Traceability tooling: the `rd_trace_*` MCP tools (tac-qlib-rd, `tac_qlib/trace.py`) — see the traceability section above