Files
tac-exp-dev/tac-qlib/skills/tradeac-rd-explain/SKILL.md
T

9.2 KiB
Raw Blame History

name, description
name description
tradeac-rd-explain Guide agents to retrieve, visualise and interpret TradeAC R&D workflow data — from the input qrun YAML to final IC / backtest metrics — via the tac-qlib-rd MCP tools (rd_exp_*) and the built-in R&D dashboard (/dashboard/rd). Use when asked about experiments, mlruns runs, workflow inputs, model hyper-parameters, IC/Rank IC evaluation, or backtest results.

tradeac-rd-explain

Every qrun workflow run is recorded into mlflow in the unified R&D store under the lake root — sqlite mlruns.db + artifact files under mlruns/<experiment_id>/<run_uuid>/ in $TAC_LAKE_DIR. This skill tells you how to pull that data out with the rd_exp_* MCP tools (from tac_qlib.rd_server), how to read the raw files directly, and how to visualise/interpret everything — either from the built-in TradeAC UI or from the raw data.

Quick map of the R&D data:

Step Where it lives rd_exp_* tool
Input config (rendered YAML) artifact config on the run rd_exp_input
Runs / experiments list mlruns.db → experiments, runs, tags, params, metrics rd_exp_list, rd_exp_get_experiment, rd_exp_get_run
Predictions & labels artifacts pred.pkl, label.pkl rd_exp_result
IC / Rank IC artifacts ic.pkl, ric.pkl + metrics IC, ICIR, Rank IC, Rank ICIR rd_exp_result
Group returns artifacts long_short_r.pkl, long_avg_r.pkl rd_exp_result
Backtest / risk artifacts portfolio_analysis/report_normal_1d.pkl, port_analysis_1d.pkl + 1day.* metrics rd_exp_result
Model + hyper-params artifact params.pkl (qlib model), config task.model rd_exp_model
Hypothesis / evaluation notes sidecar rd-notes.json rd_exp_get_notes / rd_exp_set_notes

MCP-first policy

  • Use the rd_exp_* MCP tools to read all of the above — do not reinvent them with sqlite3/pickle/pandas scripts. The tools are the canonical, JSON-safe way to pull experiment data (they fall back to the raw files automatically).
  • NEVER script directly against the MCP server (spawning tac_qlib.rd_server, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then stop and ask the user to confirm first.
  • The raw-file/sqlite snippets in §3 below are only for cases where the MCP surface is unavailable or the user explicitly asks for a direct peek.
  • If the venv is missing a runtime dep (duckdb, pyarrow, sqlite3), lazy-install it (uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow) rather than working around it.

1. Prerequisites

  • The tac-qlib-rd MCP server is registered in opencode.json (.venv/bin/python -m tac_qlib.rd_server).
  • The server resolves the unified R&D store from the lake root, so uri defaults to sqlite:///<lake>/mlruns.db (overridable via MLRUNS_URI).
  • Runs must exist first: use rd_run_workflow (or rd_train + records) to create them.

Secrets policy

  • NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (DATABASE_URL, MLRUNS_URI) in scripts, configs, notes or committed code.
  • NEVER read *.env / .env.* directly (cat/tail/grep/sed/head on .env). That pulls secrets into this session and leaks them to any agent sharing it.
  • When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned .env) and reference it by name ($VAR), never by value. If it's missing, report which variable is required instead of reading it yourself.
  • If you find a committed secret, flag it, remove it, and replace it with a placeholder.

2. Getting the data

2.1 List experiments and their runs

rd_exp_list
# -> [{experiment_id, name, run_count, latest_run: {run_id, status, headline_metrics}}]

rd_exp_get_experiment experiment_id=1
# -> experiment meta + every run: run_id, status, start/end, git, metrics, params, tags, notes, artifacts

2.2 Input configuration (what went in)

rd_exp_input experiment_id=1 run_id=<run_uuid>

Returns the saved config artifact (the fully-rendered workflow YAML: qlib_init, task.model.kwargs hyper-parameters, task.dataset.kwargs.handler universe/window/features, segments, record list) plus the resolved universe and feature_fields. If a run has no config artifact (e.g. older rd_train runs) the tool falls back to reconstructing from recorded params/tags and marks source: "reconstructed" / "partial".

Rule of thumb: the config artifact is the most complete input record; the sqlite params table alone (only cmd-sys.argv) is not enough to reconstruct the input.

2.3 Results & evaluation (what came out)

rd_exp_result experiment_id=1 run_id=<run_uuid>

Returns: headline metrics (IC / ICIR / Rank IC / Rank ICIR, l2.train/l2.valid, 1day.* risk metrics), per-day ic_series ([{date, ic, ric}]), pred_stats, group_returns (long_short / long_avg), and the backtest report (per-day cumulative return vs benchmark)

  • risk table.

2.4 Model & hyper-parameters

rd_exp_model experiment_id=1 run_id=<run_uuid>
rd_exp_model experiment_id=1 run_id=<run_uuid> tree_id=7

Returns hyperparams (from config, preferred), feature_names, feature_importances, num_trees, best_iteration, and a pruned top-layers tree for LightGBM: tree: {nodes: [{id, depth, feature, threshold, gain, leaf_value, node_count, left, right}]}.

2.5 Notes (hypothesis / evaluation)

rd_exp_get_notes experiment_id=1 run_id=<run_uuid>
rd_exp_set_notes experiment_id=1 run_id=<run_uuid> hypothesis="..." evaluation="..."
# persisted to mlruns/<experiment_id>/<run_uuid>/rd-notes.json

3. Reading the raw files directly

Everything above is a JSON view of these files (all under $TAC_LAKE_DIR):

  • mlruns.db (sqlite) — experiments, runs, tags, params, metrics, latest_metrics. Quick peek: sqlite3 $TAC_LAKE_DIR/mlruns.db "SELECT * FROM latest_metrics;".
  • mlruns/<experiment_id>/<run_uuid>/artifacts/ — pickle files:
    • config → the input YAML (dict); carries the resolved feature_fields
    • params.pkl → the trained model (qlib LGBModel; .model is a lightgbm.Booster)
    • pred.pkl, label.pkl, ic.pkl, ric.pkl, long_short_r.pkl, long_avg_r.pkl
    • portfolio_analysis/report_normal_1d.pkl, port_analysis_1d.pkl
  • mlruns/<experiment_id>/<run_uuid>/rd-notes.json — hypothesis/evaluation notes.

In Python:

import os
import pickle
from pathlib import Path

run_dir = Path(os.environ["TAC_LAKE_DIR"]) / "mlruns/1/<run_uuid>"
cfg = pickle.loads((run_dir / "artifacts/config").read_bytes())      # input config dict
ic = pickle.loads((run_dir / "artifacts/ic.pkl").read_bytes())       # per-day IC Series
import lightgbm
model = pickle.loads((run_dir / "artifacts/params.pkl").read_bytes())  # needs qlib import
tree = model.model.dump_model()["tree_info"]                         # LightGBM trees

4. Visualising & interpreting

4.1 Built-in TradeAC UI

Open the dashboard: /dashboard/rd lists experiments + runs with headline metrics and Input / Result / Model action buttons:

  • /dashboard/rd/input?expId=<id> — universe, windows, features, model settings (tables).
  • /dashboard/rd/result?expId=<id> — ECharts IC/Rank IC, cumulative group returns, backtest vs benchmark, per-day IC table, training-loss curves.
  • /dashboard/rd/model?expId=<id> — hyper-parameter table, feature importances, LightGBM tree viewer (pick a tree id).

4.2 Interpreting the numbers

  • IC / ICIR: mean per-day IC (predictive power of the signal); ICIR = mean/std × √252. |IC| ≥ ~0.02 daily with stable sign is notable for cross-sectional signals; ICIR ≥ 1 is decent, ≥ 2 strong. Rank IC is the Spearman version (more robust to outliers).
  • Training loss (l2.train/l2.valid): watch the gap — widening gap ⇒ overfitting; valid flat/rising ⇒ underfitting or stale features.
  • Group returns (long_short_r): cumulative return of top-decile-minus-bottom-decile signal baskets; steady positive slope = the ranking carries money.
  • Backtest risk (annualized_return, information_ratio, max_drawdown): IR = excess return / tracking error; max drawdown shows path risk. Compare against the benchmark column in the cumulative chart.
  • Tree viewer: root splits on the strongest features (high gain). Repeated use of a feature across the top layers ⇒ it dominates; suspicious thresholds near feature extremes often indicate leakage/sample bias.

5. Troubleshooting

Symptom Cause / fix
experiment_id not found Check rd_exp_list; ids are the mlflow experiment_id, not the name.
no recorder / empty input Run lacks a config artifact (pre-fix rd_train). Re-run via rd_run_workflow or rd_train on the fixed server to record config.
Pickle errors on params.pkl Ensure qlib + lightgbm importable (server venv). Tool returns a warning and skips the artifact rather than failing.
Empty result series Records were not run (only rd_train). Use rd_run_workflow or add SignalRecord/SigAnaRecord/PortAnaRecord.