--- name: tradeac-rd-explain description: Guide agents to retrieve, visualise and interpret TradeAC R&D workflow data — from the input qrun YAML to final IC / backtest metrics — via the tac-qlib-rd MCP tools (rd_exp_*) and the built-in R&D dashboard (/dashboard/rd). Use when asked about experiments, mlruns runs, workflow inputs, model hyper-parameters, IC/Rank IC evaluation, or backtest results. --- # tradeac-rd-explain Every `qrun` workflow run is recorded into **mlflow** in the unified R&D store under the lake root — sqlite `mlruns.db` + artifact files under `mlruns///` in `$TAC_LAKE_DIR`. This skill tells you how to pull that data out with the `rd_exp_*` MCP tools (from `tac_qlib.rd_server`), how to read the raw files directly, and how to visualise/interpret everything — either from the built-in TradeAC UI or from the raw data. Quick map of the R&D data: | Step | Where it lives | `rd_exp_*` tool | |------|----------------|-----------------| | Input config (rendered YAML) | artifact `config` on the run | `rd_exp_input` | | Runs / experiments list | `mlruns.db` → `experiments`, `runs`, `tags`, `params`, `metrics` | `rd_exp_list`, `rd_exp_get_experiment`, `rd_exp_get_run` | | Predictions & labels | artifacts `pred.pkl`, `label.pkl` | `rd_exp_result` | | IC / Rank IC | artifacts `ic.pkl`, `ric.pkl` + metrics `IC`, `ICIR`, `Rank IC`, `Rank ICIR` | `rd_exp_result` | | Group returns | artifacts `long_short_r.pkl`, `long_avg_r.pkl` | `rd_exp_result` | | Backtest / risk | artifacts `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl` + `1day.*` metrics | `rd_exp_result` | | Model + hyper-params | artifact `params.pkl` (qlib model), config `task.model` | `rd_exp_model` | | Hypothesis / evaluation notes | sidecar `rd-notes.json` | `rd_exp_get_notes` / `rd_exp_set_notes` | ## MCP-first policy - **Use the `rd_exp_*` MCP tools to read all of the above** — do not reinvent them with `sqlite3`/pickle/pandas scripts. The tools are the canonical, JSON-safe way to pull experiment data (they fall back to the raw files automatically). - **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server`, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**. - The raw-file/sqlite snippets in §3 below are **only** for cases where the MCP surface is unavailable or the user explicitly asks for a direct peek. - If the venv is missing a runtime dep (`duckdb`, `pyarrow`, `sqlite3`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) rather than working around it. ## 1. Prerequisites - The tac-qlib-rd MCP server is registered in `opencode.json` (`.venv/bin/python -m tac_qlib.rd_server`). - The server resolves the unified R&D store from the lake root, so `uri` defaults to `sqlite:////mlruns.db` (overridable via `MLRUNS_URI`). - Runs must exist first: use `rd_run_workflow` (or `rd_train` + records) to create them. ## Secrets policy - NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`) in scripts, configs, notes or committed code. - NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself. - If you find a committed secret, flag it, remove it, and replace it with a placeholder. ## 2. Getting the data ### 2.1 List experiments and their runs ```text rd_exp_list # -> [{experiment_id, name, run_count, latest_run: {run_id, status, headline_metrics}}] rd_exp_get_experiment experiment_id=1 # -> experiment meta + every run: run_id, status, start/end, git, metrics, params, tags, notes, artifacts ``` ### 2.2 Input configuration (what went in) ```text rd_exp_input experiment_id=1 run_id= ``` Returns the **saved `config` artifact** (the fully-rendered workflow YAML: `qlib_init`, `task.model.kwargs` hyper-parameters, `task.dataset.kwargs.handler` universe/window/features, `segments`, `record` list) plus the resolved `universe` and `feature_fields`. If a run has no `config` artifact (e.g. older `rd_train` runs) the tool falls back to reconstructing from recorded params/tags and marks `source: "reconstructed"` / `"partial"`. > Rule of thumb: **the `config` artifact is the most complete input record**; the sqlite > `params` table alone (only `cmd-sys.argv`) is not enough to reconstruct the input. ### 2.3 Results & evaluation (what came out) ```text rd_exp_result experiment_id=1 run_id= ``` Returns: headline `metrics` (IC / ICIR / Rank IC / Rank ICIR, `l2.train`/`l2.valid`, `1day.*` risk metrics), per-day `ic_series` (`[{date, ic, ric}]`), `pred_stats`, `group_returns` (`long_short` / `long_avg`), and the `backtest` report (per-day cumulative return vs benchmark) + `risk` table. ### 2.4 Model & hyper-parameters ```text rd_exp_model experiment_id=1 run_id= rd_exp_model experiment_id=1 run_id= tree_id=7 ``` Returns `hyperparams` (from config, preferred), `feature_names`, `feature_importances`, `num_trees`, `best_iteration`, and a **pruned top-layers tree** for LightGBM: `tree: {nodes: [{id, depth, feature, threshold, gain, leaf_value, node_count, left, right}]}`. ### 2.5 Notes (hypothesis / evaluation) ```text rd_exp_get_notes experiment_id=1 run_id= rd_exp_set_notes experiment_id=1 run_id= hypothesis="..." evaluation="..." # persisted to mlruns///rd-notes.json ``` ## 3. Reading the raw files directly Everything above is a JSON view of these files (all under `$TAC_LAKE_DIR`): - `mlruns.db` (sqlite) — `experiments`, `runs`, `tags`, `params`, `metrics`, `latest_metrics`. Quick peek: `sqlite3 $TAC_LAKE_DIR/mlruns.db "SELECT * FROM latest_metrics;"`. - `mlruns///artifacts/` — pickle files: - `config` → the input YAML (dict); carries the resolved `feature_fields` - `params.pkl` → the trained model (qlib `LGBModel`; `.model` is a `lightgbm.Booster`) - `pred.pkl`, `label.pkl`, `ic.pkl`, `ric.pkl`, `long_short_r.pkl`, `long_avg_r.pkl` - `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl` - `mlruns///rd-notes.json` — hypothesis/evaluation notes. In Python: ```python import os import pickle from pathlib import Path run_dir = Path(os.environ["TAC_LAKE_DIR"]) / "mlruns/1/" cfg = pickle.loads((run_dir / "artifacts/config").read_bytes()) # input config dict ic = pickle.loads((run_dir / "artifacts/ic.pkl").read_bytes()) # per-day IC Series import lightgbm model = pickle.loads((run_dir / "artifacts/params.pkl").read_bytes()) # needs qlib import tree = model.model.dump_model()["tree_info"] # LightGBM trees ``` ## 4. Visualising & interpreting ### 4.1 Built-in TradeAC UI Open the dashboard: `/dashboard/rd` lists experiments + runs with headline metrics and `Input` / `Result` / `Model` action buttons: - `/dashboard/rd/input?expId=` — universe, windows, features, model settings (tables). - `/dashboard/rd/result?expId=` — ECharts IC/Rank IC, cumulative group returns, backtest vs benchmark, per-day IC table, training-loss curves. - `/dashboard/rd/model?expId=` — hyper-parameter table, feature importances, LightGBM tree viewer (pick a tree id). ### 4.2 Interpreting the numbers - **IC / ICIR**: mean per-day IC (predictive power of the signal); ICIR = mean/std × √252. |IC| ≥ ~0.02 daily with stable sign is notable for cross-sectional signals; ICIR ≥ 1 is decent, ≥ 2 strong. Rank IC is the Spearman version (more robust to outliers). - **Training loss (`l2.train`/`l2.valid`)**: watch the gap — widening gap ⇒ overfitting; valid flat/rising ⇒ underfitting or stale features. - **Group returns (`long_short_r`)**: cumulative return of top-decile-minus-bottom-decile signal baskets; steady positive slope = the ranking carries money. - **Backtest risk** (`annualized_return`, `information_ratio`, `max_drawdown`): IR = excess return / tracking error; max drawdown shows path risk. Compare against the benchmark column in the cumulative chart. - **Tree viewer**: root splits on the strongest features (high gain). Repeated use of a feature across the top layers ⇒ it dominates; suspicious thresholds near feature extremes often indicate leakage/sample bias. ## 5. Troubleshooting | Symptom | Cause / fix | |---------|-------------| | `experiment_id` not found | Check `rd_exp_list`; ids are the mlflow `experiment_id`, not the name. | | `no recorder` / empty input | Run lacks a `config` artifact (pre-fix `rd_train`). Re-run via `rd_run_workflow` or `rd_train` on the fixed server to record config. | | Pickle errors on `params.pkl` | Ensure qlib + lightgbm importable (server venv). Tool returns a warning and skips the artifact rather than failing. | | Empty result series | Records were not run (only `rd_train`). Use `rd_run_workflow` or add `SignalRecord`/`SigAnaRecord`/`PortAnaRecord`. |