Files
book-tac/tac-qlib/skills/tradeac-rd-explain/SKILL.md
T

169 lines
9.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: tradeac-rd-explain
description: Guide agents to retrieve, visualise and interpret TradeAC R&D workflow data — from the input qrun YAML to final IC / backtest metrics — via the tac-qlib-rd MCP tools (rd_exp_*) and the built-in R&D dashboard (/dashboard/rd). Use when asked about experiments, mlruns runs, workflow inputs, model hyper-parameters, IC/Rank IC evaluation, or backtest results.
---
# tradeac-rd-explain
Every `qrun` workflow run is recorded into **mlflow** in the unified R&D store under the lake
root — sqlite `mlruns.db` + artifact files under `mlruns/<experiment_id>/<run_uuid>/` in
`$TAC_LAKE_DIR`. This skill tells you how to pull that data out with the
`rd_exp_*` MCP tools (from `tac_qlib.rd_server`), how to read the raw files directly, and how to
visualise/interpret everything — either from the built-in TradeAC UI or from the raw data.
Quick map of the R&D data:
| Step | Where it lives | `rd_exp_*` tool |
|------|----------------|-----------------|
| Input config (rendered YAML) | artifact `config` on the run | `rd_exp_input` |
| Runs / experiments list | `mlruns.db` → `experiments`, `runs`, `tags`, `params`, `metrics` | `rd_exp_list`, `rd_exp_get_experiment`, `rd_exp_get_run` |
| Predictions & labels | artifacts `pred.pkl`, `label.pkl` | `rd_exp_result` |
| IC / Rank IC | artifacts `ic.pkl`, `ric.pkl` + metrics `IC`, `ICIR`, `Rank IC`, `Rank ICIR` | `rd_exp_result` |
| Group returns | artifacts `long_short_r.pkl`, `long_avg_r.pkl` | `rd_exp_result` |
| Backtest / risk | artifacts `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl` + `1day.*` metrics | `rd_exp_result` |
| Model + hyper-params | artifact `params.pkl` (qlib model), config `task.model` | `rd_exp_model` |
| Hypothesis / evaluation notes | sidecar `rd-notes.json` | `rd_exp_get_notes` / `rd_exp_set_notes` |
## MCP-first policy
- **Use the `rd_exp_*` MCP tools to read all of the above** — do not reinvent them with `sqlite3`/pickle/pandas scripts. The tools are the canonical, JSON-safe way to pull experiment data (they fall back to the raw files automatically).
- **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server`, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- The raw-file/sqlite snippets in §3 below are **only** for cases where the MCP surface is unavailable or the user explicitly asks for a direct peek.
- If the venv is missing a runtime dep (`duckdb`, `pyarrow`, `sqlite3`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) rather than working around it.
## 1. Prerequisites
- The tac-qlib-rd MCP server is registered in `opencode.json` (`.venv/bin/python -m tac_qlib.rd_server`).
- The server resolves the unified R&D store from the lake root, so `uri` defaults to `sqlite:///<lake>/mlruns.db` (overridable via `MLRUNS_URI`).
- Runs must exist first: use `rd_run_workflow` (or `rd_train` + records) to create them.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`) in scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## 2. Getting the data
### 2.1 List experiments and their runs
```text
rd_exp_list
# -> [{experiment_id, name, run_count, latest_run: {run_id, status, headline_metrics}}]
rd_exp_get_experiment experiment_id=1
# -> experiment meta + every run: run_id, status, start/end, git, metrics, params, tags, notes, artifacts
```
### 2.2 Input configuration (what went in)
```text
rd_exp_input experiment_id=1 run_id=<run_uuid>
```
Returns the **saved `config` artifact** (the fully-rendered workflow YAML: `qlib_init`,
`task.model.kwargs` hyper-parameters, `task.dataset.kwargs.handler` universe/window/features,
`segments`, `record` list) plus the resolved `universe` and `feature_fields`. If a run has no
`config` artifact (e.g. older `rd_train` runs) the tool falls back to reconstructing from
recorded params/tags and marks `source: "reconstructed"` / `"partial"`.
> Rule of thumb: **the `config` artifact is the most complete input record**; the sqlite
> `params` table alone (only `cmd-sys.argv`) is not enough to reconstruct the input.
### 2.3 Results & evaluation (what came out)
```text
rd_exp_result experiment_id=1 run_id=<run_uuid>
```
Returns: headline `metrics` (IC / ICIR / Rank IC / Rank ICIR, `l2.train`/`l2.valid`, `1day.*`
risk metrics), per-day `ic_series` (`[{date, ic, ric}]`), `pred_stats`, `group_returns`
(`long_short` / `long_avg`), and the `backtest` report (per-day cumulative return vs benchmark)
+ `risk` table.
### 2.4 Model & hyper-parameters
```text
rd_exp_model experiment_id=1 run_id=<run_uuid>
rd_exp_model experiment_id=1 run_id=<run_uuid> tree_id=7
```
Returns `hyperparams` (from config, preferred), `feature_names`, `feature_importances`,
`num_trees`, `best_iteration`, and a **pruned top-layers tree** for LightGBM:
`tree: {nodes: [{id, depth, feature, threshold, gain, leaf_value, node_count, left, right}]}`.
### 2.5 Notes (hypothesis / evaluation)
```text
rd_exp_get_notes experiment_id=1 run_id=<run_uuid>
rd_exp_set_notes experiment_id=1 run_id=<run_uuid> hypothesis="..." evaluation="..."
# persisted to mlruns/<experiment_id>/<run_uuid>/rd-notes.json
```
## 3. Reading the raw files directly
Everything above is a JSON view of these files (all under `$TAC_LAKE_DIR`):
- `mlruns.db` (sqlite) — `experiments`, `runs`, `tags`, `params`, `metrics`, `latest_metrics`.
Quick peek: `sqlite3 $TAC_LAKE_DIR/mlruns.db "SELECT * FROM latest_metrics;"`.
- `mlruns/<experiment_id>/<run_uuid>/artifacts/` — pickle files:
- `config` → the input YAML (dict); carries the resolved `feature_fields`
- `params.pkl` → the trained model (qlib `LGBModel`; `.model` is a `lightgbm.Booster`)
- `pred.pkl`, `label.pkl`, `ic.pkl`, `ric.pkl`, `long_short_r.pkl`, `long_avg_r.pkl`
- `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl`
- `mlruns/<experiment_id>/<run_uuid>/rd-notes.json` — hypothesis/evaluation notes.
In Python:
```python
import os
import pickle
from pathlib import Path
run_dir = Path(os.environ["TAC_LAKE_DIR"]) / "mlruns/1/<run_uuid>"
cfg = pickle.loads((run_dir / "artifacts/config").read_bytes()) # input config dict
ic = pickle.loads((run_dir / "artifacts/ic.pkl").read_bytes()) # per-day IC Series
import lightgbm
model = pickle.loads((run_dir / "artifacts/params.pkl").read_bytes()) # needs qlib import
tree = model.model.dump_model()["tree_info"] # LightGBM trees
```
## 4. Visualising & interpreting
### 4.1 Built-in TradeAC UI
Open the dashboard: `/dashboard/rd` lists experiments + runs with headline metrics and
`Input` / `Result` / `Model` action buttons:
- `/dashboard/rd/input?expId=<id>` — universe, windows, features, model settings (tables).
- `/dashboard/rd/result?expId=<id>` — ECharts IC/Rank IC, cumulative group returns, backtest vs
benchmark, per-day IC table, training-loss curves.
- `/dashboard/rd/model?expId=<id>` — hyper-parameter table, feature importances, LightGBM tree
viewer (pick a tree id).
### 4.2 Interpreting the numbers
- **IC / ICIR**: mean per-day IC (predictive power of the signal); ICIR = mean/std × √252.
|IC| ≥ ~0.02 daily with stable sign is notable for cross-sectional signals; ICIR ≥ 1 is decent,
≥ 2 strong. Rank IC is the Spearman version (more robust to outliers).
- **Training loss (`l2.train`/`l2.valid`)**: watch the gap — widening gap ⇒ overfitting;
valid flat/rising ⇒ underfitting or stale features.
- **Group returns (`long_short_r`)**: cumulative return of top-decile-minus-bottom-decile signal
baskets; steady positive slope = the ranking carries money.
- **Backtest risk** (`annualized_return`, `information_ratio`, `max_drawdown`): IR = excess
return / tracking error; max drawdown shows path risk. Compare against the benchmark column
in the cumulative chart.
- **Tree viewer**: root splits on the strongest features (high gain). Repeated use of a feature
across the top layers ⇒ it dominates; suspicious thresholds near feature extremes often
indicate leakage/sample bias.
## 5. Troubleshooting
| Symptom | Cause / fix |
|---------|-------------|
| `experiment_id` not found | Check `rd_exp_list`; ids are the mlflow `experiment_id`, not the name. |
| `no recorder` / empty input | Run lacks a `config` artifact (pre-fix `rd_train`). Re-run via `rd_run_workflow` or `rd_train` on the fixed server to record config. |
| Pickle errors on `params.pkl` | Ensure qlib + lightgbm importable (server venv). Tool returns a warning and skips the artifact rather than failing. |
| Empty result series | Records were not run (only `rd_train`). Use `rd_run_workflow` or add `SignalRecord`/`SigAnaRecord`/`PortAnaRecord`. |