169 lines
9.2 KiB
Markdown
169 lines
9.2 KiB
Markdown
---
|
||
name: tradeac-rd-explain
|
||
description: Guide agents to retrieve, visualise and interpret TradeAC R&D workflow data — from the input qrun YAML to final IC / backtest metrics — via the tac-qlib-rd MCP tools (rd_exp_*) and the built-in R&D dashboard (/dashboard/rd). Use when asked about experiments, mlruns runs, workflow inputs, model hyper-parameters, IC/Rank IC evaluation, or backtest results.
|
||
---
|
||
|
||
# tradeac-rd-explain
|
||
|
||
Every `qrun` workflow run is recorded into **mlflow** in the unified R&D store under the lake
|
||
root — sqlite `mlruns.db` + artifact files under `mlruns/<experiment_id>/<run_uuid>/` in
|
||
`$TAC_LAKE_DIR`. This skill tells you how to pull that data out with the
|
||
`rd_exp_*` MCP tools (from `tac_qlib.rd_server`), how to read the raw files directly, and how to
|
||
visualise/interpret everything — either from the built-in TradeAC UI or from the raw data.
|
||
|
||
Quick map of the R&D data:
|
||
|
||
| Step | Where it lives | `rd_exp_*` tool |
|
||
|------|----------------|-----------------|
|
||
| Input config (rendered YAML) | artifact `config` on the run | `rd_exp_input` |
|
||
| Runs / experiments list | `mlruns.db` → `experiments`, `runs`, `tags`, `params`, `metrics` | `rd_exp_list`, `rd_exp_get_experiment`, `rd_exp_get_run` |
|
||
| Predictions & labels | artifacts `pred.pkl`, `label.pkl` | `rd_exp_result` |
|
||
| IC / Rank IC | artifacts `ic.pkl`, `ric.pkl` + metrics `IC`, `ICIR`, `Rank IC`, `Rank ICIR` | `rd_exp_result` |
|
||
| Group returns | artifacts `long_short_r.pkl`, `long_avg_r.pkl` | `rd_exp_result` |
|
||
| Backtest / risk | artifacts `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl` + `1day.*` metrics | `rd_exp_result` |
|
||
| Model + hyper-params | artifact `params.pkl` (qlib model), config `task.model` | `rd_exp_model` |
|
||
| Hypothesis / evaluation notes | sidecar `rd-notes.json` | `rd_exp_get_notes` / `rd_exp_set_notes` |
|
||
|
||
## MCP-first policy
|
||
|
||
- **Use the `rd_exp_*` MCP tools to read all of the above** — do not reinvent them with `sqlite3`/pickle/pandas scripts. The tools are the canonical, JSON-safe way to pull experiment data (they fall back to the raw files automatically).
|
||
- **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server`, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
|
||
- The raw-file/sqlite snippets in §3 below are **only** for cases where the MCP surface is unavailable or the user explicitly asks for a direct peek.
|
||
- If the venv is missing a runtime dep (`duckdb`, `pyarrow`, `sqlite3`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) rather than working around it.
|
||
|
||
## 1. Prerequisites
|
||
|
||
- The tac-qlib-rd MCP server is registered in `opencode.json` (`.venv/bin/python -m tac_qlib.rd_server`).
|
||
- The server resolves the unified R&D store from the lake root, so `uri` defaults to `sqlite:///<lake>/mlruns.db` (overridable via `MLRUNS_URI`).
|
||
- Runs must exist first: use `rd_run_workflow` (or `rd_train` + records) to create them.
|
||
|
||
## Secrets policy
|
||
|
||
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`) in scripts, configs, notes or committed code.
|
||
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
|
||
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
|
||
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
|
||
|
||
## 2. Getting the data
|
||
|
||
### 2.1 List experiments and their runs
|
||
|
||
```text
|
||
rd_exp_list
|
||
# -> [{experiment_id, name, run_count, latest_run: {run_id, status, headline_metrics}}]
|
||
|
||
rd_exp_get_experiment experiment_id=1
|
||
# -> experiment meta + every run: run_id, status, start/end, git, metrics, params, tags, notes, artifacts
|
||
```
|
||
|
||
### 2.2 Input configuration (what went in)
|
||
|
||
```text
|
||
rd_exp_input experiment_id=1 run_id=<run_uuid>
|
||
```
|
||
|
||
Returns the **saved `config` artifact** (the fully-rendered workflow YAML: `qlib_init`,
|
||
`task.model.kwargs` hyper-parameters, `task.dataset.kwargs.handler` universe/window/features,
|
||
`segments`, `record` list) plus the resolved `universe` and `feature_fields`. If a run has no
|
||
`config` artifact (e.g. older `rd_train` runs) the tool falls back to reconstructing from
|
||
recorded params/tags and marks `source: "reconstructed"` / `"partial"`.
|
||
|
||
> Rule of thumb: **the `config` artifact is the most complete input record**; the sqlite
|
||
> `params` table alone (only `cmd-sys.argv`) is not enough to reconstruct the input.
|
||
|
||
### 2.3 Results & evaluation (what came out)
|
||
|
||
```text
|
||
rd_exp_result experiment_id=1 run_id=<run_uuid>
|
||
```
|
||
|
||
Returns: headline `metrics` (IC / ICIR / Rank IC / Rank ICIR, `l2.train`/`l2.valid`, `1day.*`
|
||
risk metrics), per-day `ic_series` (`[{date, ic, ric}]`), `pred_stats`, `group_returns`
|
||
(`long_short` / `long_avg`), and the `backtest` report (per-day cumulative return vs benchmark)
|
||
+ `risk` table.
|
||
|
||
### 2.4 Model & hyper-parameters
|
||
|
||
```text
|
||
rd_exp_model experiment_id=1 run_id=<run_uuid>
|
||
rd_exp_model experiment_id=1 run_id=<run_uuid> tree_id=7
|
||
```
|
||
|
||
Returns `hyperparams` (from config, preferred), `feature_names`, `feature_importances`,
|
||
`num_trees`, `best_iteration`, and a **pruned top-layers tree** for LightGBM:
|
||
`tree: {nodes: [{id, depth, feature, threshold, gain, leaf_value, node_count, left, right}]}`.
|
||
|
||
### 2.5 Notes (hypothesis / evaluation)
|
||
|
||
```text
|
||
rd_exp_get_notes experiment_id=1 run_id=<run_uuid>
|
||
rd_exp_set_notes experiment_id=1 run_id=<run_uuid> hypothesis="..." evaluation="..."
|
||
# persisted to mlruns/<experiment_id>/<run_uuid>/rd-notes.json
|
||
```
|
||
|
||
## 3. Reading the raw files directly
|
||
|
||
Everything above is a JSON view of these files (all under `$TAC_LAKE_DIR`):
|
||
|
||
- `mlruns.db` (sqlite) — `experiments`, `runs`, `tags`, `params`, `metrics`, `latest_metrics`.
|
||
Quick peek: `sqlite3 $TAC_LAKE_DIR/mlruns.db "SELECT * FROM latest_metrics;"`.
|
||
- `mlruns/<experiment_id>/<run_uuid>/artifacts/` — pickle files:
|
||
- `config` → the input YAML (dict); carries the resolved `feature_fields`
|
||
- `params.pkl` → the trained model (qlib `LGBModel`; `.model` is a `lightgbm.Booster`)
|
||
- `pred.pkl`, `label.pkl`, `ic.pkl`, `ric.pkl`, `long_short_r.pkl`, `long_avg_r.pkl`
|
||
- `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl`
|
||
- `mlruns/<experiment_id>/<run_uuid>/rd-notes.json` — hypothesis/evaluation notes.
|
||
|
||
In Python:
|
||
|
||
```python
|
||
import os
|
||
import pickle
|
||
from pathlib import Path
|
||
|
||
run_dir = Path(os.environ["TAC_LAKE_DIR"]) / "mlruns/1/<run_uuid>"
|
||
cfg = pickle.loads((run_dir / "artifacts/config").read_bytes()) # input config dict
|
||
ic = pickle.loads((run_dir / "artifacts/ic.pkl").read_bytes()) # per-day IC Series
|
||
import lightgbm
|
||
model = pickle.loads((run_dir / "artifacts/params.pkl").read_bytes()) # needs qlib import
|
||
tree = model.model.dump_model()["tree_info"] # LightGBM trees
|
||
```
|
||
|
||
## 4. Visualising & interpreting
|
||
|
||
### 4.1 Built-in TradeAC UI
|
||
|
||
Open the dashboard: `/dashboard/rd` lists experiments + runs with headline metrics and
|
||
`Input` / `Result` / `Model` action buttons:
|
||
|
||
- `/dashboard/rd/input?expId=<id>` — universe, windows, features, model settings (tables).
|
||
- `/dashboard/rd/result?expId=<id>` — ECharts IC/Rank IC, cumulative group returns, backtest vs
|
||
benchmark, per-day IC table, training-loss curves.
|
||
- `/dashboard/rd/model?expId=<id>` — hyper-parameter table, feature importances, LightGBM tree
|
||
viewer (pick a tree id).
|
||
|
||
### 4.2 Interpreting the numbers
|
||
|
||
- **IC / ICIR**: mean per-day IC (predictive power of the signal); ICIR = mean/std × √252.
|
||
|IC| ≥ ~0.02 daily with stable sign is notable for cross-sectional signals; ICIR ≥ 1 is decent,
|
||
≥ 2 strong. Rank IC is the Spearman version (more robust to outliers).
|
||
- **Training loss (`l2.train`/`l2.valid`)**: watch the gap — widening gap ⇒ overfitting;
|
||
valid flat/rising ⇒ underfitting or stale features.
|
||
- **Group returns (`long_short_r`)**: cumulative return of top-decile-minus-bottom-decile signal
|
||
baskets; steady positive slope = the ranking carries money.
|
||
- **Backtest risk** (`annualized_return`, `information_ratio`, `max_drawdown`): IR = excess
|
||
return / tracking error; max drawdown shows path risk. Compare against the benchmark column
|
||
in the cumulative chart.
|
||
- **Tree viewer**: root splits on the strongest features (high gain). Repeated use of a feature
|
||
across the top layers ⇒ it dominates; suspicious thresholds near feature extremes often
|
||
indicate leakage/sample bias.
|
||
|
||
## 5. Troubleshooting
|
||
|
||
| Symptom | Cause / fix |
|
||
|---------|-------------|
|
||
| `experiment_id` not found | Check `rd_exp_list`; ids are the mlflow `experiment_id`, not the name. |
|
||
| `no recorder` / empty input | Run lacks a `config` artifact (pre-fix `rd_train`). Re-run via `rd_run_workflow` or `rd_train` on the fixed server to record config. |
|
||
| Pickle errors on `params.pkl` | Ensure qlib + lightgbm importable (server venv). Tool returns a warning and skips the artifact rather than failing. |
|
||
| Empty result series | Records were not run (only `rd_train`). Use `rd_run_workflow` or add `SignalRecord`/`SigAnaRecord`/`PortAnaRecord`. |
|