book: scaffold + ch00 (execution trail as spine) — evidence exp 8-31, round 3

This commit is contained in:
TradeAC Book Agent
2026-08-18 22:35:23 +00:00
commit c93424e76c
83 changed files with 17676 additions and 0 deletions
+168
View File
@@ -0,0 +1,168 @@
---
name: tradeac-rd-explain
description: Guide agents to retrieve, visualise and interpret TradeAC R&D workflow data — from the input qrun YAML to final IC / backtest metrics — via the tac-qlib-rd MCP tools (rd_exp_*) and the built-in R&D dashboard (/dashboard/rd). Use when asked about experiments, mlruns runs, workflow inputs, model hyper-parameters, IC/Rank IC evaluation, or backtest results.
---
# tradeac-rd-explain
Every `qrun` workflow run is recorded into **mlflow** in the unified R&D store under the lake
root — sqlite `mlruns.db` + artifact files under `mlruns/<experiment_id>/<run_uuid>/` in
`$TAC_LAKE_DIR`. This skill tells you how to pull that data out with the
`rd_exp_*` MCP tools (from `tac_qlib.rd_server`), how to read the raw files directly, and how to
visualise/interpret everything — either from the built-in TradeAC UI or from the raw data.
Quick map of the R&D data:
| Step | Where it lives | `rd_exp_*` tool |
|------|----------------|-----------------|
| Input config (rendered YAML) | artifact `config` on the run | `rd_exp_input` |
| Runs / experiments list | `mlruns.db` → `experiments`, `runs`, `tags`, `params`, `metrics` | `rd_exp_list`, `rd_exp_get_experiment`, `rd_exp_get_run` |
| Predictions & labels | artifacts `pred.pkl`, `label.pkl` | `rd_exp_result` |
| IC / Rank IC | artifacts `ic.pkl`, `ric.pkl` + metrics `IC`, `ICIR`, `Rank IC`, `Rank ICIR` | `rd_exp_result` |
| Group returns | artifacts `long_short_r.pkl`, `long_avg_r.pkl` | `rd_exp_result` |
| Backtest / risk | artifacts `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl` + `1day.*` metrics | `rd_exp_result` |
| Model + hyper-params | artifact `params.pkl` (qlib model), config `task.model` | `rd_exp_model` |
| Hypothesis / evaluation notes | sidecar `rd-notes.json` | `rd_exp_get_notes` / `rd_exp_set_notes` |
## MCP-first policy
- **Use the `rd_exp_*` MCP tools to read all of the above** — do not reinvent them with `sqlite3`/pickle/pandas scripts. The tools are the canonical, JSON-safe way to pull experiment data (they fall back to the raw files automatically).
- **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server`, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- The raw-file/sqlite snippets in §3 below are **only** for cases where the MCP surface is unavailable or the user explicitly asks for a direct peek.
- If the venv is missing a runtime dep (`duckdb`, `pyarrow`, `sqlite3`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) rather than working around it.
## 1. Prerequisites
- The tac-qlib-rd MCP server is registered in `opencode.json` (`.venv/bin/python -m tac_qlib.rd_server`).
- The server resolves the unified R&D store from the lake root, so `uri` defaults to `sqlite:///<lake>/mlruns.db` (overridable via `MLRUNS_URI`).
- Runs must exist first: use `rd_run_workflow` (or `rd_train` + records) to create them.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`) in scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## 2. Getting the data
### 2.1 List experiments and their runs
```text
rd_exp_list
# -> [{experiment_id, name, run_count, latest_run: {run_id, status, headline_metrics}}]
rd_exp_get_experiment experiment_id=1
# -> experiment meta + every run: run_id, status, start/end, git, metrics, params, tags, notes, artifacts
```
### 2.2 Input configuration (what went in)
```text
rd_exp_input experiment_id=1 run_id=<run_uuid>
```
Returns the **saved `config` artifact** (the fully-rendered workflow YAML: `qlib_init`,
`task.model.kwargs` hyper-parameters, `task.dataset.kwargs.handler` universe/window/features,
`segments`, `record` list) plus the resolved `universe` and `feature_fields`. If a run has no
`config` artifact (e.g. older `rd_train` runs) the tool falls back to reconstructing from
recorded params/tags and marks `source: "reconstructed"` / `"partial"`.
> Rule of thumb: **the `config` artifact is the most complete input record**; the sqlite
> `params` table alone (only `cmd-sys.argv`) is not enough to reconstruct the input.
### 2.3 Results & evaluation (what came out)
```text
rd_exp_result experiment_id=1 run_id=<run_uuid>
```
Returns: headline `metrics` (IC / ICIR / Rank IC / Rank ICIR, `l2.train`/`l2.valid`, `1day.*`
risk metrics), per-day `ic_series` (`[{date, ic, ric}]`), `pred_stats`, `group_returns`
(`long_short` / `long_avg`), and the `backtest` report (per-day cumulative return vs benchmark)
+ `risk` table.
### 2.4 Model & hyper-parameters
```text
rd_exp_model experiment_id=1 run_id=<run_uuid>
rd_exp_model experiment_id=1 run_id=<run_uuid> tree_id=7
```
Returns `hyperparams` (from config, preferred), `feature_names`, `feature_importances`,
`num_trees`, `best_iteration`, and a **pruned top-layers tree** for LightGBM:
`tree: {nodes: [{id, depth, feature, threshold, gain, leaf_value, node_count, left, right}]}`.
### 2.5 Notes (hypothesis / evaluation)
```text
rd_exp_get_notes experiment_id=1 run_id=<run_uuid>
rd_exp_set_notes experiment_id=1 run_id=<run_uuid> hypothesis="..." evaluation="..."
# persisted to mlruns/<experiment_id>/<run_uuid>/rd-notes.json
```
## 3. Reading the raw files directly
Everything above is a JSON view of these files (all under `$TAC_LAKE_DIR`):
- `mlruns.db` (sqlite) — `experiments`, `runs`, `tags`, `params`, `metrics`, `latest_metrics`.
Quick peek: `sqlite3 $TAC_LAKE_DIR/mlruns.db "SELECT * FROM latest_metrics;"`.
- `mlruns/<experiment_id>/<run_uuid>/artifacts/` — pickle files:
- `config` → the input YAML (dict); carries the resolved `feature_fields`
- `params.pkl` → the trained model (qlib `LGBModel`; `.model` is a `lightgbm.Booster`)
- `pred.pkl`, `label.pkl`, `ic.pkl`, `ric.pkl`, `long_short_r.pkl`, `long_avg_r.pkl`
- `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl`
- `mlruns/<experiment_id>/<run_uuid>/rd-notes.json` — hypothesis/evaluation notes.
In Python:
```python
import os
import pickle
from pathlib import Path
run_dir = Path(os.environ["TAC_LAKE_DIR"]) / "mlruns/1/<run_uuid>"
cfg = pickle.loads((run_dir / "artifacts/config").read_bytes()) # input config dict
ic = pickle.loads((run_dir / "artifacts/ic.pkl").read_bytes()) # per-day IC Series
import lightgbm
model = pickle.loads((run_dir / "artifacts/params.pkl").read_bytes()) # needs qlib import
tree = model.model.dump_model()["tree_info"] # LightGBM trees
```
## 4. Visualising & interpreting
### 4.1 Built-in TradeAC UI
Open the dashboard: `/dashboard/rd` lists experiments + runs with headline metrics and
`Input` / `Result` / `Model` action buttons:
- `/dashboard/rd/input?expId=<id>` — universe, windows, features, model settings (tables).
- `/dashboard/rd/result?expId=<id>` — ECharts IC/Rank IC, cumulative group returns, backtest vs
benchmark, per-day IC table, training-loss curves.
- `/dashboard/rd/model?expId=<id>` — hyper-parameter table, feature importances, LightGBM tree
viewer (pick a tree id).
### 4.2 Interpreting the numbers
- **IC / ICIR**: mean per-day IC (predictive power of the signal); ICIR = mean/std × √252.
|IC| ≥ ~0.02 daily with stable sign is notable for cross-sectional signals; ICIR ≥ 1 is decent,
≥ 2 strong. Rank IC is the Spearman version (more robust to outliers).
- **Training loss (`l2.train`/`l2.valid`)**: watch the gap — widening gap ⇒ overfitting;
valid flat/rising ⇒ underfitting or stale features.
- **Group returns (`long_short_r`)**: cumulative return of top-decile-minus-bottom-decile signal
baskets; steady positive slope = the ranking carries money.
- **Backtest risk** (`annualized_return`, `information_ratio`, `max_drawdown`): IR = excess
return / tracking error; max drawdown shows path risk. Compare against the benchmark column
in the cumulative chart.
- **Tree viewer**: root splits on the strongest features (high gain). Repeated use of a feature
across the top layers ⇒ it dominates; suspicious thresholds near feature extremes often
indicate leakage/sample bias.
## 5. Troubleshooting
| Symptom | Cause / fix |
|---------|-------------|
| `experiment_id` not found | Check `rd_exp_list`; ids are the mlflow `experiment_id`, not the name. |
| `no recorder` / empty input | Run lacks a `config` artifact (pre-fix `rd_train`). Re-run via `rd_run_workflow` or `rd_train` on the fixed server to record config. |
| Pickle errors on `params.pkl` | Ensure qlib + lightgbm importable (server venv). Tool returns a warning and skips the artifact rather than failing. |
| Empty result series | Records were not run (only `rd_train`). Use `rd_run_workflow` or add `SignalRecord`/`SigAnaRecord`/`PortAnaRecord`. |