Files

297 lines
23 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: tradeac-rd
description: Guide agents to run quant R&D on the TradeAC data lake with a Qlib-based research server exposed over MCP (tac-qlib-rd). Train GBDT (LightGBM/XGBoost) and Linear/QDA ML models on lake bars + TA features, generate cross-sectional alpha predictions, evaluate IC/Rank IC, run TopkDropout backtests with benchmark comparison, and execute one-shot YAML workflows — all through `tac_qlib.rd_server`, an MCP server in the repo venv.
---
# tradeac-rd
Quant R&D server for the TradeAC data lake. Wraps [Qlib](https://github.com/microsoft/qlib) in a local **MCP server** (`tac_qlib.rd_server` in the repo `.venv`) and uses custom qlib data providers that read directly from the lake (see `tac-engine/skills/tradeac-lake/SKILL.md` for the lake itself, and `tac-qlib/README.md` for the package).
Registered in `opencode.json` as `tac-qlib-rd` — the tools below are available directly once opencode is restarted.
## MCP-first policy
- **Prefer the tac-qlib-rd MCP tools** (`rd_train`, `rd_predict`, `rd_evaluate`, `rd_backtest`, `rd_strategy_targets`, `rd_run_workflow`, `rd_exp_*`, `rd_status`) over writing scripts that reimplement the R&D loop (custom qlib glue, own train/predict/eval/backtest, hand-rolled mlruns readers, own JSON-RPC clients).
- **NEVER script directly against the MCP server** (spawning `python -m tac_qlib.rd_server`, driving it via bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- Data prep (bars/features backfill) is done with the tac-engine lake MCP tools — see `tac-engine/skills/tradeac-lake/SKILL.md`. Inspect runs with `rd_exp_*` instead of reading `mlruns.db`/pickles directly.
- If the venv is missing a runtime dep (e.g. `duckdb`, `pyarrow`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) instead of switching to another tool.
## The R&D loop
| Tool | Purpose |
|------|---------|
| `rd_train` | Fit a model on lake data + TA features, log to MLflow, return run metadata. |
| `rd_predict` | Generate out-of-sample predictions from a trained model (by `model_path` or `run_id`). |
| `rd_evaluate` | IC / Rank IC stats of a `pred.pkl` vs `label.pkl`. |
| `rd_backtest` | TopkDropout backtest of predictions vs a benchmark, with risk metrics + artifacts. |
| `rd_strategy_targets` | Turn a prediction's signal day into a deterministic target buy list (TopkDropout selection + sizing). |
| `rd_run_workflow` | One-shot: run an entire YAML workflow (train → predict → sig-ana → backtest) and return metrics + artifacts. |
| `rd_exp_list` | List MLflow experiments with run ids on the local sqlite store. |
| `rd_exp_get_experiment` | Experiment detail: all runs (meta, metrics, notes, artifact files). |
| `rd_exp_get_run` | Single run meta + latest metrics. |
| `rd_exp_input` | What went into a run: qlib_init, model kwargs, dataset handler kwargs, segments, features, universe, label. |
| `rd_exp_result` | What came out: headline IC/ICIR/Rank IC/Rank ICIR, per-day IC series, group returns, prediction stats, backtest report + risk (benchmark-relative). |
| `rd_exp_model` | Trained model: class, hyperparameters, LightGBM tree/feature importance. |
| `rd_exp_blotter` | Execution log: account P&L summary, daily equity, current positions, trade table, signal blotter. |
| `rd_exp_get_notes` / `rd_exp_set_notes` | Read / write hypothesis + evaluation notes on a run. |
| `rd_exp_delete` | **HARD-delete** an MLflow experiment: all runs (metrics/params/tags), the traced `rd_experiments` rows that reference them (FK is `ON DELETE CASCADE`), and on-disk artifacts. Irreversible — confirm with the user first. |
| `rd_exp_delete_run` | **HARD-delete** a single MLflow run + its traced `rd_experiments` row + artifacts. Irreversible — confirm with the user first. |
| `rd_status` | Lake + qlib readiness: data window, symbols, persisted features, qlib version. |
The standard flow is `rd_train` → `rd_predict` → `rd_evaluate` → `rd_backtest`; `rd_run_workflow` replaces all of it with a YAML config. The `rd_exp_*` inspection tools read saved MLflow artifacts, so every page in the R&D app (`/rd/input`, `/rd/result`, `/rd/model`, `/rd/blotter`) is backed by an MCP call (`rd_exp_input`, `rd_exp_result`, `rd_exp_model`, `rd_exp_blotter`) keyed by `experiment_id` + `run_id`.
## Setup / env
| Var | Default | Purpose |
|-----|---------|---------|
| `TAC_LAKE_DIR` | **required** (no default) | lake root (bars + features + metadata). Local dev: absolute path (e.g. `/home/data/lake`). |
| `TAC_RD_MARKET` | `US` | market partition for lake reads |
| `DATABASE_URL` | – | Postgres tracking store for MLflow (its own `experiments`/`runs`/… tables) when set |
| `MLRUNS_URI` | postgres (`$DATABASE_URL`) or `sqlite:///<lake>/mlruns.db` | MLflow tracking URI override. Artifact files always live under `<lake>/mlruns/<exp_id>/<run_uuid>/` |
The server lives in the repo `.venv`; the MCP config is already registered. Restart opencode after editing `opencode.json`.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`, `EMBEDDING_API_KEY`) in workflow YAMLs, scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows — `rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else the lake sqlite. Never hardcode a `postgres://user:pass@…` URI.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## `rd_train`
Args: `universe` (comma-separated), `train_start/valid_end/test_end` (`YYYY-MM-DD`), `experiment_name`, `out_dir`, `model` (`lgb` default, `xgb`, `linear`, `qda`), optional `label` (default `Ref($close,-2)/Ref($close,-1)-1`, the next-day return), `topk`/`n_drop` for later backtests.
- Loads 1d bars + all persisted TA features (`features/` dir) for the universe from the lake.
- Splits into train / valid / test; fits on train with early stopping on valid.
- Logs the run to MLflow (`run_id`), saves `params.pkl` (model) + `pred.pkl` + `label.pkl` to `out_dir`.
- With `record_analysis=true` (default) also runs SignalRecord / SigAnaRecord (`ana_long_short`) / PortAnaRecord inside the run, so the result page gets IC/Rank IC series, long-short group returns, monthly IC and the portfolio backtest. `benchmark`, `topk`, `n_drop`, `account`, `risk_degree`, `open_cost`/`close_cost`/`min_cost` tune that backtest.
- Returns `run_id`, `status`, `fit_seconds`, `feature_fields`, `label`, per-segment rows/date/instrument counts, and artifact paths.
- **`wait=false`** runs the fit in a background thread and returns immediately (`status: started`, `background: true`) — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until its status is `FINISHED`, then use that `run_id`. Use it for slow windows (e.g. a 4y retrain) where a blocking MCP call can time out.
> If the lake lacks bars or features for `universe`, backfill first via the tac-engine `get_lake_bars` / `get_lake_ta` tools, or raise the training start date.
## `rd_predict`
Args: `universe`, `model_path` **or** `run_id`+`experiment_name` (artifact `params.pkl` is loaded from MLflow), same date ranges as `rd_train`, `out_dir`, optional `top` (number of top-scored rows in the `head` list).
- Rebuilds the same feature matrix for `test_start..test_end`, produces scores.
- Writes `pred.pkl` (scores) and `label.pkl` (labels) to `out_dir`.
- Returns paths, `count`, `date_min/max`, instruments, score distribution stats, and a small `head`.
## `rd_evaluate`
Args: `pred_path`, `label_path` (the two pkl files from `rd_train`/`rd_predict`).
- Returns `IC` and `RankIC` tables (`days`, `mean`, `std`, `ann_vol`, `ir`, `skew`, `kurt`, `maxdd`) and a `headline` (`IC`, `ICIR`, `Rank IC`, `Rank ICIR`).
## `rd_backtest`
Args: `pred_path`, `start_time`/`end_time`, `topk`, `n_drop`, `benchmark`, optional `out_dir`.
- TopkDropoutStrategy (topk long, n_drop drop), $100k account, `risk_degree 0.95`, day freq, benchmark comparison.
- Returns `start_time`, `end_time`, `trading_days`, `risk` (`mean`, `std`, `annualized_return`, `information_ratio`, `max_drawdown`), `benchmark`, and artifact paths (`report_normal.csv`, `positions_normal.csv`, `risk.csv`).
## `rd_strategy_targets`
Args: `pred_path`, optional `signal_date` (defaults to the last day in the prediction), `account`, `risk_degree`, `topk`, `n_drop`, optional `prices` (JSON `{symbol: price}`).
- Applies the **exact TopkDropout selection** for one signal day: rank the cross-sectional scores, drop the top `n_drop`, take the next `topk` as buys, sized at `account × risk_degree / topk` per name. Use this to chain a prediction straight into an order list — no manual strategy replication.
- With `prices`, floors each order to whole shares (`qty`) and reports `expected_price` / `invested`.
- Returns `signal_date`, `per_name_notional`, the deterministic `targets` list (`symbol`, `rank`, `score`, `side`, `notional`, `qty`), and the top-20 `ranking` for context. If fewer than `topk + n_drop` names have a score that day it returns empty `targets` with a `reason`.
## `rd_run_workflow`
Args: `config_path` (YAML, see `tac-qlib/workflows/workflow_lgb_taclake.yaml`), `experiment_name`, optional `wait` (default `false`), optional `run_in_new_process` (default `false`).
- Runs the full pipeline (qlib `signal` + `records`), returns `run_id`, `status`, the resolved `qlib_init`/`model`/`dataset`/`records` config, and `metrics` (train/valid loss, IC/ICIR/Rank IC/Rank ICIR, and the `1day.*` backtest metrics).
- `wait=false` (default) returns immediately with `status: started`; the workflow runs in a background thread — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until it finishes. `wait=true` blocks until completion (only for small windows that finish inside the MCP call timeout).
- `run_in_new_process=true` runs the workflow in a **separate OS process** instead of a thread. qlib `init` sets process-global state, so this is the safe mode for concurrent or long workflows — it isolates crashes, releases memory on exit, and avoids the thread-safety race. stdout/stderr are redirected to `<lake>/logs/rd-workflow-<exp>-<ts>.log` (returned as `log_path`; the child must never write to the MCP stdio pipe). Polling works identically because the child writes to the same mlflow store. The process `pid` is returned.
## Tracing every run started from a chat (REQUIRED)
**Every experiment you start from this chat must be traced FIRST.** The R&D
lineage (`/rd/lineage`) and the round book build on the `rd_experiments` table —
an experiment created by `rd_run_workflow` / `rd_train` without a
`rd_trace_start` is invisible there (no lineage node, no chat link). So before
triggering any run, use the `rd_trace_*` MCP tools (tac-qlib-rd):
1. Open the trace BEFORE the run (see `tac-qlib/skills/tac-qlib-custom/SKILL.md`,
"Experiment traceability" — the skill that owns the trace flow):
```
rd_trace_start rational="<what this run tests, in one line>" \
details="<universe / features / label / model / strategy sizing>" \
experiment_name=<the experiment you will run into> \
evolved_from=<predecessor traced id or auto> \
session_id="<this chat's opencode session id>"
# -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...}
```
2. Run the workflow into that **same** `experiment_name`:
```
rd_run_workflow config_path=<yaml> experiment_name=<the experiment name>
```
3. On success, **finish the trace** (links the run, copies metrics/evaluation):
```
rd_trace_finish experiment_id=<N> ref_id=<run_id> \
evaluation="<outcome>" metrics='{...headline...}' \
mlruns_dir=<lake>/mlruns/<experiment_id>/<run_id>
```
If you are NOT tracing (quick throwaway exploration), say so explicitly and note
the run will not appear in the lineage graph. The default for any run started
from a chat is to trace it.
## Building a workflow YAML and triggering a run
A workflow YAML is a qrun config: `qlib_init` (lake providers + MLflow exp manager), `task.model`, `task.dataset`, and `task.record`. Copy `tac-qlib/workflows/workflow_lgb_taclake.yaml` as the template.
```yaml
# jinja is available: {%- set LAKE = TAC_LAKE_DIR %} (TAC_LAKE_DIR is required)
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
expression_cache: null
dataset_cache: null
calendar_provider: {class: tac_qlib.data.providers.LakeCalendarProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
instrument_provider: {class: tac_qlib.data.providers.LakeInstrumentProvider, kwargs: {lake_root: "{{ LAKE }}", market: US, markets: {}}}
feature_provider: {class: tac_qlib.data.providers.LakeFeatureProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
exp_manager: {class: MLflowExpManager, module_path: qlib.workflow.expm, kwargs: {uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd"}}
task:
model:
class: LGBModel
module_path: qlib.contrib.model.gbdt
kwargs: {loss: mse, learning_rate: 0.05, num_leaves: 15, n_estimators: 200,
colsample_bytree: 0.8, subsample: 0.8, subsample_freq: 1,
reg_alpha: 0.01, reg_lambda: 0.01}
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ # universe
start_time: 2000-01-03 # lake look-back for features
end_time: 2026-08-06
fit_start_time: 2026-03-01 # normalization fit window
fit_end_time: 2026-05-31
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-2)/Ref($close,-1)-1" # next-day return
segments: # train/valid/test split
train: [2026-03-01, 2026-05-31]
valid: [2026-06-01, 2026-06-30]
test: [2026-07-01, 2026-08-06]
record:
- {class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {}}
- {class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: {ana_long_short: true, ann_scaler: 252}}
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config:
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs: {signal: "<PRED>", topk: 2, n_drop: 1, only_tradable: true, risk_degree: 0.95}
backtest:
start_time: 2026-07-01
end_time: 2026-08-06
account: 1000000
benchmark: QQQ # any symbol in the lake; empty = no benchmark
exchange_kwargs:
codes: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ
deal_price: $close
freq: day
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5.0
risk_analysis_freq: 1d
```
Then trigger it (each call = one new run in the named experiment):
```text
rd_run_workflow config_path=tac-qlib/workflows/tune_run1_wider_5d.yaml experiment_name=tac-rd-tune
# -> run_id <uuid>; save it, then inspect via rd_exp_*.
```
The saved `config` artifact (same shape as above) is what `rd_exp_input` returns, so runs are reproducible from their YAML.
## Inspecting a run given experiment_id + run_id
The URL on the R&D app is `/rd/input|result|model|blotter?expId=<id>&run=<uuid>`; the underlying MCP calls are:
| You want | MCP call | Args |
|----------|----------|------|
| Full results (IC/ICIR/Rank IC, group returns, backtest risk) | `rd_exp_result` | `experiment_id`, `run_id` |
| Input config (universe, windows, features, label, model) | `rd_exp_input` | `experiment_id`, `run_id` |
| Execution blotter (P&L, positions, trades, signals) | `rd_exp_blotter` | `experiment_id`, `run_id` |
| Model (hyperparams, tree, importances) | `rd_exp_model` | `experiment_id`, `run_id` |
| Run notes | `rd_exp_get_notes` / `rd_exp_set_notes` | `experiment_id`, `run_id` (+ `hypothesis`/`evaluation`) |
Get the run ids first: `rd_exp_list` → pick an experiment → `rd_exp_get_experiment` returns its runs (meta + latest metrics), or `rd_exp_get_run run_id=<uuid>` for one run.
## Evaluating a run and proposing the next one
Treat each run as one hypothesis. To evaluate and iterate:
1. **Read the input** (`rd_exp_input`): universe, train/valid/test windows, label expression, features, model + hyperparams. Note what was held fixed vs changed.
2. **Read the signal metrics** (`rd_exp_result.headline`): IC (predictive power), ICIR (stability — |ICIR| ≥ 0.5 strong, 0.2–0.5 weak but persistent, < 0.2 noise), Rank IC/Rank ICIR. A decent IC with Rank IC ≈ 0 means the ranking is noisy even if the mean cross-section is predictive.
3. **Read the backtest** (`rd_exp_result.backtest`): `annualized_return`, `information_ratio`, `max_drawdown` are **excess vs the benchmark** (qlib mean-daily × 238). Compare against `return_annualized` (raw strategy) and `benchmark_annualized`; check the benchmark is a sensible peer (a single high-flying stock like AAPL is a brutal benchmark for an ETF universe).
4. **Read the blotter** (`rd_exp_blotter.summary`): `n_trades`/`trading_days` reveal turnover; `total_cost` vs account is the cost drag; positions show concentration. High turnover + low topk on correlated names = cost-heavy, undiversified book.
5. **Diagnose** and pick ONE lever for the next run — change one thing, hold the rest fixed so the comparison is clean:
- *Weak/noisy signal* (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day `Ref($close,-6)/Ref($close,-1)-1`), stronger regularization (`reg_alpha`/`reg_lambda` up, `subsample`/`colsample` down), or a cleaner universe (drop leveraged/duplicate names).
- *Good signal, bad book* (high IC but poor excess return): raise `topk` for diversification, tune `n_drop` for rotation, reduce turnover, check `total_cost`.
- *Benchmark mismatch*: pick an index ETF (QQQ/IVV) the universe tracks instead of a single stock.
- *Data window*: a 3-month fit window is short; consider rolling/expanding if the lake history allows.
6. **Write the next run as a YAML** (see section above), **trace it first** (`rd_trace_start experiment_name=<exp>`), then trigger with `rd_run_workflow` into that **same new experiment** (e.g. `tac-rd-tune`), and `rd_trace_finish experiment_id=<N> ref_id=<run_id>` when it succeeds. Then `rd_exp_get_experiment` to compare run-to-run. Record the hypothesis/evaluation via `rd_exp_set_notes`.
Example: the baseline `Exp-1 Run-f29f5446` shows IC 0.071 / ICIR 0.17 / Rank IC 0.014 with excess return −0.94 ann (IR −2.23) vs a +89% ann benchmark — the 1-day signal is unstable, the topk=2 book turned 24 trades in 27 days (~1.1% cost drag) on correlated ETFs + leveraged hedges, and AAPL is an unfair benchmark. Two improvement runs are ready in `tac-qlib/workflows/tune_run1_wider_5d.yaml` (5-day label, topk=5, deduped 10-name universe, benchmark QQQ) and `tune_run2_regularized.yaml` (stronger regularization, topk=3/n_drop=2, same-day label) — trigger both into `experiment_name=tac-rd-tune` and compare.
## Example session
```text
# 1) train
rd_train universe=AAPL,MSFT,TSLA,USO,SLV,TLT train_start=2026-03-01 train_end=2026-05-31
valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06
experiment_name=tac-rd-mcp out_dir=/tmp/rd_out
# -> run_id ...
# 2) predict on the test window (from the mlflow run)
rd_predict universe=AAPL,MSFT,TSLA,USO,SLV,TLT run_id=<run_id> experiment_name=tac-rd-mcp
train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30
test_start=2026-07-01 test_end=2026-08-06 out_dir=/tmp/rd_out top=5
# 3) evaluate the alpha
rd_evaluate pred_path=/tmp/rd_out/pred.pkl label_path=/tmp/rd_out/label.pkl
# 4) backtest the signal
rd_backtest pred_path=/tmp/rd_out/pred.pkl start_time=2026-07-01 end_time=2026-08-06 topk=2 n_drop=1 benchmark=AAPL
# 5) one-shot equivalent — trace first, then run, then finish
rd_trace_start --rational "<hypothesis>" --experiment-name tac-rd-one-shot --evolved-from auto
rd_run_workflow config_path=tac-qlib/workflows/workflow_lgb_taclake.yaml experiment_name=tac-rd-one-shot
rd_trace_finish --id <EXPERIMENT_ID> --ref-id <run_id> --evaluation "<outcome>"
# 6) inspect that run later — given experiment_id + run_id (the /rd pages call exactly these)
rd_exp_get_experiment experiment_id=1 # -> runs with meta + latest metrics
rd_exp_input experiment_id=1 run_id=<run_id> # what went in: universe, windows, features, label, model
rd_exp_result experiment_id=1 run_id=<run_id> # what came out: IC/ICIR/Rank IC, backtest risk
rd_exp_blotter experiment_id=1 run_id=<run_id> # execution: P&L, positions, trades, signals
rd_exp_model experiment_id=1 run_id=<run_id> # hyperparameters + tree / importances
rd_exp_set_notes experiment_id=1 run_id=<run_id> hypothesis="5d label + topk5" evaluation="ICIR 0.5, ann +12%"
```
## Notes
- The server reads the lake lazily via the custom `LakeCalendarProvider` / `LakeInstrumentProvider` / `LakeFeatureProvider`; if data is missing the relevant provider raises a clear error — backfill through the tac-engine lake tools first.
- All tools return JSON via stdio (MCP). Diagnostics/logs go to stderr.
- MLflow runs are stored in the tracking store at `$DATABASE_URL` (Postgres) when
set, else the unified lake sqlite `mlruns.db`; artifact files always live under
`<lake>/mlruns/<exp_id>/<run_uuid>/`. Override the tracking URI with `MLRUNS_URI` if needed.
- `rd_run_workflow` / `rd_train` pin each experiment's MLflow `artifact_location` to `<lake>/mlruns` so DB and artifacts stay co-located even when the server process runs from another cwd. Readers (`rd_exp_*`) resolve each run's artifact dir from its recorded `artifact_uri`, falling back to the side-by-side `<lake>/mlruns` layout — so runs whose artifacts were written elsewhere (e.g. `<cwd>/mlruns`) still display.