--- name: tradeac-rd description: Guide agents to run quant R&D on the TradeAC data lake with a Qlib-based research server exposed over MCP (tac-qlib-rd). Train GBDT (LightGBM/XGBoost) and Linear/QDA ML models on lake bars + TA features, generate cross-sectional alpha predictions, evaluate IC/Rank IC, run TopkDropout backtests with benchmark comparison, and execute one-shot YAML workflows — all through `tac_qlib.rd_server`, an MCP server in the repo venv. --- # tradeac-rd Quant R&D server for the TradeAC data lake. Wraps [Qlib](https://github.com/microsoft/qlib) in a local **MCP server** (`tac_qlib.rd_server` in the repo `.venv`) and uses custom qlib data providers that read directly from the lake (see `tac-engine/skills/tradeac-lake/SKILL.md` for the lake itself, and `tac-qlib/README.md` for the package). Registered in `opencode.json` as `tac-qlib-rd` — the tools below are available directly once opencode is restarted. ## MCP-first policy - **Prefer the tac-qlib-rd MCP tools** (`rd_train`, `rd_predict`, `rd_evaluate`, `rd_backtest`, `rd_strategy_targets`, `rd_run_workflow`, `rd_exp_*`, `rd_status`) over writing scripts that reimplement the R&D loop (custom qlib glue, own train/predict/eval/backtest, hand-rolled mlruns readers, own JSON-RPC clients). - **NEVER script directly against the MCP server** (spawning `python -m tac_qlib.rd_server`, driving it via bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**. - Data prep (bars/features backfill) is done with the tac-engine lake MCP tools — see `tac-engine/skills/tradeac-lake/SKILL.md`. Inspect runs with `rd_exp_*` instead of reading `mlruns.db`/pickles directly. - If the venv is missing a runtime dep (e.g. `duckdb`, `pyarrow`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) instead of switching to another tool. ## The R&D loop | Tool | Purpose | |------|---------| | `rd_train` | Fit a model on lake data + TA features, log to MLflow, return run metadata. | | `rd_predict` | Generate out-of-sample predictions from a trained model (by `model_path` or `run_id`). | | `rd_evaluate` | IC / Rank IC stats of a `pred.pkl` vs `label.pkl`. | | `rd_backtest` | TopkDropout backtest of predictions vs a benchmark, with risk metrics + artifacts. | | `rd_strategy_targets` | Turn a prediction's signal day into a deterministic target buy list (TopkDropout selection + sizing). | | `rd_run_workflow` | One-shot: run an entire YAML workflow (train → predict → sig-ana → backtest) and return metrics + artifacts. | | `rd_exp_list` | List MLflow experiments with run ids on the local sqlite store. | | `rd_exp_get_experiment` | Experiment detail: all runs (meta, metrics, notes, artifact files). | | `rd_exp_get_run` | Single run meta + latest metrics. | | `rd_exp_input` | What went into a run: qlib_init, model kwargs, dataset handler kwargs, segments, features, universe, label. | | `rd_exp_result` | What came out: headline IC/ICIR/Rank IC/Rank ICIR, per-day IC series, group returns, prediction stats, backtest report + risk (benchmark-relative). | | `rd_exp_model` | Trained model: class, hyperparameters, LightGBM tree/feature importance. | | `rd_exp_blotter` | Execution log: account P&L summary, daily equity, current positions, trade table, signal blotter. | | `rd_exp_get_notes` / `rd_exp_set_notes` | Read / write hypothesis + evaluation notes on a run. | | `rd_exp_delete` | **HARD-delete** an MLflow experiment: all runs (metrics/params/tags), the traced `rd_experiments` rows that reference them (FK is `ON DELETE CASCADE`), and on-disk artifacts. Irreversible — confirm with the user first. | | `rd_exp_delete_run` | **HARD-delete** a single MLflow run + its traced `rd_experiments` row + artifacts. Irreversible — confirm with the user first. | | `rd_status` | Lake + qlib readiness: data window, symbols, persisted features, qlib version. | The standard flow is `rd_train` → `rd_predict` → `rd_evaluate` → `rd_backtest`; `rd_run_workflow` replaces all of it with a YAML config. The `rd_exp_*` inspection tools read saved MLflow artifacts, so every page in the R&D app (`/rd/input`, `/rd/result`, `/rd/model`, `/rd/blotter`) is backed by an MCP call (`rd_exp_input`, `rd_exp_result`, `rd_exp_model`, `rd_exp_blotter`) keyed by `experiment_id` + `run_id`. ## Setup / env | Var | Default | Purpose | |-----|---------|---------| | `TAC_LAKE_DIR` | **required** (no default) | lake root (bars + features + metadata). Local dev: absolute path (e.g. `/home/data/lake`). | | `TAC_RD_MARKET` | `US` | market partition for lake reads | | `DATABASE_URL` | – | Postgres tracking store for MLflow (its own `experiments`/`runs`/… tables) when set | | `MLRUNS_URI` | postgres (`$DATABASE_URL`) or `sqlite:////mlruns.db` | MLflow tracking URI override. Artifact files always live under `/mlruns///` | The server lives in the repo `.venv`; the MCP config is already registered. Restart opencode after editing `opencode.json`. ## Secrets policy - NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`, `EMBEDDING_API_KEY`) in workflow YAMLs, scripts, configs, notes or committed code. - NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself. - Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows — `rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else the lake sqlite. Never hardcode a `postgres://user:pass@…` URI. - If you find a committed secret, flag it, remove it, and replace it with a placeholder. ## `rd_train` Args: `universe` (comma-separated), `train_start/valid_end/test_end` (`YYYY-MM-DD`), `experiment_name`, `out_dir`, `model` (`lgb` default, `xgb`, `linear`, `qda`), optional `label` (default `Ref($close,-2)/Ref($close,-1)-1`, the next-day return), `topk`/`n_drop` for later backtests. - Loads 1d bars + all persisted TA features (`features/` dir) for the universe from the lake. - Splits into train / valid / test; fits on train with early stopping on valid. - Logs the run to MLflow (`run_id`), saves `params.pkl` (model) + `pred.pkl` + `label.pkl` to `out_dir`. - With `record_analysis=true` (default) also runs SignalRecord / SigAnaRecord (`ana_long_short`) / PortAnaRecord inside the run, so the result page gets IC/Rank IC series, long-short group returns, monthly IC and the portfolio backtest. `benchmark`, `topk`, `n_drop`, `account`, `risk_degree`, `open_cost`/`close_cost`/`min_cost` tune that backtest. - Returns `run_id`, `status`, `fit_seconds`, `feature_fields`, `label`, per-segment rows/date/instrument counts, and artifact paths. - **`wait=false`** runs the fit in a background thread and returns immediately (`status: started`, `background: true`) — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until its status is `FINISHED`, then use that `run_id`. Use it for slow windows (e.g. a 4y retrain) where a blocking MCP call can time out. > If the lake lacks bars or features for `universe`, backfill first via the tac-engine `get_lake_bars` / `get_lake_ta` tools, or raise the training start date. ## `rd_predict` Args: `universe`, `model_path` **or** `run_id`+`experiment_name` (artifact `params.pkl` is loaded from MLflow), same date ranges as `rd_train`, `out_dir`, optional `top` (number of top-scored rows in the `head` list). - Rebuilds the same feature matrix for `test_start..test_end`, produces scores. - Writes `pred.pkl` (scores) and `label.pkl` (labels) to `out_dir`. - Returns paths, `count`, `date_min/max`, instruments, score distribution stats, and a small `head`. ## `rd_evaluate` Args: `pred_path`, `label_path` (the two pkl files from `rd_train`/`rd_predict`). - Returns `IC` and `RankIC` tables (`days`, `mean`, `std`, `ann_vol`, `ir`, `skew`, `kurt`, `maxdd`) and a `headline` (`IC`, `ICIR`, `Rank IC`, `Rank ICIR`). ## `rd_backtest` Args: `pred_path`, `start_time`/`end_time`, `topk`, `n_drop`, `benchmark`, optional `out_dir`. - TopkDropoutStrategy (topk long, n_drop drop), $100k account, `risk_degree 0.95`, day freq, benchmark comparison. - Returns `start_time`, `end_time`, `trading_days`, `risk` (`mean`, `std`, `annualized_return`, `information_ratio`, `max_drawdown`), `benchmark`, and artifact paths (`report_normal.csv`, `positions_normal.csv`, `risk.csv`). ## `rd_strategy_targets` Args: `pred_path`, optional `signal_date` (defaults to the last day in the prediction), `account`, `risk_degree`, `topk`, `n_drop`, optional `prices` (JSON `{symbol: price}`). - Applies the **exact TopkDropout selection** for one signal day: rank the cross-sectional scores, drop the top `n_drop`, take the next `topk` as buys, sized at `account × risk_degree / topk` per name. Use this to chain a prediction straight into an order list — no manual strategy replication. - With `prices`, floors each order to whole shares (`qty`) and reports `expected_price` / `invested`. - Returns `signal_date`, `per_name_notional`, the deterministic `targets` list (`symbol`, `rank`, `score`, `side`, `notional`, `qty`), and the top-20 `ranking` for context. If fewer than `topk + n_drop` names have a score that day it returns empty `targets` with a `reason`. ## `rd_run_workflow` Args: `config_path` (YAML, see `tac-qlib/workflows/workflow_lgb_taclake.yaml`), `experiment_name`, optional `wait` (default `false`), optional `run_in_new_process` (default `false`). - Runs the full pipeline (qlib `signal` + `records`), returns `run_id`, `status`, the resolved `qlib_init`/`model`/`dataset`/`records` config, and `metrics` (train/valid loss, IC/ICIR/Rank IC/Rank ICIR, and the `1day.*` backtest metrics). - `wait=false` (default) returns immediately with `status: started`; the workflow runs in a background thread — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until it finishes. `wait=true` blocks until completion (only for small windows that finish inside the MCP call timeout). - `run_in_new_process=true` runs the workflow in a **separate OS process** instead of a thread. qlib `init` sets process-global state, so this is the safe mode for concurrent or long workflows — it isolates crashes, releases memory on exit, and avoids the thread-safety race. stdout/stderr are redirected to `/logs/rd-workflow--.log` (returned as `log_path`; the child must never write to the MCP stdio pipe). Polling works identically because the child writes to the same mlflow store. The process `pid` is returned. ## Tracing every run started from a chat (REQUIRED) **Every experiment you start from this chat must be traced FIRST.** The R&D lineage (`/rd/lineage`) and the round book build on the `rd_experiments` table — an experiment created by `rd_run_workflow` / `rd_train` without a `rd_trace_start` is invisible there (no lineage node, no chat link). So before triggering any run, use the `rd_trace_*` MCP tools (tac-qlib-rd): 1. Open the trace BEFORE the run (see `tac-qlib/skills/tac-qlib-custom/SKILL.md`, "Experiment traceability" — the skill that owns the trace flow): ``` rd_trace_start rational="" \ details="" \ experiment_name= \ evolved_from= \ session_id="" # -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...} ``` 2. Run the workflow into that **same** `experiment_name`: ``` rd_run_workflow config_path= experiment_name= ``` 3. On success, **finish the trace** (links the run, copies metrics/evaluation): ``` rd_trace_finish experiment_id= ref_id= \ evaluation="" metrics='{...headline...}' \ mlruns_dir=/mlruns// ``` If you are NOT tracing (quick throwaway exploration), say so explicitly and note the run will not appear in the lineage graph. The default for any run started from a chat is to trace it. ## Building a workflow YAML and triggering a run A workflow YAML is a qrun config: `qlib_init` (lake providers + MLflow exp manager), `task.model`, `task.dataset`, and `task.record`. Copy `tac-qlib/workflows/workflow_lgb_taclake.yaml` as the template. ```yaml # jinja is available: {%- set LAKE = TAC_LAKE_DIR %} (TAC_LAKE_DIR is required) qlib_init: provider_uri: "{{ LAKE }}" region: us expression_cache: null dataset_cache: null calendar_provider: {class: tac_qlib.data.providers.LakeCalendarProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}} instrument_provider: {class: tac_qlib.data.providers.LakeInstrumentProvider, kwargs: {lake_root: "{{ LAKE }}", market: US, markets: {}}} feature_provider: {class: tac_qlib.data.providers.LakeFeatureProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}} exp_manager: {class: MLflowExpManager, module_path: qlib.workflow.expm, kwargs: {uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd"}} task: model: class: LGBModel module_path: qlib.contrib.model.gbdt kwargs: {loss: mse, learning_rate: 0.05, num_leaves: 15, n_estimators: 200, colsample_bytree: 0.8, subsample: 0.8, subsample_freq: 1, reg_alpha: 0.01, reg_lambda: 0.01} dataset: class: DatasetH module_path: qlib.data.dataset kwargs: handler: class: TACHandler module_path: tac_qlib.contrib.data.handler kwargs: instruments: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ # universe start_time: 2000-01-03 # lake look-back for features end_time: 2026-08-06 fit_start_time: 2026-03-01 # normalization fit window fit_end_time: 2026-05-31 freq: day lake_root: "{{ LAKE }}" market: US label: "Ref($close,-2)/Ref($close,-1)-1" # next-day return segments: # train/valid/test split train: [2026-03-01, 2026-05-31] valid: [2026-06-01, 2026-06-30] test: [2026-07-01, 2026-08-06] record: - {class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {}} - {class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: {ana_long_short: true, ann_scaler: 252}} - class: PortAnaRecord module_path: qlib.workflow.record_temp kwargs: config: strategy: class: TopkDropoutStrategy module_path: qlib.contrib.strategy kwargs: {signal: "", topk: 2, n_drop: 1, only_tradable: true, risk_degree: 0.95} backtest: start_time: 2026-07-01 end_time: 2026-08-06 account: 1000000 benchmark: QQQ # any symbol in the lake; empty = no benchmark exchange_kwargs: codes: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ deal_price: $close freq: day open_cost: 0.0005 close_cost: 0.0015 min_cost: 5.0 risk_analysis_freq: 1d ``` Then trigger it (each call = one new run in the named experiment): ```text rd_run_workflow config_path=tac-qlib/workflows/tune_run1_wider_5d.yaml experiment_name=tac-rd-tune # -> run_id ; save it, then inspect via rd_exp_*. ``` The saved `config` artifact (same shape as above) is what `rd_exp_input` returns, so runs are reproducible from their YAML. ## Inspecting a run given experiment_id + run_id The URL on the R&D app is `/rd/input|result|model|blotter?expId=&run=`; the underlying MCP calls are: | You want | MCP call | Args | |----------|----------|------| | Full results (IC/ICIR/Rank IC, group returns, backtest risk) | `rd_exp_result` | `experiment_id`, `run_id` | | Input config (universe, windows, features, label, model) | `rd_exp_input` | `experiment_id`, `run_id` | | Execution blotter (P&L, positions, trades, signals) | `rd_exp_blotter` | `experiment_id`, `run_id` | | Model (hyperparams, tree, importances) | `rd_exp_model` | `experiment_id`, `run_id` | | Run notes | `rd_exp_get_notes` / `rd_exp_set_notes` | `experiment_id`, `run_id` (+ `hypothesis`/`evaluation`) | Get the run ids first: `rd_exp_list` → pick an experiment → `rd_exp_get_experiment` returns its runs (meta + latest metrics), or `rd_exp_get_run run_id=` for one run. ## Evaluating a run and proposing the next one Treat each run as one hypothesis. To evaluate and iterate: 1. **Read the input** (`rd_exp_input`): universe, train/valid/test windows, label expression, features, model + hyperparams. Note what was held fixed vs changed. 2. **Read the signal metrics** (`rd_exp_result.headline`): IC (predictive power), ICIR (stability — |ICIR| ≥ 0.5 strong, 0.2–0.5 weak but persistent, < 0.2 noise), Rank IC/Rank ICIR. A decent IC with Rank IC ≈ 0 means the ranking is noisy even if the mean cross-section is predictive. 3. **Read the backtest** (`rd_exp_result.backtest`): `annualized_return`, `information_ratio`, `max_drawdown` are **excess vs the benchmark** (qlib mean-daily × 238). Compare against `return_annualized` (raw strategy) and `benchmark_annualized`; check the benchmark is a sensible peer (a single high-flying stock like AAPL is a brutal benchmark for an ETF universe). 4. **Read the blotter** (`rd_exp_blotter.summary`): `n_trades`/`trading_days` reveal turnover; `total_cost` vs account is the cost drag; positions show concentration. High turnover + low topk on correlated names = cost-heavy, undiversified book. 5. **Diagnose** and pick ONE lever for the next run — change one thing, hold the rest fixed so the comparison is clean: - *Weak/noisy signal* (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day `Ref($close,-6)/Ref($close,-1)-1`), stronger regularization (`reg_alpha`/`reg_lambda` up, `subsample`/`colsample` down), or a cleaner universe (drop leveraged/duplicate names). - *Good signal, bad book* (high IC but poor excess return): raise `topk` for diversification, tune `n_drop` for rotation, reduce turnover, check `total_cost`. - *Benchmark mismatch*: pick an index ETF (QQQ/IVV) the universe tracks instead of a single stock. - *Data window*: a 3-month fit window is short; consider rolling/expanding if the lake history allows. 6. **Write the next run as a YAML** (see section above), **trace it first** (`rd_trace_start experiment_name=`), then trigger with `rd_run_workflow` into that **same new experiment** (e.g. `tac-rd-tune`), and `rd_trace_finish experiment_id= ref_id=` when it succeeds. Then `rd_exp_get_experiment` to compare run-to-run. Record the hypothesis/evaluation via `rd_exp_set_notes`. Example: the baseline `Exp-1 Run-f29f5446` shows IC 0.071 / ICIR 0.17 / Rank IC 0.014 with excess return −0.94 ann (IR −2.23) vs a +89% ann benchmark — the 1-day signal is unstable, the topk=2 book turned 24 trades in 27 days (~1.1% cost drag) on correlated ETFs + leveraged hedges, and AAPL is an unfair benchmark. Two improvement runs are ready in `tac-qlib/workflows/tune_run1_wider_5d.yaml` (5-day label, topk=5, deduped 10-name universe, benchmark QQQ) and `tune_run2_regularized.yaml` (stronger regularization, topk=3/n_drop=2, same-day label) — trigger both into `experiment_name=tac-rd-tune` and compare. ## Example session ```text # 1) train rd_train universe=AAPL,MSFT,TSLA,USO,SLV,TLT train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06 experiment_name=tac-rd-mcp out_dir=/tmp/rd_out # -> run_id ... # 2) predict on the test window (from the mlflow run) rd_predict universe=AAPL,MSFT,TSLA,USO,SLV,TLT run_id= experiment_name=tac-rd-mcp train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06 out_dir=/tmp/rd_out top=5 # 3) evaluate the alpha rd_evaluate pred_path=/tmp/rd_out/pred.pkl label_path=/tmp/rd_out/label.pkl # 4) backtest the signal rd_backtest pred_path=/tmp/rd_out/pred.pkl start_time=2026-07-01 end_time=2026-08-06 topk=2 n_drop=1 benchmark=AAPL # 5) one-shot equivalent — trace first, then run, then finish rd_trace_start --rational "" --experiment-name tac-rd-one-shot --evolved-from auto rd_run_workflow config_path=tac-qlib/workflows/workflow_lgb_taclake.yaml experiment_name=tac-rd-one-shot rd_trace_finish --id --ref-id --evaluation "" # 6) inspect that run later — given experiment_id + run_id (the /rd pages call exactly these) rd_exp_get_experiment experiment_id=1 # -> runs with meta + latest metrics rd_exp_input experiment_id=1 run_id= # what went in: universe, windows, features, label, model rd_exp_result experiment_id=1 run_id= # what came out: IC/ICIR/Rank IC, backtest risk rd_exp_blotter experiment_id=1 run_id= # execution: P&L, positions, trades, signals rd_exp_model experiment_id=1 run_id= # hyperparameters + tree / importances rd_exp_set_notes experiment_id=1 run_id= hypothesis="5d label + topk5" evaluation="ICIR 0.5, ann +12%" ``` ## Notes - The server reads the lake lazily via the custom `LakeCalendarProvider` / `LakeInstrumentProvider` / `LakeFeatureProvider`; if data is missing the relevant provider raises a clear error — backfill through the tac-engine lake tools first. - All tools return JSON via stdio (MCP). Diagnostics/logs go to stderr. - MLflow runs are stored in the tracking store at `$DATABASE_URL` (Postgres) when set, else the unified lake sqlite `mlruns.db`; artifact files always live under `/mlruns///`. Override the tracking URI with `MLRUNS_URI` if needed. - `rd_run_workflow` / `rd_train` pin each experiment's MLflow `artifact_location` to `/mlruns` so DB and artifacts stay co-located even when the server process runs from another cwd. Readers (`rd_exp_*`) resolve each run's artifact dir from its recorded `artifact_uri`, falling back to the side-by-side `/mlruns` layout — so runs whose artifacts were written elsewhere (e.g. `/mlruns`) still display.