23 KiB
name, description
| name | description |
|---|---|
| tradeac-rd | Guide agents to run quant R&D on the TradeAC data lake with a Qlib-based research server exposed over MCP (tac-qlib-rd). Train GBDT (LightGBM/XGBoost) and Linear/QDA ML models on lake bars + TA features, generate cross-sectional alpha predictions, evaluate IC/Rank IC, run TopkDropout backtests with benchmark comparison, and execute one-shot YAML workflows — all through `tac_qlib.rd_server`, an MCP server in the repo venv. |
tradeac-rd
Quant R&D server for the TradeAC data lake. Wraps Qlib in a local MCP server (tac_qlib.rd_server in the repo .venv) and uses custom qlib data providers that read directly from the lake (see tac-engine/skills/tradeac-lake/SKILL.md for the lake itself, and tac-qlib/README.md for the package).
Registered in opencode.json as tac-qlib-rd — the tools below are available directly once opencode is restarted.
MCP-first policy
- Prefer the tac-qlib-rd MCP tools (
rd_train,rd_predict,rd_evaluate,rd_backtest,rd_strategy_targets,rd_run_workflow,rd_exp_*,rd_status) over writing scripts that reimplement the R&D loop (custom qlib glue, own train/predict/eval/backtest, hand-rolled mlruns readers, own JSON-RPC clients). - NEVER script directly against the MCP server (spawning
python -m tac_qlib.rd_server, driving it via bash/curl/stdio) unless a tool genuinely can't do the job — then stop and ask the user to confirm first. - Data prep (bars/features backfill) is done with the tac-engine lake MCP tools — see
tac-engine/skills/tradeac-lake/SKILL.md. Inspect runs withrd_exp_*instead of readingmlruns.db/pickles directly. - If the venv is missing a runtime dep (e.g.
duckdb,pyarrow), lazy-install it (uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow) instead of switching to another tool.
The R&D loop
| Tool | Purpose |
|---|---|
rd_train |
Fit a model on lake data + TA features, log to MLflow, return run metadata. |
rd_predict |
Generate out-of-sample predictions from a trained model (by model_path or run_id). |
rd_evaluate |
IC / Rank IC stats of a pred.pkl vs label.pkl. |
rd_backtest |
TopkDropout backtest of predictions vs a benchmark, with risk metrics + artifacts. |
rd_strategy_targets |
Turn a prediction's signal day into a deterministic target buy list (TopkDropout selection + sizing). |
rd_run_workflow |
One-shot: run an entire YAML workflow (train → predict → sig-ana → backtest) and return metrics + artifacts. |
rd_exp_list |
List MLflow experiments with run ids on the local sqlite store. |
rd_exp_get_experiment |
Experiment detail: all runs (meta, metrics, notes, artifact files). |
rd_exp_get_run |
Single run meta + latest metrics. |
rd_exp_input |
What went into a run: qlib_init, model kwargs, dataset handler kwargs, segments, features, universe, label. |
rd_exp_result |
What came out: headline IC/ICIR/Rank IC/Rank ICIR, per-day IC series, group returns, prediction stats, backtest report + risk (benchmark-relative). |
rd_exp_model |
Trained model: class, hyperparameters, LightGBM tree/feature importance. |
rd_exp_blotter |
Execution log: account P&L summary, daily equity, current positions, trade table, signal blotter. |
rd_exp_get_notes / rd_exp_set_notes |
Read / write hypothesis + evaluation notes on a run. |
rd_exp_delete |
HARD-delete an MLflow experiment: all runs (metrics/params/tags), the traced rd_experiments rows that reference them (FK is ON DELETE CASCADE), and on-disk artifacts. Irreversible — confirm with the user first. |
rd_exp_delete_run |
HARD-delete a single MLflow run + its traced rd_experiments row + artifacts. Irreversible — confirm with the user first. |
rd_status |
Lake + qlib readiness: data window, symbols, persisted features, qlib version. |
The standard flow is rd_train → rd_predict → rd_evaluate → rd_backtest; rd_run_workflow replaces all of it with a YAML config. The rd_exp_* inspection tools read saved MLflow artifacts, so every page in the R&D app (/rd/input, /rd/result, /rd/model, /rd/blotter) is backed by an MCP call (rd_exp_input, rd_exp_result, rd_exp_model, rd_exp_blotter) keyed by experiment_id + run_id.
Setup / env
| Var | Default | Purpose |
|---|---|---|
TAC_LAKE_DIR |
required (no default) | lake root (bars + features + metadata). Local dev: absolute path (e.g. /home/data/lake). |
TAC_RD_MARKET |
US |
market partition for lake reads |
DATABASE_URL |
– | Postgres tracking store for MLflow (its own experiments/runs/… tables) when set |
MLRUNS_URI |
postgres ($DATABASE_URL) or sqlite:///<lake>/mlruns.db |
MLflow tracking URI override. Artifact files always live under <lake>/mlruns/<exp_id>/<run_uuid>/ |
The server lives in the repo .venv; the MCP config is already registered. Restart opencode after editing opencode.json.
Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (
DATABASE_URL,MLRUNS_URI,EMBEDDING_API_KEY) in workflow YAMLs, scripts, configs, notes or committed code. - NEVER read
*.env/.env.*directly (cat/tail/grep/sed/headon.env). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned
.env) and reference it by name ($VAR), never by value. If it's missing, report which variable is required instead of reading it yourself. - Tracking store: use
uri: "sqlite:///mlruns.db"(relative) in workflows —rd_run_workflownormalizes it to Postgres when$DATABASE_URLis set, else the lake sqlite. Never hardcode apostgres://user:pass@…URI. - If you find a committed secret, flag it, remove it, and replace it with a placeholder.
rd_train
Args: universe (comma-separated), train_start/valid_end/test_end (YYYY-MM-DD), experiment_name, out_dir, model (lgb default, xgb, linear, qda), optional label (default Ref($close,-2)/Ref($close,-1)-1, the next-day return), topk/n_drop for later backtests.
- Loads 1d bars + all persisted TA features (
features/dir) for the universe from the lake. - Splits into train / valid / test; fits on train with early stopping on valid.
- Logs the run to MLflow (
run_id), savesparams.pkl(model) +pred.pkl+label.pkltoout_dir. - With
record_analysis=true(default) also runs SignalRecord / SigAnaRecord (ana_long_short) / PortAnaRecord inside the run, so the result page gets IC/Rank IC series, long-short group returns, monthly IC and the portfolio backtest.benchmark,topk,n_drop,account,risk_degree,open_cost/close_cost/min_costtune that backtest. - Returns
run_id,status,fit_seconds,feature_fields,label, per-segment rows/date/instrument counts, and artifact paths. wait=falseruns the fit in a background thread and returns immediately (status: started,background: true) — pollrd_exp_get_run/rd_exp_listfor the newest run ofexperiment_nameuntil its status isFINISHED, then use thatrun_id. Use it for slow windows (e.g. a 4y retrain) where a blocking MCP call can time out.
If the lake lacks bars or features for
universe, backfill first via the tac-engineget_lake_bars/get_lake_tatools, or raise the training start date.
rd_predict
Args: universe, model_path or run_id+experiment_name (artifact params.pkl is loaded from MLflow), same date ranges as rd_train, out_dir, optional top (number of top-scored rows in the head list).
- Rebuilds the same feature matrix for
test_start..test_end, produces scores. - Writes
pred.pkl(scores) andlabel.pkl(labels) toout_dir. - Returns paths,
count,date_min/max, instruments, score distribution stats, and a smallhead.
rd_evaluate
Args: pred_path, label_path (the two pkl files from rd_train/rd_predict).
- Returns
ICandRankICtables (days,mean,std,ann_vol,ir,skew,kurt,maxdd) and aheadline(IC,ICIR,Rank IC,Rank ICIR).
rd_backtest
Args: pred_path, start_time/end_time, topk, n_drop, benchmark, optional out_dir.
- TopkDropoutStrategy (topk long, n_drop drop), $100k account,
risk_degree 0.95, day freq, benchmark comparison. - Returns
start_time,end_time,trading_days,risk(mean,std,annualized_return,information_ratio,max_drawdown),benchmark, and artifact paths (report_normal.csv,positions_normal.csv,risk.csv).
rd_strategy_targets
Args: pred_path, optional signal_date (defaults to the last day in the prediction), account, risk_degree, topk, n_drop, optional prices (JSON {symbol: price}).
- Applies the exact TopkDropout selection for one signal day: rank the cross-sectional scores, drop the top
n_drop, take the nexttopkas buys, sized ataccount × risk_degree / topkper name. Use this to chain a prediction straight into an order list — no manual strategy replication. - With
prices, floors each order to whole shares (qty) and reportsexpected_price/invested. - Returns
signal_date,per_name_notional, the deterministictargetslist (symbol,rank,score,side,notional,qty), and the top-20rankingfor context. If fewer thantopk + n_dropnames have a score that day it returns emptytargetswith areason.
rd_run_workflow
Args: config_path (YAML, see tac-qlib/workflows/workflow_lgb_taclake.yaml), experiment_name, optional wait (default false), optional run_in_new_process (default false).
- Runs the full pipeline (qlib
signal+records), returnsrun_id,status, the resolvedqlib_init/model/dataset/recordsconfig, andmetrics(train/valid loss, IC/ICIR/Rank IC/Rank ICIR, and the1day.*backtest metrics). wait=false(default) returns immediately withstatus: started; the workflow runs in a background thread — pollrd_exp_get_run/rd_exp_listfor the newest run ofexperiment_nameuntil it finishes.wait=trueblocks until completion (only for small windows that finish inside the MCP call timeout).run_in_new_process=trueruns the workflow in a separate OS process instead of a thread. qlibinitsets process-global state, so this is the safe mode for concurrent or long workflows — it isolates crashes, releases memory on exit, and avoids the thread-safety race. stdout/stderr are redirected to<lake>/logs/rd-workflow-<exp>-<ts>.log(returned aslog_path; the child must never write to the MCP stdio pipe). Polling works identically because the child writes to the same mlflow store. The processpidis returned.
Tracing every run started from a chat (REQUIRED)
Every experiment you start from this chat must be traced FIRST. The R&D
lineage (/rd/lineage) and the round book build on the rd_experiments table —
an experiment created by rd_run_workflow / rd_train without a
rd_trace_start is invisible there (no lineage node, no chat link). So before
triggering any run, use the rd_trace_* MCP tools (tac-qlib-rd):
- Open the trace BEFORE the run (see
tac-qlib/skills/tac-qlib-custom/SKILL.md, "Experiment traceability" — the skill that owns the trace flow):rd_trace_start rational="<what this run tests, in one line>" \ details="<universe / features / label / model / strategy sizing>" \ experiment_name=<the experiment you will run into> \ evolved_from=<predecessor traced id or auto> \ session_id="<this chat's opencode session id>" # -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...} - Run the workflow into that same
experiment_name:rd_run_workflow config_path=<yaml> experiment_name=<the experiment name> - On success, finish the trace (links the run, copies metrics/evaluation):
rd_trace_finish experiment_id=<N> ref_id=<run_id> \ evaluation="<outcome>" metrics='{...headline...}' \ mlruns_dir=<lake>/mlruns/<experiment_id>/<run_id>
If you are NOT tracing (quick throwaway exploration), say so explicitly and note the run will not appear in the lineage graph. The default for any run started from a chat is to trace it.
Building a workflow YAML and triggering a run
A workflow YAML is a qrun config: qlib_init (lake providers + MLflow exp manager), task.model, task.dataset, and task.record. Copy tac-qlib/workflows/workflow_lgb_taclake.yaml as the template.
# jinja is available: {%- set LAKE = TAC_LAKE_DIR %} (TAC_LAKE_DIR is required)
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
expression_cache: null
dataset_cache: null
calendar_provider: {class: tac_qlib.data.providers.LakeCalendarProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
instrument_provider: {class: tac_qlib.data.providers.LakeInstrumentProvider, kwargs: {lake_root: "{{ LAKE }}", market: US, markets: {}}}
feature_provider: {class: tac_qlib.data.providers.LakeFeatureProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
exp_manager: {class: MLflowExpManager, module_path: qlib.workflow.expm, kwargs: {uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd"}}
task:
model:
class: LGBModel
module_path: qlib.contrib.model.gbdt
kwargs: {loss: mse, learning_rate: 0.05, num_leaves: 15, n_estimators: 200,
colsample_bytree: 0.8, subsample: 0.8, subsample_freq: 1,
reg_alpha: 0.01, reg_lambda: 0.01}
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ # universe
start_time: 2000-01-03 # lake look-back for features
end_time: 2026-08-06
fit_start_time: 2026-03-01 # normalization fit window
fit_end_time: 2026-05-31
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-2)/Ref($close,-1)-1" # next-day return
segments: # train/valid/test split
train: [2026-03-01, 2026-05-31]
valid: [2026-06-01, 2026-06-30]
test: [2026-07-01, 2026-08-06]
record:
- {class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {}}
- {class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: {ana_long_short: true, ann_scaler: 252}}
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config:
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs: {signal: "<PRED>", topk: 2, n_drop: 1, only_tradable: true, risk_degree: 0.95}
backtest:
start_time: 2026-07-01
end_time: 2026-08-06
account: 1000000
benchmark: QQQ # any symbol in the lake; empty = no benchmark
exchange_kwargs:
codes: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ
deal_price: $close
freq: day
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5.0
risk_analysis_freq: 1d
Then trigger it (each call = one new run in the named experiment):
rd_run_workflow config_path=tac-qlib/workflows/tune_run1_wider_5d.yaml experiment_name=tac-rd-tune
# -> run_id <uuid>; save it, then inspect via rd_exp_*.
The saved config artifact (same shape as above) is what rd_exp_input returns, so runs are reproducible from their YAML.
Inspecting a run given experiment_id + run_id
The URL on the R&D app is /rd/input|result|model|blotter?expId=<id>&run=<uuid>; the underlying MCP calls are:
| You want | MCP call | Args |
|---|---|---|
| Full results (IC/ICIR/Rank IC, group returns, backtest risk) | rd_exp_result |
experiment_id, run_id |
| Input config (universe, windows, features, label, model) | rd_exp_input |
experiment_id, run_id |
| Execution blotter (P&L, positions, trades, signals) | rd_exp_blotter |
experiment_id, run_id |
| Model (hyperparams, tree, importances) | rd_exp_model |
experiment_id, run_id |
| Run notes | rd_exp_get_notes / rd_exp_set_notes |
experiment_id, run_id (+ hypothesis/evaluation) |
Get the run ids first: rd_exp_list → pick an experiment → rd_exp_get_experiment returns its runs (meta + latest metrics), or rd_exp_get_run run_id=<uuid> for one run.
Evaluating a run and proposing the next one
Treat each run as one hypothesis. To evaluate and iterate:
- Read the input (
rd_exp_input): universe, train/valid/test windows, label expression, features, model + hyperparams. Note what was held fixed vs changed. - Read the signal metrics (
rd_exp_result.headline): IC (predictive power), ICIR (stability — |ICIR| ≥ 0.5 strong, 0.2–0.5 weak but persistent, < 0.2 noise), Rank IC/Rank ICIR. A decent IC with Rank IC ≈ 0 means the ranking is noisy even if the mean cross-section is predictive. - Read the backtest (
rd_exp_result.backtest):annualized_return,information_ratio,max_drawdownare excess vs the benchmark (qlib mean-daily × 238). Compare againstreturn_annualized(raw strategy) andbenchmark_annualized; check the benchmark is a sensible peer (a single high-flying stock like AAPL is a brutal benchmark for an ETF universe). - Read the blotter (
rd_exp_blotter.summary):n_trades/trading_daysreveal turnover;total_costvs account is the cost drag; positions show concentration. High turnover + low topk on correlated names = cost-heavy, undiversified book. - Diagnose and pick ONE lever for the next run — change one thing, hold the rest fixed so the comparison is clean:
- Weak/noisy signal (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day
Ref($close,-6)/Ref($close,-1)-1), stronger regularization (reg_alpha/reg_lambdaup,subsample/colsampledown), or a cleaner universe (drop leveraged/duplicate names). - Good signal, bad book (high IC but poor excess return): raise
topkfor diversification, tunen_dropfor rotation, reduce turnover, checktotal_cost. - Benchmark mismatch: pick an index ETF (QQQ/IVV) the universe tracks instead of a single stock.
- Data window: a 3-month fit window is short; consider rolling/expanding if the lake history allows.
- Weak/noisy signal (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day
- Write the next run as a YAML (see section above), trace it first (
rd_trace_start experiment_name=<exp>), then trigger withrd_run_workflowinto that same new experiment (e.g.tac-rd-tune), andrd_trace_finish experiment_id=<N> ref_id=<run_id>when it succeeds. Thenrd_exp_get_experimentto compare run-to-run. Record the hypothesis/evaluation viard_exp_set_notes.
Example: the baseline Exp-1 Run-f29f5446 shows IC 0.071 / ICIR 0.17 / Rank IC 0.014 with excess return −0.94 ann (IR −2.23) vs a +89% ann benchmark — the 1-day signal is unstable, the topk=2 book turned 24 trades in 27 days (~1.1% cost drag) on correlated ETFs + leveraged hedges, and AAPL is an unfair benchmark. Two improvement runs are ready in tac-qlib/workflows/tune_run1_wider_5d.yaml (5-day label, topk=5, deduped 10-name universe, benchmark QQQ) and tune_run2_regularized.yaml (stronger regularization, topk=3/n_drop=2, same-day label) — trigger both into experiment_name=tac-rd-tune and compare.
Example session
# 1) train
rd_train universe=AAPL,MSFT,TSLA,USO,SLV,TLT train_start=2026-03-01 train_end=2026-05-31
valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06
experiment_name=tac-rd-mcp out_dir=/tmp/rd_out
# -> run_id ...
# 2) predict on the test window (from the mlflow run)
rd_predict universe=AAPL,MSFT,TSLA,USO,SLV,TLT run_id=<run_id> experiment_name=tac-rd-mcp
train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30
test_start=2026-07-01 test_end=2026-08-06 out_dir=/tmp/rd_out top=5
# 3) evaluate the alpha
rd_evaluate pred_path=/tmp/rd_out/pred.pkl label_path=/tmp/rd_out/label.pkl
# 4) backtest the signal
rd_backtest pred_path=/tmp/rd_out/pred.pkl start_time=2026-07-01 end_time=2026-08-06 topk=2 n_drop=1 benchmark=AAPL
# 5) one-shot equivalent — trace first, then run, then finish
rd_trace_start --rational "<hypothesis>" --experiment-name tac-rd-one-shot --evolved-from auto
rd_run_workflow config_path=tac-qlib/workflows/workflow_lgb_taclake.yaml experiment_name=tac-rd-one-shot
rd_trace_finish --id <EXPERIMENT_ID> --ref-id <run_id> --evaluation "<outcome>"
# 6) inspect that run later — given experiment_id + run_id (the /rd pages call exactly these)
rd_exp_get_experiment experiment_id=1 # -> runs with meta + latest metrics
rd_exp_input experiment_id=1 run_id=<run_id> # what went in: universe, windows, features, label, model
rd_exp_result experiment_id=1 run_id=<run_id> # what came out: IC/ICIR/Rank IC, backtest risk
rd_exp_blotter experiment_id=1 run_id=<run_id> # execution: P&L, positions, trades, signals
rd_exp_model experiment_id=1 run_id=<run_id> # hyperparameters + tree / importances
rd_exp_set_notes experiment_id=1 run_id=<run_id> hypothesis="5d label + topk5" evaluation="ICIR 0.5, ann +12%"
Notes
- The server reads the lake lazily via the custom
LakeCalendarProvider/LakeInstrumentProvider/LakeFeatureProvider; if data is missing the relevant provider raises a clear error — backfill through the tac-engine lake tools first. - All tools return JSON via stdio (MCP). Diagnostics/logs go to stderr.
- MLflow runs are stored in the tracking store at
$DATABASE_URL(Postgres) when set, else the unified lake sqlitemlruns.db; artifact files always live under<lake>/mlruns/<exp_id>/<run_uuid>/. Override the tracking URI withMLRUNS_URIif needed. rd_run_workflow/rd_trainpin each experiment's MLflowartifact_locationto<lake>/mlrunsso DB and artifacts stay co-located even when the server process runs from another cwd. Readers (rd_exp_*) resolve each run's artifact dir from its recordedartifact_uri, falling back to the side-by-side<lake>/mlrunslayout — so runs whose artifacts were written elsewhere (e.g.<cwd>/mlruns) still display.