Files

23 KiB
Raw Permalink Blame History

name, description
name description
tradeac-rd Guide agents to run quant R&D on the TradeAC data lake with a Qlib-based research server exposed over MCP (tac-qlib-rd). Train GBDT (LightGBM/XGBoost) and Linear/QDA ML models on lake bars + TA features, generate cross-sectional alpha predictions, evaluate IC/Rank IC, run TopkDropout backtests with benchmark comparison, and execute one-shot YAML workflows — all through `tac_qlib.rd_server`, an MCP server in the repo venv.

tradeac-rd

Quant R&D server for the TradeAC data lake. Wraps Qlib in a local MCP server (tac_qlib.rd_server in the repo .venv) and uses custom qlib data providers that read directly from the lake (see tac-engine/skills/tradeac-lake/SKILL.md for the lake itself, and tac-qlib/README.md for the package).

Registered in opencode.json as tac-qlib-rd — the tools below are available directly once opencode is restarted.

MCP-first policy

  • Prefer the tac-qlib-rd MCP tools (rd_train, rd_predict, rd_evaluate, rd_backtest, rd_strategy_targets, rd_run_workflow, rd_exp_*, rd_status) over writing scripts that reimplement the R&D loop (custom qlib glue, own train/predict/eval/backtest, hand-rolled mlruns readers, own JSON-RPC clients).
  • NEVER script directly against the MCP server (spawning python -m tac_qlib.rd_server, driving it via bash/curl/stdio) unless a tool genuinely can't do the job — then stop and ask the user to confirm first.
  • Data prep (bars/features backfill) is done with the tac-engine lake MCP tools — see tac-engine/skills/tradeac-lake/SKILL.md. Inspect runs with rd_exp_* instead of reading mlruns.db/pickles directly.
  • If the venv is missing a runtime dep (e.g. duckdb, pyarrow), lazy-install it (uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow) instead of switching to another tool.

The R&D loop

Tool Purpose
rd_train Fit a model on lake data + TA features, log to MLflow, return run metadata.
rd_predict Generate out-of-sample predictions from a trained model (by model_path or run_id).
rd_evaluate IC / Rank IC stats of a pred.pkl vs label.pkl.
rd_backtest TopkDropout backtest of predictions vs a benchmark, with risk metrics + artifacts.
rd_strategy_targets Turn a prediction's signal day into a deterministic target buy list (TopkDropout selection + sizing).
rd_run_workflow One-shot: run an entire YAML workflow (train → predict → sig-ana → backtest) and return metrics + artifacts.
rd_exp_list List MLflow experiments with run ids on the local sqlite store.
rd_exp_get_experiment Experiment detail: all runs (meta, metrics, notes, artifact files).
rd_exp_get_run Single run meta + latest metrics.
rd_exp_input What went into a run: qlib_init, model kwargs, dataset handler kwargs, segments, features, universe, label.
rd_exp_result What came out: headline IC/ICIR/Rank IC/Rank ICIR, per-day IC series, group returns, prediction stats, backtest report + risk (benchmark-relative).
rd_exp_model Trained model: class, hyperparameters, LightGBM tree/feature importance.
rd_exp_blotter Execution log: account P&L summary, daily equity, current positions, trade table, signal blotter.
rd_exp_get_notes / rd_exp_set_notes Read / write hypothesis + evaluation notes on a run.
rd_exp_delete HARD-delete an MLflow experiment: all runs (metrics/params/tags), the traced rd_experiments rows that reference them (FK is ON DELETE CASCADE), and on-disk artifacts. Irreversible — confirm with the user first.
rd_exp_delete_run HARD-delete a single MLflow run + its traced rd_experiments row + artifacts. Irreversible — confirm with the user first.
rd_status Lake + qlib readiness: data window, symbols, persisted features, qlib version.

The standard flow is rd_train → rd_predict → rd_evaluate → rd_backtest; rd_run_workflow replaces all of it with a YAML config. The rd_exp_* inspection tools read saved MLflow artifacts, so every page in the R&D app (/rd/input, /rd/result, /rd/model, /rd/blotter) is backed by an MCP call (rd_exp_input, rd_exp_result, rd_exp_model, rd_exp_blotter) keyed by experiment_id + run_id.

Setup / env

Var Default Purpose
TAC_LAKE_DIR required (no default) lake root (bars + features + metadata). Local dev: absolute path (e.g. /home/data/lake).
TAC_RD_MARKET US market partition for lake reads
DATABASE_URL – Postgres tracking store for MLflow (its own experiments/runs/… tables) when set
MLRUNS_URI postgres ($DATABASE_URL) or sqlite:///<lake>/mlruns.db MLflow tracking URI override. Artifact files always live under <lake>/mlruns/<exp_id>/<run_uuid>/

The server lives in the repo .venv; the MCP config is already registered. Restart opencode after editing opencode.json.

Secrets policy

  • NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (DATABASE_URL, MLRUNS_URI, EMBEDDING_API_KEY) in workflow YAMLs, scripts, configs, notes or committed code.
  • NEVER read *.env / .env.* directly (cat/tail/grep/sed/head on .env). That pulls secrets into this session and leaks them to any agent sharing it.
  • When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned .env) and reference it by name ($VAR), never by value. If it's missing, report which variable is required instead of reading it yourself.
  • Tracking store: use uri: "sqlite:///mlruns.db" (relative) in workflows — rd_run_workflow normalizes it to Postgres when $DATABASE_URL is set, else the lake sqlite. Never hardcode a postgres://user:pass@… URI.
  • If you find a committed secret, flag it, remove it, and replace it with a placeholder.

rd_train

Args: universe (comma-separated), train_start/valid_end/test_end (YYYY-MM-DD), experiment_name, out_dir, model (lgb default, xgb, linear, qda), optional label (default Ref($close,-2)/Ref($close,-1)-1, the next-day return), topk/n_drop for later backtests.

  • Loads 1d bars + all persisted TA features (features/ dir) for the universe from the lake.
  • Splits into train / valid / test; fits on train with early stopping on valid.
  • Logs the run to MLflow (run_id), saves params.pkl (model) + pred.pkl + label.pkl to out_dir.
  • With record_analysis=true (default) also runs SignalRecord / SigAnaRecord (ana_long_short) / PortAnaRecord inside the run, so the result page gets IC/Rank IC series, long-short group returns, monthly IC and the portfolio backtest. benchmark, topk, n_drop, account, risk_degree, open_cost/close_cost/min_cost tune that backtest.
  • Returns run_id, status, fit_seconds, feature_fields, label, per-segment rows/date/instrument counts, and artifact paths.
  • wait=false runs the fit in a background thread and returns immediately (status: started, background: true) — poll rd_exp_get_run / rd_exp_list for the newest run of experiment_name until its status is FINISHED, then use that run_id. Use it for slow windows (e.g. a 4y retrain) where a blocking MCP call can time out.

If the lake lacks bars or features for universe, backfill first via the tac-engine get_lake_bars / get_lake_ta tools, or raise the training start date.

rd_predict

Args: universe, model_path or run_id+experiment_name (artifact params.pkl is loaded from MLflow), same date ranges as rd_train, out_dir, optional top (number of top-scored rows in the head list).

  • Rebuilds the same feature matrix for test_start..test_end, produces scores.
  • Writes pred.pkl (scores) and label.pkl (labels) to out_dir.
  • Returns paths, count, date_min/max, instruments, score distribution stats, and a small head.

rd_evaluate

Args: pred_path, label_path (the two pkl files from rd_train/rd_predict).

  • Returns IC and RankIC tables (days, mean, std, ann_vol, ir, skew, kurt, maxdd) and a headline (IC, ICIR, Rank IC, Rank ICIR).

rd_backtest

Args: pred_path, start_time/end_time, topk, n_drop, benchmark, optional out_dir.

  • TopkDropoutStrategy (topk long, n_drop drop), $100k account, risk_degree 0.95, day freq, benchmark comparison.
  • Returns start_time, end_time, trading_days, risk (mean, std, annualized_return, information_ratio, max_drawdown), benchmark, and artifact paths (report_normal.csv, positions_normal.csv, risk.csv).

rd_strategy_targets

Args: pred_path, optional signal_date (defaults to the last day in the prediction), account, risk_degree, topk, n_drop, optional prices (JSON {symbol: price}).

  • Applies the exact TopkDropout selection for one signal day: rank the cross-sectional scores, drop the top n_drop, take the next topk as buys, sized at account × risk_degree / topk per name. Use this to chain a prediction straight into an order list — no manual strategy replication.
  • With prices, floors each order to whole shares (qty) and reports expected_price / invested.
  • Returns signal_date, per_name_notional, the deterministic targets list (symbol, rank, score, side, notional, qty), and the top-20 ranking for context. If fewer than topk + n_drop names have a score that day it returns empty targets with a reason.

rd_run_workflow

Args: config_path (YAML, see tac-qlib/workflows/workflow_lgb_taclake.yaml), experiment_name, optional wait (default false), optional run_in_new_process (default false).

  • Runs the full pipeline (qlib signal + records), returns run_id, status, the resolved qlib_init/model/dataset/records config, and metrics (train/valid loss, IC/ICIR/Rank IC/Rank ICIR, and the 1day.* backtest metrics).
  • wait=false (default) returns immediately with status: started; the workflow runs in a background thread — poll rd_exp_get_run / rd_exp_list for the newest run of experiment_name until it finishes. wait=true blocks until completion (only for small windows that finish inside the MCP call timeout).
  • run_in_new_process=true runs the workflow in a separate OS process instead of a thread. qlib init sets process-global state, so this is the safe mode for concurrent or long workflows — it isolates crashes, releases memory on exit, and avoids the thread-safety race. stdout/stderr are redirected to <lake>/logs/rd-workflow-<exp>-<ts>.log (returned as log_path; the child must never write to the MCP stdio pipe). Polling works identically because the child writes to the same mlflow store. The process pid is returned.

Tracing every run started from a chat (REQUIRED)

Every experiment you start from this chat must be traced FIRST. The R&D lineage (/rd/lineage) and the round book build on the rd_experiments table — an experiment created by rd_run_workflow / rd_train without a rd_trace_start is invisible there (no lineage node, no chat link). So before triggering any run, use the rd_trace_* MCP tools (tac-qlib-rd):

  1. Open the trace BEFORE the run (see tac-qlib/skills/tac-qlib-custom/SKILL.md, "Experiment traceability" — the skill that owns the trace flow):
    rd_trace_start  rational="<what this run tests, in one line>" \
                    details="<universe / features / label / model / strategy sizing>" \
                    experiment_name=<the experiment you will run into> \
                    evolved_from=<predecessor traced id or auto> \
                    session_id="<this chat's opencode session id>"
    # -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...}
    
  2. Run the workflow into that same experiment_name:
    rd_run_workflow config_path=<yaml> experiment_name=<the experiment name>
    
  3. On success, finish the trace (links the run, copies metrics/evaluation):
    rd_trace_finish  experiment_id=<N>  ref_id=<run_id> \
                     evaluation="<outcome>" metrics='{...headline...}' \
                     mlruns_dir=<lake>/mlruns/<experiment_id>/<run_id>
    

If you are NOT tracing (quick throwaway exploration), say so explicitly and note the run will not appear in the lineage graph. The default for any run started from a chat is to trace it.

Building a workflow YAML and triggering a run

A workflow YAML is a qrun config: qlib_init (lake providers + MLflow exp manager), task.model, task.dataset, and task.record. Copy tac-qlib/workflows/workflow_lgb_taclake.yaml as the template.

# jinja is available: {%- set LAKE = TAC_LAKE_DIR %} (TAC_LAKE_DIR is required)
qlib_init:
    provider_uri: "{{ LAKE }}"
    region: us
    expression_cache: null
    dataset_cache: null
    calendar_provider: {class: tac_qlib.data.providers.LakeCalendarProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
    instrument_provider: {class: tac_qlib.data.providers.LakeInstrumentProvider, kwargs: {lake_root: "{{ LAKE }}", market: US, markets: {}}}
    feature_provider: {class: tac_qlib.data.providers.LakeFeatureProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
    exp_manager: {class: MLflowExpManager, module_path: qlib.workflow.expm, kwargs: {uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd"}}

task:
    model:
        class: LGBModel
        module_path: qlib.contrib.model.gbdt
        kwargs: {loss: mse, learning_rate: 0.05, num_leaves: 15, n_estimators: 200,
                 colsample_bytree: 0.8, subsample: 0.8, subsample_freq: 1,
                 reg_alpha: 0.01, reg_lambda: 0.01}
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: TACHandler
                module_path: tac_qlib.contrib.data.handler
                kwargs:
                    instruments: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ   # universe
                    start_time: 2000-01-03      # lake look-back for features
                    end_time: 2026-08-06
                    fit_start_time: 2026-03-01  # normalization fit window
                    fit_end_time: 2026-05-31
                    freq: day
                    lake_root: "{{ LAKE }}"
                    market: US
                    label: "Ref($close,-2)/Ref($close,-1)-1"   # next-day return
            segments:                                  # train/valid/test split
                train: [2026-03-01, 2026-05-31]
                valid: [2026-06-01, 2026-06-30]
                test:  [2026-07-01, 2026-08-06]
    record:
        - {class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {}}
        - {class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: {ana_long_short: true, ann_scaler: 252}}
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs:
              config:
                  strategy:
                      class: TopkDropoutStrategy
                      module_path: qlib.contrib.strategy
                      kwargs: {signal: "<PRED>", topk: 2, n_drop: 1, only_tradable: true, risk_degree: 0.95}
                  backtest:
                      start_time: 2026-07-01
                      end_time: 2026-08-06
                      account: 1000000
                      benchmark: QQQ                     # any symbol in the lake; empty = no benchmark
                      exchange_kwargs:
                          codes: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ
                          deal_price: $close
                          freq: day
                          open_cost: 0.0005
                          close_cost: 0.0015
                          min_cost: 5.0
              risk_analysis_freq: 1d

Then trigger it (each call = one new run in the named experiment):

rd_run_workflow config_path=tac-qlib/workflows/tune_run1_wider_5d.yaml experiment_name=tac-rd-tune
# -> run_id <uuid>; save it, then inspect via rd_exp_*.

The saved config artifact (same shape as above) is what rd_exp_input returns, so runs are reproducible from their YAML.

Inspecting a run given experiment_id + run_id

The URL on the R&D app is /rd/input|result|model|blotter?expId=<id>&run=<uuid>; the underlying MCP calls are:

You want MCP call Args
Full results (IC/ICIR/Rank IC, group returns, backtest risk) rd_exp_result experiment_id, run_id
Input config (universe, windows, features, label, model) rd_exp_input experiment_id, run_id
Execution blotter (P&L, positions, trades, signals) rd_exp_blotter experiment_id, run_id
Model (hyperparams, tree, importances) rd_exp_model experiment_id, run_id
Run notes rd_exp_get_notes / rd_exp_set_notes experiment_id, run_id (+ hypothesis/evaluation)

Get the run ids first: rd_exp_list → pick an experiment → rd_exp_get_experiment returns its runs (meta + latest metrics), or rd_exp_get_run run_id=<uuid> for one run.

Evaluating a run and proposing the next one

Treat each run as one hypothesis. To evaluate and iterate:

  1. Read the input (rd_exp_input): universe, train/valid/test windows, label expression, features, model + hyperparams. Note what was held fixed vs changed.
  2. Read the signal metrics (rd_exp_result.headline): IC (predictive power), ICIR (stability — |ICIR| ≥ 0.5 strong, 0.2–0.5 weak but persistent, < 0.2 noise), Rank IC/Rank ICIR. A decent IC with Rank IC ≈ 0 means the ranking is noisy even if the mean cross-section is predictive.
  3. Read the backtest (rd_exp_result.backtest): annualized_return, information_ratio, max_drawdown are excess vs the benchmark (qlib mean-daily × 238). Compare against return_annualized (raw strategy) and benchmark_annualized; check the benchmark is a sensible peer (a single high-flying stock like AAPL is a brutal benchmark for an ETF universe).
  4. Read the blotter (rd_exp_blotter.summary): n_trades/trading_days reveal turnover; total_cost vs account is the cost drag; positions show concentration. High turnover + low topk on correlated names = cost-heavy, undiversified book.
  5. Diagnose and pick ONE lever for the next run — change one thing, hold the rest fixed so the comparison is clean:
    • Weak/noisy signal (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day Ref($close,-6)/Ref($close,-1)-1), stronger regularization (reg_alpha/reg_lambda up, subsample/colsample down), or a cleaner universe (drop leveraged/duplicate names).
    • Good signal, bad book (high IC but poor excess return): raise topk for diversification, tune n_drop for rotation, reduce turnover, check total_cost.
    • Benchmark mismatch: pick an index ETF (QQQ/IVV) the universe tracks instead of a single stock.
    • Data window: a 3-month fit window is short; consider rolling/expanding if the lake history allows.
  6. Write the next run as a YAML (see section above), trace it first (rd_trace_start experiment_name=<exp>), then trigger with rd_run_workflow into that same new experiment (e.g. tac-rd-tune), and rd_trace_finish experiment_id=<N> ref_id=<run_id> when it succeeds. Then rd_exp_get_experiment to compare run-to-run. Record the hypothesis/evaluation via rd_exp_set_notes.

Example: the baseline Exp-1 Run-f29f5446 shows IC 0.071 / ICIR 0.17 / Rank IC 0.014 with excess return −0.94 ann (IR −2.23) vs a +89% ann benchmark — the 1-day signal is unstable, the topk=2 book turned 24 trades in 27 days (~1.1% cost drag) on correlated ETFs + leveraged hedges, and AAPL is an unfair benchmark. Two improvement runs are ready in tac-qlib/workflows/tune_run1_wider_5d.yaml (5-day label, topk=5, deduped 10-name universe, benchmark QQQ) and tune_run2_regularized.yaml (stronger regularization, topk=3/n_drop=2, same-day label) — trigger both into experiment_name=tac-rd-tune and compare.

Example session

# 1) train
rd_train universe=AAPL,MSFT,TSLA,USO,SLV,TLT train_start=2026-03-01 train_end=2026-05-31
        valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06
        experiment_name=tac-rd-mcp out_dir=/tmp/rd_out
# -> run_id ...

# 2) predict on the test window (from the mlflow run)
rd_predict universe=AAPL,MSFT,TSLA,USO,SLV,TLT run_id=<run_id> experiment_name=tac-rd-mcp
        train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30
        test_start=2026-07-01 test_end=2026-08-06 out_dir=/tmp/rd_out top=5

# 3) evaluate the alpha
rd_evaluate pred_path=/tmp/rd_out/pred.pkl label_path=/tmp/rd_out/label.pkl

# 4) backtest the signal
rd_backtest pred_path=/tmp/rd_out/pred.pkl start_time=2026-07-01 end_time=2026-08-06 topk=2 n_drop=1 benchmark=AAPL

# 5) one-shot equivalent — trace first, then run, then finish
rd_trace_start --rational "<hypothesis>" --experiment-name tac-rd-one-shot --evolved-from auto
rd_run_workflow config_path=tac-qlib/workflows/workflow_lgb_taclake.yaml experiment_name=tac-rd-one-shot
rd_trace_finish --id <EXPERIMENT_ID> --ref-id <run_id> --evaluation "<outcome>"

# 6) inspect that run later — given experiment_id + run_id (the /rd pages call exactly these)
rd_exp_get_experiment experiment_id=1            # -> runs with meta + latest metrics
rd_exp_input    experiment_id=1 run_id=<run_id>  # what went in: universe, windows, features, label, model
rd_exp_result   experiment_id=1 run_id=<run_id>  # what came out: IC/ICIR/Rank IC, backtest risk
rd_exp_blotter  experiment_id=1 run_id=<run_id>  # execution: P&L, positions, trades, signals
rd_exp_model    experiment_id=1 run_id=<run_id>  # hyperparameters + tree / importances
rd_exp_set_notes experiment_id=1 run_id=<run_id> hypothesis="5d label + topk5" evaluation="ICIR 0.5, ann +12%"

Notes

  • The server reads the lake lazily via the custom LakeCalendarProvider / LakeInstrumentProvider / LakeFeatureProvider; if data is missing the relevant provider raises a clear error — backfill through the tac-engine lake tools first.
  • All tools return JSON via stdio (MCP). Diagnostics/logs go to stderr.
  • MLflow runs are stored in the tracking store at $DATABASE_URL (Postgres) when set, else the unified lake sqlite mlruns.db; artifact files always live under <lake>/mlruns/<exp_id>/<run_uuid>/. Override the tracking URI with MLRUNS_URI if needed.
  • rd_run_workflow / rd_train pin each experiment's MLflow artifact_location to <lake>/mlruns so DB and artifacts stay co-located even when the server process runs from another cwd. Readers (rd_exp_*) resolve each run's artifact dir from its recorded artifact_uri, falling back to the side-by-side <lake>/mlruns layout — so runs whose artifacts were written elsewhere (e.g. <cwd>/mlruns) still display.