book: scaffold + ch00 (execution trail as spine) — evidence exp 8-31, round 3
This commit is contained in:
@@ -0,0 +1,296 @@
|
||||
---
|
||||
name: tradeac-rd
|
||||
description: Guide agents to run quant R&D on the TradeAC data lake with a Qlib-based research server exposed over MCP (tac-qlib-rd). Train GBDT (LightGBM/XGBoost) and Linear/QDA ML models on lake bars + TA features, generate cross-sectional alpha predictions, evaluate IC/Rank IC, run TopkDropout backtests with benchmark comparison, and execute one-shot YAML workflows — all through `tac_qlib.rd_server`, an MCP server in the repo venv.
|
||||
---
|
||||
|
||||
# tradeac-rd
|
||||
|
||||
Quant R&D server for the TradeAC data lake. Wraps [Qlib](https://github.com/microsoft/qlib) in a local **MCP server** (`tac_qlib.rd_server` in the repo `.venv`) and uses custom qlib data providers that read directly from the lake (see `tac-engine/skills/tradeac-lake/SKILL.md` for the lake itself, and `tac-qlib/README.md` for the package).
|
||||
|
||||
Registered in `opencode.json` as `tac-qlib-rd` — the tools below are available directly once opencode is restarted.
|
||||
|
||||
## MCP-first policy
|
||||
|
||||
- **Prefer the tac-qlib-rd MCP tools** (`rd_train`, `rd_predict`, `rd_evaluate`, `rd_backtest`, `rd_strategy_targets`, `rd_run_workflow`, `rd_exp_*`, `rd_status`) over writing scripts that reimplement the R&D loop (custom qlib glue, own train/predict/eval/backtest, hand-rolled mlruns readers, own JSON-RPC clients).
|
||||
- **NEVER script directly against the MCP server** (spawning `python -m tac_qlib.rd_server`, driving it via bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
|
||||
- Data prep (bars/features backfill) is done with the tac-engine lake MCP tools — see `tac-engine/skills/tradeac-lake/SKILL.md`. Inspect runs with `rd_exp_*` instead of reading `mlruns.db`/pickles directly.
|
||||
- If the venv is missing a runtime dep (e.g. `duckdb`, `pyarrow`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) instead of switching to another tool.
|
||||
|
||||
## The R&D loop
|
||||
|
||||
| Tool | Purpose |
|
||||
|------|---------|
|
||||
| `rd_train` | Fit a model on lake data + TA features, log to MLflow, return run metadata. |
|
||||
| `rd_predict` | Generate out-of-sample predictions from a trained model (by `model_path` or `run_id`). |
|
||||
| `rd_evaluate` | IC / Rank IC stats of a `pred.pkl` vs `label.pkl`. |
|
||||
| `rd_backtest` | TopkDropout backtest of predictions vs a benchmark, with risk metrics + artifacts. |
|
||||
| `rd_strategy_targets` | Turn a prediction's signal day into a deterministic target buy list (TopkDropout selection + sizing). |
|
||||
| `rd_run_workflow` | One-shot: run an entire YAML workflow (train → predict → sig-ana → backtest) and return metrics + artifacts. |
|
||||
| `rd_exp_list` | List MLflow experiments with run ids on the local sqlite store. |
|
||||
| `rd_exp_get_experiment` | Experiment detail: all runs (meta, metrics, notes, artifact files). |
|
||||
| `rd_exp_get_run` | Single run meta + latest metrics. |
|
||||
| `rd_exp_input` | What went into a run: qlib_init, model kwargs, dataset handler kwargs, segments, features, universe, label. |
|
||||
| `rd_exp_result` | What came out: headline IC/ICIR/Rank IC/Rank ICIR, per-day IC series, group returns, prediction stats, backtest report + risk (benchmark-relative). |
|
||||
| `rd_exp_model` | Trained model: class, hyperparameters, LightGBM tree/feature importance. |
|
||||
| `rd_exp_blotter` | Execution log: account P&L summary, daily equity, current positions, trade table, signal blotter. |
|
||||
| `rd_exp_get_notes` / `rd_exp_set_notes` | Read / write hypothesis + evaluation notes on a run. |
|
||||
| `rd_exp_delete` | **HARD-delete** an MLflow experiment: all runs (metrics/params/tags), the traced `rd_experiments` rows that reference them (FK is `ON DELETE CASCADE`), and on-disk artifacts. Irreversible — confirm with the user first. |
|
||||
| `rd_exp_delete_run` | **HARD-delete** a single MLflow run + its traced `rd_experiments` row + artifacts. Irreversible — confirm with the user first. |
|
||||
| `rd_status` | Lake + qlib readiness: data window, symbols, persisted features, qlib version. |
|
||||
|
||||
The standard flow is `rd_train` → `rd_predict` → `rd_evaluate` → `rd_backtest`; `rd_run_workflow` replaces all of it with a YAML config. The `rd_exp_*` inspection tools read saved MLflow artifacts, so every page in the R&D app (`/rd/input`, `/rd/result`, `/rd/model`, `/rd/blotter`) is backed by an MCP call (`rd_exp_input`, `rd_exp_result`, `rd_exp_model`, `rd_exp_blotter`) keyed by `experiment_id` + `run_id`.
|
||||
|
||||
## Setup / env
|
||||
|
||||
| Var | Default | Purpose |
|
||||
|-----|---------|---------|
|
||||
| `TAC_LAKE_DIR` | **required** (no default) | lake root (bars + features + metadata). Local dev: absolute path (e.g. `/home/data/lake`). |
|
||||
| `TAC_RD_MARKET` | `US` | market partition for lake reads |
|
||||
| `DATABASE_URL` | – | Postgres tracking store for MLflow (its own `experiments`/`runs`/… tables) when set |
|
||||
| `MLRUNS_URI` | postgres (`$DATABASE_URL`) or `sqlite:///<lake>/mlruns.db` | MLflow tracking URI override. Artifact files always live under `<lake>/mlruns/<exp_id>/<run_uuid>/` |
|
||||
|
||||
The server lives in the repo `.venv`; the MCP config is already registered. Restart opencode after editing `opencode.json`.
|
||||
|
||||
## Secrets policy
|
||||
|
||||
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`, `EMBEDDING_API_KEY`) in workflow YAMLs, scripts, configs, notes or committed code.
|
||||
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
|
||||
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
|
||||
- Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows — `rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else the lake sqlite. Never hardcode a `postgres://user:pass@…` URI.
|
||||
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
|
||||
|
||||
## `rd_train`
|
||||
|
||||
Args: `universe` (comma-separated), `train_start/valid_end/test_end` (`YYYY-MM-DD`), `experiment_name`, `out_dir`, `model` (`lgb` default, `xgb`, `linear`, `qda`), optional `label` (default `Ref($close,-2)/Ref($close,-1)-1`, the next-day return), `topk`/`n_drop` for later backtests.
|
||||
|
||||
- Loads 1d bars + all persisted TA features (`features/` dir) for the universe from the lake.
|
||||
- Splits into train / valid / test; fits on train with early stopping on valid.
|
||||
- Logs the run to MLflow (`run_id`), saves `params.pkl` (model) + `pred.pkl` + `label.pkl` to `out_dir`.
|
||||
- With `record_analysis=true` (default) also runs SignalRecord / SigAnaRecord (`ana_long_short`) / PortAnaRecord inside the run, so the result page gets IC/Rank IC series, long-short group returns, monthly IC and the portfolio backtest. `benchmark`, `topk`, `n_drop`, `account`, `risk_degree`, `open_cost`/`close_cost`/`min_cost` tune that backtest.
|
||||
- Returns `run_id`, `status`, `fit_seconds`, `feature_fields`, `label`, per-segment rows/date/instrument counts, and artifact paths.
|
||||
- **`wait=false`** runs the fit in a background thread and returns immediately (`status: started`, `background: true`) — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until its status is `FINISHED`, then use that `run_id`. Use it for slow windows (e.g. a 4y retrain) where a blocking MCP call can time out.
|
||||
|
||||
> If the lake lacks bars or features for `universe`, backfill first via the tac-engine `get_lake_bars` / `get_lake_ta` tools, or raise the training start date.
|
||||
|
||||
## `rd_predict`
|
||||
|
||||
Args: `universe`, `model_path` **or** `run_id`+`experiment_name` (artifact `params.pkl` is loaded from MLflow), same date ranges as `rd_train`, `out_dir`, optional `top` (number of top-scored rows in the `head` list).
|
||||
|
||||
- Rebuilds the same feature matrix for `test_start..test_end`, produces scores.
|
||||
- Writes `pred.pkl` (scores) and `label.pkl` (labels) to `out_dir`.
|
||||
- Returns paths, `count`, `date_min/max`, instruments, score distribution stats, and a small `head`.
|
||||
|
||||
## `rd_evaluate`
|
||||
|
||||
Args: `pred_path`, `label_path` (the two pkl files from `rd_train`/`rd_predict`).
|
||||
|
||||
- Returns `IC` and `RankIC` tables (`days`, `mean`, `std`, `ann_vol`, `ir`, `skew`, `kurt`, `maxdd`) and a `headline` (`IC`, `ICIR`, `Rank IC`, `Rank ICIR`).
|
||||
|
||||
## `rd_backtest`
|
||||
|
||||
Args: `pred_path`, `start_time`/`end_time`, `topk`, `n_drop`, `benchmark`, optional `out_dir`.
|
||||
|
||||
- TopkDropoutStrategy (topk long, n_drop drop), $100k account, `risk_degree 0.95`, day freq, benchmark comparison.
|
||||
- Returns `start_time`, `end_time`, `trading_days`, `risk` (`mean`, `std`, `annualized_return`, `information_ratio`, `max_drawdown`), `benchmark`, and artifact paths (`report_normal.csv`, `positions_normal.csv`, `risk.csv`).
|
||||
|
||||
## `rd_strategy_targets`
|
||||
|
||||
Args: `pred_path`, optional `signal_date` (defaults to the last day in the prediction), `account`, `risk_degree`, `topk`, `n_drop`, optional `prices` (JSON `{symbol: price}`).
|
||||
|
||||
- Applies the **exact TopkDropout selection** for one signal day: rank the cross-sectional scores, drop the top `n_drop`, take the next `topk` as buys, sized at `account × risk_degree / topk` per name. Use this to chain a prediction straight into an order list — no manual strategy replication.
|
||||
- With `prices`, floors each order to whole shares (`qty`) and reports `expected_price` / `invested`.
|
||||
- Returns `signal_date`, `per_name_notional`, the deterministic `targets` list (`symbol`, `rank`, `score`, `side`, `notional`, `qty`), and the top-20 `ranking` for context. If fewer than `topk + n_drop` names have a score that day it returns empty `targets` with a `reason`.
|
||||
|
||||
## `rd_run_workflow`
|
||||
|
||||
Args: `config_path` (YAML, see `tac-qlib/workflows/workflow_lgb_taclake.yaml`), `experiment_name`, optional `wait` (default `false`), optional `run_in_new_process` (default `false`).
|
||||
|
||||
- Runs the full pipeline (qlib `signal` + `records`), returns `run_id`, `status`, the resolved `qlib_init`/`model`/`dataset`/`records` config, and `metrics` (train/valid loss, IC/ICIR/Rank IC/Rank ICIR, and the `1day.*` backtest metrics).
|
||||
- `wait=false` (default) returns immediately with `status: started`; the workflow runs in a background thread — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until it finishes. `wait=true` blocks until completion (only for small windows that finish inside the MCP call timeout).
|
||||
- `run_in_new_process=true` runs the workflow in a **separate OS process** instead of a thread. qlib `init` sets process-global state, so this is the safe mode for concurrent or long workflows — it isolates crashes, releases memory on exit, and avoids the thread-safety race. stdout/stderr are redirected to `<lake>/logs/rd-workflow-<exp>-<ts>.log` (returned as `log_path`; the child must never write to the MCP stdio pipe). Polling works identically because the child writes to the same mlflow store. The process `pid` is returned.
|
||||
|
||||
## Tracing every run started from a chat (REQUIRED)
|
||||
|
||||
**Every experiment you start from this chat must be traced FIRST.** The R&D
|
||||
lineage (`/rd/lineage`) and the round book build on the `rd_experiments` table —
|
||||
an experiment created by `rd_run_workflow` / `rd_train` without a
|
||||
`rd_trace_start` is invisible there (no lineage node, no chat link). So before
|
||||
triggering any run, use the `rd_trace_*` MCP tools (tac-qlib-rd):
|
||||
|
||||
1. Open the trace BEFORE the run (see `tac-qlib/skills/tac-qlib-custom/SKILL.md`,
|
||||
"Experiment traceability" — the skill that owns the trace flow):
|
||||
```
|
||||
rd_trace_start rational="<what this run tests, in one line>" \
|
||||
details="<universe / features / label / model / strategy sizing>" \
|
||||
experiment_name=<the experiment you will run into> \
|
||||
evolved_from=<predecessor traced id or auto> \
|
||||
session_id="<this chat's opencode session id>"
|
||||
# -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...}
|
||||
```
|
||||
2. Run the workflow into that **same** `experiment_name`:
|
||||
```
|
||||
rd_run_workflow config_path=<yaml> experiment_name=<the experiment name>
|
||||
```
|
||||
3. On success, **finish the trace** (links the run, copies metrics/evaluation):
|
||||
```
|
||||
rd_trace_finish experiment_id=<N> ref_id=<run_id> \
|
||||
evaluation="<outcome>" metrics='{...headline...}' \
|
||||
mlruns_dir=<lake>/mlruns/<experiment_id>/<run_id>
|
||||
```
|
||||
|
||||
If you are NOT tracing (quick throwaway exploration), say so explicitly and note
|
||||
the run will not appear in the lineage graph. The default for any run started
|
||||
from a chat is to trace it.
|
||||
|
||||
## Building a workflow YAML and triggering a run
|
||||
|
||||
A workflow YAML is a qrun config: `qlib_init` (lake providers + MLflow exp manager), `task.model`, `task.dataset`, and `task.record`. Copy `tac-qlib/workflows/workflow_lgb_taclake.yaml` as the template.
|
||||
|
||||
```yaml
|
||||
# jinja is available: {%- set LAKE = TAC_LAKE_DIR %} (TAC_LAKE_DIR is required)
|
||||
qlib_init:
|
||||
provider_uri: "{{ LAKE }}"
|
||||
region: us
|
||||
expression_cache: null
|
||||
dataset_cache: null
|
||||
calendar_provider: {class: tac_qlib.data.providers.LakeCalendarProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
|
||||
instrument_provider: {class: tac_qlib.data.providers.LakeInstrumentProvider, kwargs: {lake_root: "{{ LAKE }}", market: US, markets: {}}}
|
||||
feature_provider: {class: tac_qlib.data.providers.LakeFeatureProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
|
||||
exp_manager: {class: MLflowExpManager, module_path: qlib.workflow.expm, kwargs: {uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd"}}
|
||||
|
||||
task:
|
||||
model:
|
||||
class: LGBModel
|
||||
module_path: qlib.contrib.model.gbdt
|
||||
kwargs: {loss: mse, learning_rate: 0.05, num_leaves: 15, n_estimators: 200,
|
||||
colsample_bytree: 0.8, subsample: 0.8, subsample_freq: 1,
|
||||
reg_alpha: 0.01, reg_lambda: 0.01}
|
||||
dataset:
|
||||
class: DatasetH
|
||||
module_path: qlib.data.dataset
|
||||
kwargs:
|
||||
handler:
|
||||
class: TACHandler
|
||||
module_path: tac_qlib.contrib.data.handler
|
||||
kwargs:
|
||||
instruments: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ # universe
|
||||
start_time: 2000-01-03 # lake look-back for features
|
||||
end_time: 2026-08-06
|
||||
fit_start_time: 2026-03-01 # normalization fit window
|
||||
fit_end_time: 2026-05-31
|
||||
freq: day
|
||||
lake_root: "{{ LAKE }}"
|
||||
market: US
|
||||
label: "Ref($close,-2)/Ref($close,-1)-1" # next-day return
|
||||
segments: # train/valid/test split
|
||||
train: [2026-03-01, 2026-05-31]
|
||||
valid: [2026-06-01, 2026-06-30]
|
||||
test: [2026-07-01, 2026-08-06]
|
||||
record:
|
||||
- {class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {}}
|
||||
- {class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: {ana_long_short: true, ann_scaler: 252}}
|
||||
- class: PortAnaRecord
|
||||
module_path: qlib.workflow.record_temp
|
||||
kwargs:
|
||||
config:
|
||||
strategy:
|
||||
class: TopkDropoutStrategy
|
||||
module_path: qlib.contrib.strategy
|
||||
kwargs: {signal: "<PRED>", topk: 2, n_drop: 1, only_tradable: true, risk_degree: 0.95}
|
||||
backtest:
|
||||
start_time: 2026-07-01
|
||||
end_time: 2026-08-06
|
||||
account: 1000000
|
||||
benchmark: QQQ # any symbol in the lake; empty = no benchmark
|
||||
exchange_kwargs:
|
||||
codes: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ
|
||||
deal_price: $close
|
||||
freq: day
|
||||
open_cost: 0.0005
|
||||
close_cost: 0.0015
|
||||
min_cost: 5.0
|
||||
risk_analysis_freq: 1d
|
||||
```
|
||||
|
||||
Then trigger it (each call = one new run in the named experiment):
|
||||
|
||||
```text
|
||||
rd_run_workflow config_path=tac-qlib/workflows/tune_run1_wider_5d.yaml experiment_name=tac-rd-tune
|
||||
# -> run_id <uuid>; save it, then inspect via rd_exp_*.
|
||||
```
|
||||
|
||||
The saved `config` artifact (same shape as above) is what `rd_exp_input` returns, so runs are reproducible from their YAML.
|
||||
|
||||
## Inspecting a run given experiment_id + run_id
|
||||
|
||||
The URL on the R&D app is `/rd/input|result|model|blotter?expId=<id>&run=<uuid>`; the underlying MCP calls are:
|
||||
|
||||
| You want | MCP call | Args |
|
||||
|----------|----------|------|
|
||||
| Full results (IC/ICIR/Rank IC, group returns, backtest risk) | `rd_exp_result` | `experiment_id`, `run_id` |
|
||||
| Input config (universe, windows, features, label, model) | `rd_exp_input` | `experiment_id`, `run_id` |
|
||||
| Execution blotter (P&L, positions, trades, signals) | `rd_exp_blotter` | `experiment_id`, `run_id` |
|
||||
| Model (hyperparams, tree, importances) | `rd_exp_model` | `experiment_id`, `run_id` |
|
||||
| Run notes | `rd_exp_get_notes` / `rd_exp_set_notes` | `experiment_id`, `run_id` (+ `hypothesis`/`evaluation`) |
|
||||
|
||||
Get the run ids first: `rd_exp_list` → pick an experiment → `rd_exp_get_experiment` returns its runs (meta + latest metrics), or `rd_exp_get_run run_id=<uuid>` for one run.
|
||||
|
||||
## Evaluating a run and proposing the next one
|
||||
|
||||
Treat each run as one hypothesis. To evaluate and iterate:
|
||||
|
||||
1. **Read the input** (`rd_exp_input`): universe, train/valid/test windows, label expression, features, model + hyperparams. Note what was held fixed vs changed.
|
||||
2. **Read the signal metrics** (`rd_exp_result.headline`): IC (predictive power), ICIR (stability — |ICIR| ≥ 0.5 strong, 0.2–0.5 weak but persistent, < 0.2 noise), Rank IC/Rank ICIR. A decent IC with Rank IC ≈ 0 means the ranking is noisy even if the mean cross-section is predictive.
|
||||
3. **Read the backtest** (`rd_exp_result.backtest`): `annualized_return`, `information_ratio`, `max_drawdown` are **excess vs the benchmark** (qlib mean-daily × 238). Compare against `return_annualized` (raw strategy) and `benchmark_annualized`; check the benchmark is a sensible peer (a single high-flying stock like AAPL is a brutal benchmark for an ETF universe).
|
||||
4. **Read the blotter** (`rd_exp_blotter.summary`): `n_trades`/`trading_days` reveal turnover; `total_cost` vs account is the cost drag; positions show concentration. High turnover + low topk on correlated names = cost-heavy, undiversified book.
|
||||
5. **Diagnose** and pick ONE lever for the next run — change one thing, hold the rest fixed so the comparison is clean:
|
||||
- *Weak/noisy signal* (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day `Ref($close,-6)/Ref($close,-1)-1`), stronger regularization (`reg_alpha`/`reg_lambda` up, `subsample`/`colsample` down), or a cleaner universe (drop leveraged/duplicate names).
|
||||
- *Good signal, bad book* (high IC but poor excess return): raise `topk` for diversification, tune `n_drop` for rotation, reduce turnover, check `total_cost`.
|
||||
- *Benchmark mismatch*: pick an index ETF (QQQ/IVV) the universe tracks instead of a single stock.
|
||||
- *Data window*: a 3-month fit window is short; consider rolling/expanding if the lake history allows.
|
||||
6. **Write the next run as a YAML** (see section above), **trace it first** (`rd_trace_start experiment_name=<exp>`), then trigger with `rd_run_workflow` into that **same new experiment** (e.g. `tac-rd-tune`), and `rd_trace_finish experiment_id=<N> ref_id=<run_id>` when it succeeds. Then `rd_exp_get_experiment` to compare run-to-run. Record the hypothesis/evaluation via `rd_exp_set_notes`.
|
||||
|
||||
Example: the baseline `Exp-1 Run-f29f5446` shows IC 0.071 / ICIR 0.17 / Rank IC 0.014 with excess return −0.94 ann (IR −2.23) vs a +89% ann benchmark — the 1-day signal is unstable, the topk=2 book turned 24 trades in 27 days (~1.1% cost drag) on correlated ETFs + leveraged hedges, and AAPL is an unfair benchmark. Two improvement runs are ready in `tac-qlib/workflows/tune_run1_wider_5d.yaml` (5-day label, topk=5, deduped 10-name universe, benchmark QQQ) and `tune_run2_regularized.yaml` (stronger regularization, topk=3/n_drop=2, same-day label) — trigger both into `experiment_name=tac-rd-tune` and compare.
|
||||
|
||||
## Example session
|
||||
|
||||
```text
|
||||
# 1) train
|
||||
rd_train universe=AAPL,MSFT,TSLA,USO,SLV,TLT train_start=2026-03-01 train_end=2026-05-31
|
||||
valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06
|
||||
experiment_name=tac-rd-mcp out_dir=/tmp/rd_out
|
||||
# -> run_id ...
|
||||
|
||||
# 2) predict on the test window (from the mlflow run)
|
||||
rd_predict universe=AAPL,MSFT,TSLA,USO,SLV,TLT run_id=<run_id> experiment_name=tac-rd-mcp
|
||||
train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30
|
||||
test_start=2026-07-01 test_end=2026-08-06 out_dir=/tmp/rd_out top=5
|
||||
|
||||
# 3) evaluate the alpha
|
||||
rd_evaluate pred_path=/tmp/rd_out/pred.pkl label_path=/tmp/rd_out/label.pkl
|
||||
|
||||
# 4) backtest the signal
|
||||
rd_backtest pred_path=/tmp/rd_out/pred.pkl start_time=2026-07-01 end_time=2026-08-06 topk=2 n_drop=1 benchmark=AAPL
|
||||
|
||||
# 5) one-shot equivalent — trace first, then run, then finish
|
||||
rd_trace_start --rational "<hypothesis>" --experiment-name tac-rd-one-shot --evolved-from auto
|
||||
rd_run_workflow config_path=tac-qlib/workflows/workflow_lgb_taclake.yaml experiment_name=tac-rd-one-shot
|
||||
rd_trace_finish --id <EXPERIMENT_ID> --ref-id <run_id> --evaluation "<outcome>"
|
||||
|
||||
# 6) inspect that run later — given experiment_id + run_id (the /rd pages call exactly these)
|
||||
rd_exp_get_experiment experiment_id=1 # -> runs with meta + latest metrics
|
||||
rd_exp_input experiment_id=1 run_id=<run_id> # what went in: universe, windows, features, label, model
|
||||
rd_exp_result experiment_id=1 run_id=<run_id> # what came out: IC/ICIR/Rank IC, backtest risk
|
||||
rd_exp_blotter experiment_id=1 run_id=<run_id> # execution: P&L, positions, trades, signals
|
||||
rd_exp_model experiment_id=1 run_id=<run_id> # hyperparameters + tree / importances
|
||||
rd_exp_set_notes experiment_id=1 run_id=<run_id> hypothesis="5d label + topk5" evaluation="ICIR 0.5, ann +12%"
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
- The server reads the lake lazily via the custom `LakeCalendarProvider` / `LakeInstrumentProvider` / `LakeFeatureProvider`; if data is missing the relevant provider raises a clear error — backfill through the tac-engine lake tools first.
|
||||
- All tools return JSON via stdio (MCP). Diagnostics/logs go to stderr.
|
||||
- MLflow runs are stored in the tracking store at `$DATABASE_URL` (Postgres) when
|
||||
set, else the unified lake sqlite `mlruns.db`; artifact files always live under
|
||||
`<lake>/mlruns/<exp_id>/<run_uuid>/`. Override the tracking URI with `MLRUNS_URI` if needed.
|
||||
- `rd_run_workflow` / `rd_train` pin each experiment's MLflow `artifact_location` to `<lake>/mlruns` so DB and artifacts stay co-located even when the server process runs from another cwd. Readers (`rd_exp_*`) resolve each run's artifact dir from its recorded `artifact_uri`, falling back to the side-by-side `<lake>/mlruns` layout — so runs whose artifacts were written elsewhere (e.g. `<cwd>/mlruns`) still display.
|
||||
Reference in New Issue
Block a user