book: scaffold + ch00 (execution trail as spine) — evidence exp 8-31, round 3

This commit is contained in:
TradeAC Book Agent
2026-08-18 22:35:23 +00:00
commit c93424e76c
83 changed files with 17676 additions and 0 deletions
+325
View File
@@ -0,0 +1,325 @@
---
name: tac-algo-trade
description: Guide agents to run the TradeAC scheduled algo-trading flow end-to-end. First backfill the data lake for all symbols up to the latest completed trading day, then — given a reference MLflow run (experiment_name + run_id) — re-train the same model configuration on a rolling window (4 years up to the latest completed trading day), generate fresh signals, run the selected strategy into a target order list, and place the orders on the Alpaca paper account — chaining every step from the previous one's output. Uses the `tac-engine` lake MCP tools (get_lake_bars, backfill_lake_calendar, get_lake_coverage) for data, the `tac-qlib-rd` MCP tools (rd_train, rd_predict, rd_strategy_targets, rd_exp_*) for the quant side, and the `tac-engine` MCP tools (place_order, list_orders, list_positions, get_news, ...) for execution.
---
# tac-algo-trade
Scheduled algo trading for the TradeAC paper account. Each scheduled execution (1) backfills the lake so it is current through the latest completed trading day, (2) re-trains the reference run's configuration on the most recent 4 years of lake data, (3) predicts, (4) derives an order list from the strategy, and (5) places orders on Alpaca.
This is the skill the app's scheduler invokes (`/dashboard/scheduler`). It chains strict step-to-step outputs: **do not skip ahead, do not fabricate outputs — every step consumes the artifact path returned by the previous one.**
## MCP tools
- Data side: `tac-engine` lake tools (see `tac-engine/skills/tradeac-lake/SKILL.md`) — `get_lake_coverage`, `backfill_lake_calendar`, `get_lake_bars` (lazy backfill), `get_lake_status`.
- Quant side: `tac-qlib-rd` (see `tradeac-rd/SKILL.md`) — `rd_status`, `rd_exp_get_experiment`, `rd_exp_input`, `rd_train`, `rd_predict`, `rd_strategy_targets`, `rd_exp_get_run`.
- Execution side: `tac-engine` (see `tac-engine/skills/tradeac-alpaca/SKILL.md`) — `get_account`, `list_positions`, `list_orders`, `place_order`, `get_stock_snapshot`, `get_stock_latest_quotes`, `get_news`.
- Round book: `tac-rd-book` — the execution trail (see the "Round book" section below). `round_create` / `round_update`, `fact_record`, `intent_set`, `decision_record`, `round_sync_fills`, `round_update_status`, `book_reconcile`, `trail_funnel`.
All steps use the MCP tools directly. Never hand-compute scores, read mlruns files directly (`mlruns.db` / pickles — the store is Postgres via `DATABASE_URL` when set), or script the MCP servers yourself. This is a **paper** account — trade normally, size each order per the strategy's target weight × live equity (whole shares), capped by available buying power, and skip anything untradeable.
## Inputs
- `experiment_name` — MLflow experiment of the reference run.
- `run_id` — the reference run inside that experiment (its saved `config` artifact is the source of truth for the whole pipeline).
- `strategy` (optional) — a workflow YAML from `tac-qlib/workflows/*.yaml` defining the strategy sizing (e.g. `topk` / `n_drop` / `risk_degree`, benchmark, costs). Default: the reference run's own backtest config.
- Time context: today's date in the scheduling city's timezone.
## Step 1 — Backfill the lake (data currency)
The retrain must see every symbol up to the latest available bar — do **not** train on stale data.
1. `get_lake_coverage` `{"market":"US","timeframe":"1d"}` → read each symbol's loaded window; the **last loaded date** across the universe is your backfill start.
2. `backfill_lake_calendar` `{"market":"US","symbols":"all","start":<backfill start>,"end":<today>}` → seed the trading-day set first so the 1d completeness check knows which days to expect.
3. `get_lake_bars` `{"market":"US","symbols":"all","timeframe":"1d","start":<backfill start>,"end":<today>,"lazy":true,"quiet":true}` → backfill every symbol's gap (Alpaca historical bars) and persist to the lake. Use `quiet: true` so the tool returns a per-symbol `{count, first_t, last_t}` summary instead of echoing back thousands of bar rows. Alpaca has no bar for today until the session closes, so the latest bar landed is the **latest completed trading day** `D` (for a Monday run this is Friday).
Confirm with `rd_status` (calendar range + coverage) that the lake is populated through `D`. **Output: `D`, the latest completed trading day.**
## Step 2 — Inspect the reference run
`rd_exp_get_experiment` with `experiment_id` (or `rd_exp_input` with `run_id`) → extract from the run's `config` artifact:
- handler config: `universe` (instruments), `features`, `label`, `freq`
- model kwargs: `learning_rate`, `num_leaves`, `n_estimators`, `colsample_bytree`, `subsample`, `subsample_freq`, `reg_alpha`, `reg_lambda`, `seed`
- strategy sizing: `topk` / `n_drop` / `risk_degree`, costs, benchmark
Record these — they define the retrain. **Output: config values above.**
## Step 3 — Open the traced experiment (git lineage)
Every scheduled run is a **traced experiment** on the tac-qlib-custom lineage: a row in the
`rd_experiments` table plus a per-experiment git branch in the `experiments` submodule,
forked from the predecessor's branch. **This is part of the run — do it automatically, do
not wait for the user to prompt** (see `tac-qlib/skills/tac-qlib-custom/SKILL.md`,
"Experiment traceability", for the full procedure and env vars).
1. Resolve the predecessor: if the reference run (`experiment_name` / `run_id` from Step 2)
is itself traced, reuse its traced id as `evolved_from`; otherwise use `--evolved-from auto`
(semantic search over existing rationals).
2. Open the trace — this inserts the row, forks the branch from the predecessor and pushes it (via the `rd_trace_*` MCP tools on tac-qlib-rd):
```
rd_trace_init
rd_trace_start rational="scheduled algo retrain on <D>: <ref exp>/<ref run> re-trained on 4y -> live paper orders" \
details="<universe / features / label / model / strategy sizing from the reference run config>" \
experiment_name=<THE RUN'S experiment name — see naming below> \
evolved_from=<predecessor id or auto> \
session_id="<this chat's opencode session id>"
# -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...}
```
**Experiment naming (unique per run):** every retrain runs into its OWN
experiment — `<reference experiment name>-<epoch seconds>` (e.g.
`tac-basic-short-1786883261`). The scheduler prompt names the exact
experiment for you; use that name for `rd_trace_start experiment_name`,
`rd_train`'s `experiment_name`, and the round's `experiment_name`. **Never**
reuse the reference experiment name for this run's trace node — reusing it
creates duplicate lineage entries with the same name and a wrong parent
chain (seen with `tac-basic-short`).
The tool returns `experiment_id` / `branch` as JSON — record them; every
later `rd_trace_*` call uses the id. Commit the run's files (workflow YAML /
notes) with `rd_trace_commit experiment_id=<N> message="..."` as you go.
3. **Round window** — the execution trail for `D`:
- **If the scheduler pre-created it** (your instructions name a `ROUND_ID` / `target_date` / `source`) — **skip `round_create`** and use that `ROUND_ID`. If the `D` you computed in Step 1 differs from the given `target_date`, correct it first with `round_update {round_id:<ROUND_ID>, target_date:<D>}` (weekday rule can't see NYSE holidays; the agent reconciles).
- Otherwise create it yourself (idempotent: a second scheduled run for the same day reuses the open window):
```
round_create {target_date:<D>, signal_date:<D>, source:"scheduled",
rd_experiment_id:<EXPERIMENT_ID>, experiment_name:<the run's unique experiment name>}
# -> round_id (record it; every round-book call below uses it)
```
**Output: `EXPERIMENT_ID` (and its branch), `ROUND_ID`.**
## Step 4 — Re-train with the rolling window
Call `rd_train` with the **exact same configuration** from Step 2, only the dates change:
- `train_start` = 4 years before `D` (same day-of-month), `train_end` = `D`
- **Validation is optional** — qlib supports omitting it, so omit `valid_start`/`valid_end`/`test_start`/`test_end` (pass them empty). If the tool/your run requires a holdout for sanity, use a short recent `valid` window only; never reserve data the live model needs.
- `record_analysis=false` (we only need the model; no SignalRecord/PortAnaRecord on a holdout we don't use)
- `wait=false` (recommended) — `rd_train` returns immediately and the fit runs in the background; poll `rd_exp_get_run` (or `rd_exp_list` filtered to the experiment) until the newest run's status is `FINISHED`, then take its `run_id`. With `wait=true` the call blocks until the fit completes — fine when the window is small, but a 4y LightGBM fit can outlive the MCP call timeout, which forced manual recovery in an earlier run.
- `out_dir` — the working directory for this run (e.g. `tac-algo-output`)
- `experiment_name` — the **run's unique experiment name** (the scheduler prompt names it: `<reference experiment name>-<epoch seconds>`). This is the SAME name used for `rd_trace_start experiment_name` and the round's `experiment_name`. Do not reuse the reference experiment name.
Keep the same `universe`, `features`, `label`, and every model hyper-parameter. **Output: the new run's `model_path` (and its `run_id`).**
> If a 4-year window is slower than the schedule allows, use the largest trailing window you can complete and say so in the summary — never silently shrink the horizon.
Pin the new training run to the round window:
```
round_update {round_id:<ROUND_ID>, run_id:<new run_id>, model_path:<params.pkl path>}
```
## Step 5 — Generate predictions (the signal)
Call `rd_predict` with `model_path` = the path returned by Step 4 (preferred over `run_id` since it is the freshly-trained artifact):
- `test_start` = `D`, `test_end` = `D` (the just-completed trading day — this is the signal we trade on)
- same `universe` / `features` / `label` as Step 2
- `out_dir` = the same working directory
**Output: `pred_path` (pred.pkl) and the score ranking.** The model's predicted score per instrument IS the alpha signal for day `D` — top-scored names are candidates.
Record the signal into the round book (one `fact_record` per top-scored name, plus the strategy config and the market snapshot at prediction time):
```
fact_record {round_id:<ROUND_ID>, kind:"signal_score", symbol:<ticker>, payload:{"pred":<score>, "rank":<rank>}, source:"rd_predict"}
fact_record {round_id:<ROUND_ID>, kind:"strategy_config", payload:{...strategy sizing...}, source:"reference config"}
fact_record {round_id:<ROUND_ID>, kind:"market_snapshot", payload:{<ticker>: {last:<px>, change_pct:<%>, vol:<vol>, updated:<ts>}, ...}, source:"get_stock_snapshots / get_stock_latest_quotes"}
```
`market_snapshot` freezes the market state **when the prediction was made** — the latest price / % change / volume per universe name, so the signal can later be judged against what the market looked like at that moment.
## Step 6 — Run the configured strategy, derive the target order list
**First pull the current portfolio — it is an input to the strategy step** (the order list is a delta, not a full rebuild):
- `get_account` → cash / buying power **and total equity** (equity sizes the positions; buying power caps total buys)
- `list_positions` → current holdings and their market value
Then run the strategy **exactly as it was configured in the reference run** — this works for any model/strategy, not just TopkDropout. The reference run's saved `config` artifact (from `rd_exp_input`, Step 2) carries the strategy configuration from its backtest/record block (e.g. `TopkDropoutStrategy` kwargs: `topk`, `n_drop`, `risk_degree`, or any custom strategy's own kwargs, plus costs, `account`, `benchmark`). **Use those values — not tool defaults.** The model's score is the signal the strategy consumes; the strategy's config decides allocation.
Call `rd_strategy_targets` with:
- `pred_path` = the signal from Step 5
- the run-configured `topk` / `n_drop` / `risk_degree` (from the reference run config)
- `account` = the **live account equity** from `get_account` (a new account is not a $1M book — sizing against `$1M` when equity is far smaller produces oversized orders)
- `prices` = a JSON `{symbol: price}` of latest quotes (from `get_stock_latest_quotes`) so the tool floors each order to whole shares (`qty`) and reports `expected_price` / `invested`
- `risk_limits` = the round's risk-limit spec JSON (see below) — the SAME spec that `rd_backtest` uses, so live gating is provable against backtest
- `equity` / `peak_equity` = live equity and its trailing peak (from `get_portfolio_history`) when `risk_limits.drawdown_pause_pct` is set
The tool applies the exact TopkDropout selection on day `D`: rank the cross-sectional scores, **drop the top `n_drop`**, take the next `topk` as buys, sized at `account × risk_degree / topk` per name. It then applies `risk_limits` as pre-gates — liquidity floor (drops names with avg daily dollar volume below `liquidity_floor_adv`), per-name `size_cap_pct` of equity, `concentration_cap_pct` of equity on total deployed, and `drawdown_pause_pct` (equity ≤ (1−pause)×peak ⇒ no buys). **Output: the deterministic target buy list** (`symbol`, `rank`, `score`, `side`, `notional`, `qty`), the full `ranking`, and `risk_limits_applied` (which limits cut what — record it). **No manual strategy replication** (an earlier run's hand-rolled sizing silently dropped the n_drop and bought the wrong names).
> If the strategy in the run/workflow config does not fit TopkDropout's `topk`/`n_drop`/`risk_degree`, apply the strategy's own rules to the Step 5 scores directly to derive the target portfolio, still bounded by `get_account` buying power and today's `list_positions`.
Then convert the target portfolio into an order list against the current holdings:
- For each target ticker compute the **delta** vs. what the account already holds: buy the shortfall, sell the excess. Do not blindly re-buy names already held, and do not sell names that are not in the portfolio.
- **Fresh account (no positions):** the target portfolio is entirely new buys — emit no sell orders, and size each buy from the tool's `qty` (or `notional` ÷ latest quote), capped by buying power.
- Skip any ticker whose delta is ~0 (already at target) so you don't churn held names.
- Cap total buy size to available buying power. Drop any ticker with no score in Step 5 or no tradable quote.
**Output: the explicit order list** (ticker, side, qty, order type).
**Write the target into the round book** — this is the intent the round reconciles against (versions auto-increment; a second strategy pass for the same round supersedes the first):
```
fact_record {round_id:<ROUND_ID>, kind:"account_state", payload:{"equity":<live equity>, "buying_power":<bp>}, source:"get_account"}
fact_record {round_id:<ROUND_ID>, kind:"position_state", symbol:<ticker>, payload:{"shares":<held>}, source:"list_positions"}
fact_record {round_id:<ROUND_ID>, kind:"risk_check", payload:{"risk_limits":{...spec...}, "applied":{...risk_limits_applied from the tool...}, "equity":<equity>, "peak_equity":<peak>}, source:"rd_strategy_targets"}
round_update {round_id:<ROUND_ID>, account_equity_at_sizing:<live equity>, strategy_snapshot:{topk, n_drop, risk_degree, costs, benchmark, risk_limits:{liquidity_floor_adv?, size_cap_pct?, concentration_cap_pct?, drawdown_pause_pct?}}}
intent_set {round_id:<ROUND_ID>, target_portfolio:[{symbol, side, qty, notional, expected_price, score, rank}...],
raw_strategy_output:{...the strategy output as computed...}, reason:"topk<N> from <ref run>"}
```
**Risk-limit spec (B)**: the round's `risk_limits` (a JSON map with any of `liquidity_floor_adv`, `size_cap_pct`, `concentration_cap_pct`, `drawdown_pause_pct`) is the single source of truth — **the same spec is passed to `rd_backtest` when calibrating** (B2), folded into `rd_train`'s PortAnaRecord via `risk_degree`, and consulted by `rd_strategy_targets` live. Store it verbatim in `strategy_snapshot.risk_limits`. When the tool's `risk_limits_applied` reports a limit that cut targets (dropped liquidity / capped sizing / drawdown pause), record it — the audit trail proves the limit fired live exactly as the calibration predicted. If `drawdown_pause_pct` fired and produced an empty target list, **settle the round as open→settled with no orders** rather than forcing buys (that is the intended behavior).
**Record the evidence behind each selected name** — the feature snapshot and the decision rationale, so the fact table can answer *why this symbol was ranked top-K*:
- `symbol_features` — the model-input feature values that produced the score on day `D` (the top features by `rd_exp_model` importance, plus the handful most relevant for that name — e.g. trend slopes, RSI, volume/vol ratios, MACD):
```
get_lake_ta {symbol:<ticker>, timeframe:"1d", start:<~60d before D>, end:<D>, persist:true, quiet:true} # (re)compute TA + sp_* columns up to D
get_lake_features {symbol:<ticker>, timeframe:"1d", start:<D>, end:<D>} # read the D row; if 0 rows, the persisted features are stale -> persist first as above
rd_exp_model {run_id:<new training run_id>, tree_id:0, max_depth:4} # feature_importances + tree nodes
fact_record {round_id:<ROUND_ID>, kind:"symbol_features", symbol:<ticker>,
payload:{"score":<score>, "rank":<rank>, "features":{<top feature>:<value>, ...}}, source:"get_lake_features / rd_exp_model"}
```
`get_lake_features` returns 0 rows for day `D` when the persisted feature files were last written before `D` (they are per-symbol parquet files that only extend to the last time they were computed). In that case **first persist** with `get_lake_ta ... persist:true` (and `get_lake_sp` when the model uses `sp_*` columns — the rd_train feature list from Step 2 tells you which), then read `get_lake_features` for `D` again — it must return a row per ticker.
- `decision_justification` — **concise** (under 500 words total, aim for 2–4 sentences per name): why the model ranked the symbol top-K. Ground it in the actual data — the `rd_exp_model` tree path (which feature conditions led the row down the high-score branch) and the `symbol_features` values — not generic commentary:
```
fact_record {round_id:<ROUND_ID>, kind:"decision_justification", symbol:<ticker>,
payload:{"score":<score>, "rank":<rank>, "why": "<2-4 sentences, e.g. 'strong 5d trend slope + rising volume ratio put TSLA above $sp_trend_slope_60 threshold, sending it down the high-score branch (leaf value +0.0545); RSI recovering but not overbought.'>"},
source:"rd_exp_model tree + feature snapshot"}
```
## Step 7 — Execution context + news sentiment gate
Before placing anything, per candidate ticker:
1. `get_account` (buying power), `list_orders` (open orders), `list_positions` (current holdings).
2. `get_stock_snapshot` / `get_stock_latest_quotes` → sanity-check each quote: skip tickers with no quote, a stale/illiquid quote (wide spread or near-zero volume), or a halt. Use the latest quote, not just the model score, for sizing and order type.
3. `get_news` with `symbols=<ticker>`, `limit=20`, `include_content=true` → assign a sentiment score **−3 (strongly negative) … +3 (strongly positive)**.
**Sentiment gate:** if sentiment strongly contradicts the signal — a **BUY** with sentiment ≤ −2 or a **SELL** with sentiment ≥ +2 — **cancel** that order and record it as `cancelled: sentiment conflict`. Tickers with no news or neutral sentiment (−1..+1) trade normally.
Record the evidence per candidate into the round book (so the reconcile step can explain every skip):
```
fact_record {round_id:<ROUND_ID>, kind:"quote", symbol:<ticker>, payload:{bid, ask, last, spread_bps}, source:"get_stock_snapshot"}
fact_record {round_id:<ROUND_ID>, kind:"news_sentiment", symbol:<ticker>, payload:{"sentiment":<−3..+3>, "headline":<top headline>}, source:"get_news"}
```
## Step 8 — Place orders on Alpaca
For each surviving order in the Step 6 list (respecting the gate): call the `tac-engine` `place_order` tool with the ticker, side, qty and order type. Then verify with `list_orders` / `list_positions` that the intended changes went through.
**Record every decision in the round book** — placed orders AND deliberate skips, each with its reason (this is what the reconcile / funnel view reads):
```
# each placed order (order id from the place_order response):
decision_record {round_id:<ROUND_ID>, symbol:<ticker>, side:<buy|sell>, qty:<qty>, order_type:<type>,
expected_price:<last quote px>, status:"placed", reason:"placed",
intent_id:<intent id from intent_set>, alpaca_order_id:<alpaca order id>, client_order_id:<cl id>}
# each gate cancel / skip (delta≈0, no quote, illiquid, halt, bp cap, sentiment conflict, no score, risk limit):
decision_record {round_id:<ROUND_ID>, symbol:<ticker>, side:<side>, qty:<qty>, status:"skipped",
reason:"sentiment_conflict"|"illiquid"|"no_quote"|"halt"|"delta_zero"|"bp_cap"|"no_score"|"risk_limit",
reason_detail:<short why>, intent_id:<intent id>}
```
**Sync fills** — pull Alpaca's order state into the round (pass the `list_orders` output as `orders` so no API call is needed; unmatched orders are reported back):
```
round_sync_fills {round_id:<ROUND_ID>, orders:[{id, client_order_id, symbol, side, qty, filled_qty, filled_avg_price, status}...]}
```
## Step 9 — Evidence check, close the traced experiment + summarize
**Evidence gate — run this BEFORE committing/closing. Do not skip, do not "summarize only".** Query the round and confirm every evidence kind is present; record anything missing right now, then re-query:
```
fact_query {round_id:<ROUND_ID>} # or per-kind: fact_query {round_id:<ROUND_ID>, kind:"<kind>"}
```
For each of the per-universe kinds (`signal_score`, `market_snapshot`, `quote`, `news_sentiment`, `symbol_features`, `decision_justification`) count that you recorded one per symbol you processed; `strategy_config`, `account_state`, `position_state` once each. If any kind is missing or any target symbol is missing from a kind, **go back and `fact_record` it now** (use the persist→read recipe in Step 6 for `symbol_features`). Only when every kind above is present, proceed:
1. Commit the run artifacts to the experiment branch: `rd_trace_commit experiment_id=<EXPERIMENT_ID> message="algo run <D>: orders placed"`.
2. Close the lineage — re-embeds the rational/details, records metrics/evaluation, commits + pushes:
```
rd_trace_finish experiment_id=<EXPERIMENT_ID> \
ref_id=<new training run_id from Step 4> \
evaluation="<outcome of today's trade: target vs placed, cancellations>" \
metrics='{"n_buys":N,"n_sells":M,"n_cancelled":K}' \
mlruns_dir=<lake>/mlruns/<exp_id>/<run_id>
```
3. **Settle the round** — reconcile and close the window:
```
book_reconcile {round_id:<ROUND_ID>} # residual vs target, per-symbol reasons
trail_funnel {round_id:<ROUND_ID>} # targets -> decided -> placed -> filled, skips by reason
round_update_status {round_id:<ROUND_ID>, status:"settled", summary_metrics:{...funnel + invested...}}
```
4. **Close the loop** — record the round's execution economics for the next run's tuning:
- `round_metrics` → the round's invested notional, turnover, slippage bps, estimated cost, cost-as-% of gross (the `fetchPriorRoundFeedback` in the scheduler injects these into the NEXT run's prompt automatically).
- If the round had fills, run `rd_factor_attribution` over the round window (pass the `get_portfolio_history` equity curve as `portfolio_equity`, benchmark e.g. `IVV`, realized slippage+cost bps from `round_metrics`, expected values from the calibration) and record the result:
```
fact_record {round_id:<ROUND_ID>, kind:"attribution", payload:{beta, alpha_annualized_pct, pnl_beta, pnl_alpha, drift_alarm}, source:"rd_factor_attribution"}
```
- A `drift_alarm` in the attribution means live execution cost is deviating from the backtest assumption — re-run `rd_risk_calibrate` before the next round and tighten sizing/limits.
5. End your reply with the compact summary: date `D`, reference run (`experiment_name` / `run_id`), new training run (`run_id` / `model_path`), window (4y → `D`), number of scores, top names, per-ticker sentiment scores, what was bought/sold, which orders were cancelled by the sentiment gate (and why), and any skipped trades (with reasons).
## Example
```
experiment_name=tac-rd run_id=<ref-uuid> strategy=tune_run1_wider_5d.yaml
1. get_lake_coverage {US,1d} -> last loaded date; backfill_lake_calendar; get_lake_bars lazy -> lake current -> D
2. rd_exp_input run_id=<ref-uuid> -> universe=all, features=KR..(ta fields), label=Ref($close,-2)/Ref($close,-1)-1, lr=0.05, leaves=15 ... topk/n_drop from the run's backtest config
3. EXP_NEW=<ref exp>-<epoch seconds> # unique per run (scheduler names it)
rd_trace_init && rd_trace_start experiment_name=$EXP_NEW evolved_from=auto -> experiment_id / branch
# scheduler usually pre-creates the round (ROUND_ID in the instructions) -> skip round_create, use it
round_create {target_date:<D>, signal_date:<D>, source:"scheduled", rd_experiment_id:<EXPERIMENT_ID>, experiment_name:$EXP_NEW} -> ROUND_ID
4. rd_train experiment_name=$EXP_NEW train_start=<D-4y> train_end=<D> record_analysis=false wait=false out_dir=tac-algo-output
# -> returns immediately; poll rd_exp_get_run until status FINISHED -> run_id <new-uuid>, model_path tac-algo-output/params.pkl
round_update {round_id:<ROUND_ID>, run_id:<new-uuid>, model_path:"tac-algo-output/params.pkl"}
5. rd_predict model_path=tac-algo-output/params.pkl test_start=<D> test_end=<D>
# -> pred_path tac-algo-output/pred.pkl, score head ...
fact_record {kind:"signal_score", symbol:<ticker>, payload:{pred, rank}} per top name
6. get_account + list_positions # current portfolio as strategy input; account=live equity
get_stock_latest_quotes -> prices JSON for sizing
rd_strategy_targets pred_path=tac-algo-output/pred.pkl signal_date=<D> \
topk=<from run config> n_drop=<from run config> risk_degree=<from run config> account=<live equity> prices='{...}' \
risk_limits='{"liquidity_floor_adv":5000000,"size_cap_pct":8,"concentration_cap_pct":30}' equity=<equity> peak_equity=<peak>
# -> deterministic target buys (symbol/rank/score/notional/qty); delta vs list_positions -> order list (fresh account = all buys)
fact_record {kind:"risk_check", payload:{risk_limits:{...}, applied:{...risk_limits_applied...}, equity, peak_equity}}
round_update {round_id:<ROUND_ID>, account_equity_at_sizing:<equity>, strategy_snapshot:{topk, n_drop, risk_degree, costs, benchmark, risk_limits:{...}}}
intent_set {round_id:<ROUND_ID>, target_portfolio:[{symbol, side, qty, expected_price, score, rank}]} -> intent_id
7. get_news per ticker -> sentiment gate; fact_record quote + news_sentiment per ticker
8. place_order ... per surviving delta; decision_record per placed + skipped (with reason)
round_sync_fills {round_id:<ROUND_ID>, orders:[...list_orders output...]}
9. rd_trace_commit experiment_id=<EXPERIMENT_ID> + rd_trace_finish experiment_id=<EXPERIMENT_ID> ref_id=<new-uuid>
book_reconcile + trail_funnel; round_update_status {status:"settled", summary_metrics:{...}}; summary
```
## Round book — the execution trail
Every scheduled run writes its decision→fill trail to Postgres via the `tac-rd-book`
tools, mirroring the `/dashboard/rounds` UI. The round is the link between the scheduler
run, the traced experiment, and the actual account activity:
```
scheduler_runs ──► ROUND ──► rd_experiments
│ fact_events evidence: signal_score / market_snapshot / quote / news_sentiment / account_state / position_state / symbol_features / decision_justification / risk_check
│ round_intents versioned target portfolios (new version supersedes old)
│ round_decisions per-symbol: placed OR skipped, each with a reason (incl. risk_limit)
└──► round_orders execution rows (Alpaca order id + fills), synced via round_sync_fills
```
`book_reconcile` returns the per-symbol residual (target qty − filled qty, with the reason
it did not fill) plus cash/BP impact, slippage bps and estimated cost — that is the answer
to "why is the account not at the target portfolio". `round_metrics` reports the round
roll-ups (invested notional, turnover, slippage bps, estimated cost, cost-as-% of gross);
`trail_funnel` gives the counts (targets → decided → placed → filled, skips by reason).
All surface unchanged in the UI. When a round fires `drawdown_pause_pct`, its `round_metrics`
will show `invested_notional: 0` — that is the pause working, not a broken round.
+529
View File
@@ -0,0 +1,529 @@
---
name: tac-qlib-custom
description: "Guide agents to customize and extend Qlib on the TradeAC R&D stack — how to configure workflow YAMLs (qlib_init, model, dataset/handler, processors, records, PortAnaRecord strategies), how to extend Qlib classes wired into those workflows (custom Model, BaseStrategy, DataHandler, Record), and the empirically-tested knobs from this repo (RankIC early-stopping, stochastic-control strategies, stochastic-process features, catch22/GARCH/Hurst/signature). Also encodes the experiment traceability loop: every backtest runs as a workflow-with-recorder, is recorded in the Postgres experiments table (rationale/details/evaluation/metrics with pgvector embeddings, evolution chain) and on a per-experiment git branch that is committed + pushed. Companion to tradeac-rd (MCP run tools) and tradeac-lake (parquet lake)."
---
# tac-qlib-custom
Customizing and extending Qlib on the TradeAC stack. This skill encodes what was
learned from actual experiments in this repo: how a workflow YAML maps to Qlib
classes, how to write a custom class that the YAML can load, and which training /
strategy / feature knobs measurably moved IC, RankIC and the backtest.
Read `tac-qlib/skills/tradeac-rd/SKILL.md` for the MCP run/inspect tools and
`tac-qlib/README.md` for the package layout. The venv is `/app/.venv`
(qlib 0.1.dev2066); `tac_qlib` is installed into the venv's `site-packages`
(editable copy under `/opt/venv/.../tac_qlib/`), so **any new module must be
copied to `/opt/venv/lib/python3.12/site-packages/tac_qlib/...` too** (or use an
editable install) before `rd_run_workflow` can import it.
## MCP-first policy
- **Drive every backtest and run through the `tac-qlib-rd` MCP tools** (`rd_run_workflow`,
`rd_train`, `rd_predict`, `rd_exp_*`) and the tac-engine lake tools for data prep. Do not
reimplement them with ad-hoc scripts (custom qlib glue, own mlruns readers, direct
JSON-RPC/stdio clients).
- **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server` /
`tac-engine`, bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and
ask the user to confirm first**.
- The traceability bookkeeping (Postgres `rd_experiments` row + pgvector embeddings +
branch-per-experiment git) is exposed as the **`rd_trace_*` MCP tools** on the tac-qlib-rd
server — use those, not bash scripts. Data prep, training, evaluation and backtests also go
through MCP tools.
- If the venv is missing a runtime dep (`duckdb`, `pyarrow`, feature libs), lazy-install it
(`uv pip install --python $VIRTUAL_ENV/bin/python <pkg>`) instead of switching tools.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or
credential-bearing URLs (`DATABASE_URL`, `GIT_PASS`, `EMBEDDING_API_KEY`) in
workflow YAMLs, scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on
`.env`). That pulls secrets into this session and leaks them to any agent
sharing it.
- When a tool or command needs an env var, ASK the user to set it in the
environment (shell/container env, or the user-owned `.env`) and reference it
by name (`$VAR`), never by value. If it's missing, report which variable is
required instead of reading it yourself.
- Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows —
`rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else
the lake sqlite. Never hardcode a `postgres://user:pass@…` URI.
- If you find a committed secret, flag it, remove it, and replace it with a
placeholder. (The `rd_trace_*` MCP tools' commit guard blocks adding
credential-shaped lines.)
## How a workflow YAML maps to Qlib classes
A workflow YAML (`tac-qlib/workflows/*.yaml`) is rendered by Jinja (vars like
`{{ LAKE }}` from `TAC_LAKE_DIR`) then executed by `qrun` / `rd_run_workflow`.
Every block is a Qlib class reference resolved by `module_path` + `class`:
```yaml
{%- set LAKE = TAC_LAKE_DIR %}
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
calendar_provider: # custom tac-qlib providers read the parquet lake
class: LakeCalendarProvider
module_path: tac_qlib.data.providers
instrument_provider: # ... (markets: {} => lake universe)
feature_provider: # LakeFeatureProvider: routes $open..$volume from bars,
class: LakeFeatureProvider # $<ta-lib/sp_*> from features parquet, $amount derived
exp_manager:
class: MLflowExpManager
module_path: qlib.workflow.expm
kwargs: { uri: "sqlite:///{{ LAKE }}/mlruns.db", default_exp_name: "my-exp" }
task:
model: # <MODEL BLOCK> — custom model → new module_path
class: RankICLGBModel
module_path: tac_qlib.contrib.model.rank_gbdt
kwargs: { loss: mse, learning_rate: 0.02, num_leaves: 31, ... }
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler: # <HANDLER BLOCK> — feature selection + processors live here
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: "SPY,QQQ,..."
start_time: 2015-01-03
end_time: 2026-08-10
fit_start_time: 2015-01-03 # processors fit on this window
fit_end_time: 2025-09-01
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-6)/Ref($close,-1)-1" # 5d forward return
feature_fields: "$open,$high,$low,$close,$vwap,$volume,sp_ret,sp_ou_zscore,..."
infer_processors: # feature-time transforms, fit on fit_*
- { class: DropAllNaN, kwargs: {} }
- { class: ProcessInf, kwargs: {} }
- { class: CSRankNorm, kwargs: {} } # per-day cross-sectional rank
- { class: ZScoreNorm, kwargs: {} }
- { class: Fillna, kwargs: {} }
segments:
train: [2015-01-03, 2025-09-01]
valid: [2025-09-03, 2026-01-03]
test: [2026-01-04, 2026-08-10]
record: # each entry records one artifact type to the run
- { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} }
- { class: SigAnaRecord, module_path: qlib.workflow.record_temp,
kwargs: { ana_long_short: true, ann_scaler: 252 } }
- { class: PortAnaRecord, module_path: qlib.workflow.record_temp,
kwargs: { config: { strategy: <STRATEGY BLOCK>, backtest: {...} }, risk_analysis_freq: 1d } }
```
`rd_run_workflow config_path=<yaml> experiment_name=<exp>` runs it; the MCP call
may time out for long runs (RankIC tuning, heavy feature sets) — the run keeps
executing; poll via `rd_exp_list` / `rd_exp_get_run` on the returned experiment.
## Experiment traceability (DB + git + embeddings)
Every backtest you run as an agent MUST be tracked: it runs as a workflow with the
`record` block (SignalRecord/SigAnaRecord/PortAnaRecord → MLflow artifacts on disk
under `<lake>/mlruns/<exp_id>/<run_id>`), and a row is written to the Postgres
`experiments` table plus a git branch per experiment. The `tac-app` UI owns the
schema (Drizzle migrations in `tac-app/drizzle/`); this skill's `lib/` scripts are
the executor the agent drives.
**Trigger the lineage as part of the run — automatically, not on prompt.** Any
time you execute a qlib workflow (`rd_run_workflow`) or a train/predict pipeline
on this stack, the traceability bookkeeping is part of that run, not a separate
step the user must ask for: open the traced experiment with `rd_trace_start`
before running, commit intermediates with `rd_trace_commit`, and close it with
`rd_trace_finish` after — without waiting to be prompted (see "The
per-experiment procedure" below).
### Env vars
| Var | Purpose |
|-----|---------|
| `DATABASE_URL` | Postgres URL for the `rd_experiments` table AND the MLflow tracking store (set in repo `.env`) |
| `EMBEDDING_API_BASE_URL` | embedding POST endpoint (e.g. `https://embd.h.lizhao.net/embeddings`) |
| `EMBEDDING_API_KEY` | basic-auth credential (`user:pass` form is supported) |
| `GIT_USER` / `GIT_PASS` | git remote credentials for push/fetch |
| `GIT_REPO_URL` | experiment git repo tracked by the `experiments` submodule (branches are pushed here) |
| `TAC_LAKE_DIR` | lake root (mlruns artifact files live under it) |
The experiment repo is the **`experiments` git submodule** at the workspace root
(`<repo-root>/experiments`), always tracking `$GIT_REPO_URL`. `rd_trace_init`
creates/validates it; it errors if `experiments/` exists but points at a
different URL. There is no `TAC_EXP_GIT_DIR` — the submodule path IS the
experiment repo, and ALL experiment/backtest changes (workflow YAMLs, notes,
outputs) must live inside it, never in the parent tradeac repo.
### The `rd_experiments` table
Owned by tac-app's Drizzle schema (`tac-app/src/db/schema.ts`); `rd_trace_init`
can `init` it idempotently. The table is named **`rd_experiments`** (NOT
`experiments`) because MLflow's Postgres tracking store creates its own
`experiments` table in the same database. Key columns: `id` (PK), `rational` +
`rational_embedding` (pgvector `vector(384)`), `details` + `details_embedding`,
`evaluation`, `metrics` (jsonb), `evolved_from` (FK → rd_experiments.id),
`start_ts`/`end_ts`, `git_branch`, `experiment_ref_id`, `mlruns_dir`, `status`.
`experiment_ref_id` holds the **mlflow run id** returned by `rd_run_workflow` and
is an FK to MLflow's `runs(run_uuid)` (added by `rd_trace_init` after the
mlflow store tables exist — MLflow creates `runs` lazily).
Tracking store: **Postgres `$DATABASE_URL`** (MLflow's own tables) when set,
falling back to the unified lake sqlite `sqlite:///<lake>/mlruns.db`. Artifact
files always stay on disk under `<lake>/mlruns/<exp_id>/<run_id>/artifacts`.
Embedding model: `michaelfeil/bge-small-en-v1.5` (384-dim, **512-token context**).
Rational/details are written paper-summary style (≤512 tokens) and embedded verbatim —
NEVER truncate; if a text is longer, summarize it first (the embed helper rejects
over-limit input).
### Git repo + branch-per-experiment
The experiment repo is the `experiments` submodule at the workspace root
(`<repo-root>/experiments`, tracking `$GIT_REPO_URL`). The `rd_trace_*` MCP
tools handle it, and every git operation is scoped to that submodule —
experiments NEVER stage or push parent-repo (tradeac) files.
- `rd_trace_init` creates/validates the submodule and the base branch. If
`experiments/` does not exist it runs `git clone $GIT_REPO_URL experiments`;
if it exists but tracks a different URL, init errors out.
- Base branch: `main` (or `master`). If the submodule is empty, a seed commit is
made and pushed so there are commits to fork from.
- Every experiment runs on its own branch `exp/<id>-<slug>`.
- `evolved_from` resolution (in order):
1. If the wizard prompt explicitly says `evolved_from=<id>` (run wizard click on an
existing experiment) — use that id directly.
2. Otherwise `--evolved-from auto`: the user prompt / rational is embedded and
cosine-searched over the `experiments.rational_embedding` column; the top hit
above the similarity threshold (0.5) becomes `evolved_from`.
3. Otherwise (first experiment, or a new chat with no predecessor) — no evolved_from;
fork from `main`'s latest commits.
- The new branch is forked from the **evolved-from experiment's branch** (its latest
commits), or from `main` when there is no predecessor — so experiment lineages form
a git branch chain.
- On every finish, and for intermediate steps, changes are committed + pushed.
### Custom code is part of the lineage (code snapshot)
Custom contrib modules (`tac_qlib/contrib/model/`, `tac_qlib/contrib/strategy/`,
`tac_qlib/contrib/data/`, `tac_qlib/data/providers.py`) live in the **parent**
tradeac repo, not in the `experiments/` submodule — so they are normally invisible
to the experiment branch and a descendant forking from it would reinvent them.
The lineage tooling fixes this: **every experiment branch carries a `code/`
snapshot of exactly the qlib extension code that run depended on**, so descendants
reuse it instead of re-authoring it.
- `rd_trace_start` and `rd_trace_finish` automatically snapshot the default paths
(`tac-qlib/tac_qlib/contrib`, `tac-qlib/tac_qlib/data`) into
`<experiments>/code/<parent-relative-path>` on the experiment branch.
- `rd_trace_snapshot` snapshots mid-run (e.g. after writing a
new custom model) without waiting for finish.
- The snapshot also writes `code/MANIFEST.txt` recording the **parent-repo HEAD
commit** and the per-file blob hashes it was taken from — so a run can be traced
back to the exact parent commit that produced its custom code.
- Descendants: the custom modules your run needs are under `code/tac_qlib/...` on the
evolved-from branch. Reuse them (copy/`git show`) instead of writing new ones; check
`code/MANIFEST.txt` to see which parent commit they came from and port fixes back.
- Guardrail exception: parent-repo changes under `tac_qlib/tac_qlib/contrib` and
`tac_qlib/tac_qlib/data` are **expected** (they are the snapshotted code);
`parent_changes` reports them as a note, not a violation. Any OTHER parent change
is still a guardrail violation.
Guardrail — experiments must NOT introduce side effects to the parent repo:
- Write workflow YAMLs, notes and experiment outputs ONLY inside
`<repo-root>/experiments/` (they are committed on the experiment branch).
- Never `git add`/commit/stage anything in the parent tradeac repo.
- Run `rd_trace_guard` to list any parent
changes outside the submodule pointer; `rd_trace_finish` also surfaces them.
Revert any accidental parent edits before finishing.
- If an experiment reveals a PRODUCT change (workflow template, skill, tac-app),
propose it separately for the tradeac repo — do not mix it into the experiment
branch.
The `rd_trace_*` MCP tools perform git operations with the mandated credential
helper (from `GIT_USER` / `GIT_PASS`), so you do not need to construct it by hand.
### The per-experiment procedure
**Use the `rd_trace_*` MCP tools (tac-qlib-rd)** — they replace the old
`trace.sh`/`trace_db.py` scripts. The server is long-lived (psycopg imported
once, DB connection reused per call) and every tool returns one JSON object, so
no output parsing is needed:
```text
# 0. ensure ready (rd_experiments table + experiments git repo + base main)
rd_trace_init
# 1. start — inserts the row, resolves evolved_from, forks+pushes the branch.
# Returns {experiment_id, branch, evolved_from, base_branch} as JSON.
rd_trace_start rational="5-day forward label, RankIC early stop, 50-ETF universe" \
details="LGBModel mse lr=0.02 num_leaves=15 num_boost_round=3000; TopkDropout topk=2; benchmark QQQ" \
experiment_name="tac-rd-expN" \
evolved_from="auto" \
session_id="<this chat's opencode session id, if started from a chat>"
# -> {"experiment_id": N, "branch": "exp/N-...", "evolved_from": ..., "base_branch": ...}
# 2. write the workflow YAML INSIDE the experiments submodule
# (e.g. <repo-root>/experiments/workflows/<exp>/workflow.yaml), then commit it:
rd_trace_commit experiment_id=<N> message="add workflow yaml"
# 2b. if the workflow uses a NEW custom module, snapshot it onto the branch
# (start/finish auto-snapshot contrib+data; do this to capture mid-run):
rd_trace_snapshot experiment_id=<N> # default contrib+data
# or: rd_trace_snapshot experiment_id=<N> paths="tac-qlib/tac_qlib/contrib/model/rank_gbdt.py"
# 3. run the backtest through the WORKFLOW with the recorder (MUST write mlruns):
rd_run_workflow config_path=<repo-root>/experiments/workflows/<exp>/workflow.yaml experiment_name=tac-rd-expN
# -> returns run_id (= experiment_ref_id) + metrics
# 4. inspect with rd_exp_result / rd_exp_blotter, then finish — updates the row
# (re-embeds rational/details, sets metrics/eval/end_ts), snapshots the custom
# code, and commits+pushes. finish also surfaces parent-repo side effects.
rd_trace_finish experiment_id=<N> \
ref_id=<mlflow-run-id> \
evaluation="IC 0.0645, RankIC 0.075; net excess +0.85% ann" \
metrics='{"IC":0.0645,"RankIC":0.075,"ann_excess":0.85}' \
mlruns_dir=<lake>/mlruns/<exp_id>/<run_id>
```
Helpers (MCP tools): `rd_trace_search` (semantic), `rd_trace_get` (one row),
`rd_trace_list`, `rd_trace_mlruns_dir` (resolves the mlruns dir for an
experiment name), `rd_trace_guard` (parent-repo side-effect check).
Rules:
- **Always** run backtests as workflows with the `record` block (req 2) — never a bare
`rd_backtest` for a traced experiment.
- **Always** open the lineage (`rd_trace_start`) BEFORE the run and **Always**
`rd_trace_finish` + push after it completes (req 5) — this happens as part of the run,
do not wait for the user to ask; intermediate `rd_trace_commit` is encouraged (req 5).
- **Always** snapshot the custom qlib code (`rd_trace_snapshot`, or rely on the
auto-snapshot at start/finish) so the experiment branch carries the exact contrib/data
modules the run used — descendants fork and reuse `code/` instead of reinventing it.
- Keep rational/details ≤ 512 tokens (paper-summary style) so embeddings are exact —
no truncation.
- **Confine experiments to the `experiments/` submodule** — never write to, stage, or
commit parent tradeac repo files; run `rd_trace_guard` to check for side effects.
(Custom code edits under `tac-qlib/tac_qlib/contrib` and `.../data` are the sanctioned
exception — they are the snapshotted modules; see "Custom code is part of the lineage".)
- **Follow the Secrets policy above** — no secrets in files, no reading `.env*`, ask the
user to set env vars; use `uri: "sqlite:///mlruns.db"` for the tracking store.
- Workflow YAMLs are jinja-rendered with `os.environ` as the context, so env-var
placeholders work (`{%- set LAKE = TAC_LAKE_DIR %}` then `{{ LAKE }}`). Use them for
paths/config — never for secrets that get committed.
## Extending Qlib — the 4 class families you can override
### 1. Custom Model (train-time) — `tac_qlib/contrib/model/`
Subclass `qlib.contrib.model.gbdt.LGBModel` (or `qlib.model.base.BaseModel`) and
implement `fit(dataset, ...)` + `predict(dataset)`. `LGBModel.fit` calls
`self._prepare_data(dataset)` → `lgb.Dataset`s, then `lgb.train` with
`early_stopping` on the valid set. Override points that matter:
- `_prepare_data` → build the `lgb.Dataset` with `group=` (per-day query groups)
when you need ranking metrics per trading day.
- `fit` → change what early-stops training (the biggest IC/backtest lever, see §Knobs).
- `predict` → return the Series keyed (datetime, instrument).
Reference: `tac_qlib/tac_qlib/contrib/model/rank_gbdt.py` — `RankICLGBModel`
subclasses `LGBModel`, adds per-day `group` in `_prepare_data`, injects
`feval=rankic_feval` (mean per-day Spearman) into `lgb.train`, and forces
`metric='None'` + `first_metric_only=True` so early-stopping tracks RankIC only.
### 2. Custom Strategy (backtest-time) — `tac_qlib/contrib/strategy/`
Subclass `qlib.contrib.strategy.signal_strategy.BaseSignalStrategy` (which wraps
`qlib.strategy.base.BaseStrategy`) and implement:
```python
def generate_trade_decision(self, execute_result=None):
# trade_step, trade_start/end = self.trade_calendar.get_step_time(trade_step)
# pred = self.signal.get_signal(start_time=pred_shift, end_time=pred_shift) # shift=-1 => signal known at t-1
# self.trade_position / self.trade_exchange / self.trade_calendar injected by the executor
# build qlib.backtest.Order(stock_id, amount, start_time, end_time, direction=Order.BUY/SELL)
# return TradeDecisionWO(orders, self)
```
Wire it into the YAML under `PortAnaRecord.config.strategy`:
```yaml
strategy:
class: OptimalStopControl
module_path: tac_qlib.contrib.strategy.optimal_stop
kwargs:
signal: "<PRED>" # placeholder replaced with the recorded pred
topk: 10
entry_pct: 0.85
exit_pct: 0.7
max_hold_days: 10
min_hold_days: 2
sl: -0.08
risk_degree: 0.95
```
Reference: `tac_qlib/tac_qlib/contrib/strategy/optimal_stop.py`
(`OptimalStopControl` — entry gated by cross-sectional signal percentile, exits
by percentile/time/stop-loss, equal-weight control sizing).
### 3. Custom DataHandler / processors — `tac_qlib/contrib/data/handler.py`
`TACHandler(DataHandlerLP)` already wraps the lake via `QlibDataLoader` +
`LakeFeatureProvider`. Key config surface (all usable from YAML without new code):
- `feature_fields` — explicit list; the handler prefixes `$` and de-dups. Anything
the provider can route is usable: bar fields, `$amount` (v*vw), and any column
present in the lake `features/.../symbol=*.parquet` files.
- `infer_processors` / `learn_processors` — add `CSRankNorm`, `CSZScoreNorm`
(label), `ZScoreNorm`, `DropnaLabel`, `Fillna`, etc. `DropAllNaN` is a
tac-qlib processor (drops all-NaN columns on the fit window).
- `label` — any qlib expression, e.g. `Ref($close,-6)/Ref($close,-1)-1`.
To add a *new feature family*: compute it once (see `examples/sp_features.py` +
`examples/persist_sp_features.py`), persist extra columns into
`features/market=US/timeframe=1d/symbol=*.parquet` (drop stale `sp_*` columns
first on re-runs), then reference them in `feature_fields`.
**The Rust engine already ships the SP feature pipeline as a lake MCP tool**:
`get_lake_sp` (tac-engine, stochastic-rs) computes `sp_ou_*`, `sp_hmm_*`,
`sp_jump_*`, `sp_rv*`/`sp_vol_ratio_*` (+ `sp_rv_ac1`, `sp_rv_cv_22`),
`sp_max_up`/`sp_max_down`, `sp_trend_slope_*`, `sp_logp`,
`sp_hurst_exponent`, `sp_sig_*` (levels 1/2 at lag 1 and 5),
`sp_rskew_*`/`sp_rkurt_*`/`sp_dsv_*` (realized moments via stochastic-rs
`realized`) + `sp_ret` from lake bars and persists them into
the feature parquets (replacing stale `sp_*`), all in one call:
```json
{"symbol": "AAPL", "timeframe": "1d", "start": "2015-01-03", "end": "2026-08-10", "fit_end": "2025-09-01"}
```
`fit_end` pins the Gaussian-HMM fit to the train window (no lookahead), matching
the `FIT_END` convention. **Deferred families** (`garch`, `entropy`, `catch22`)
are still computed with the Python `sp_features.py` path until their ports land.
Note two deliberate differences vs the Python reference: the Rust HMM uses the
causal *forward filter* (`filtered_state_probs`) rather than hmmlearn's smoothed
`predict_proba`, and `hurst` is estimated on the returns series directly
(`take_differences=false`) rather than the reference's double-differenced
`kind="random_walk"` — regime *state* assignments agree, probability levels are
comparable but not identical.
### 4. Custom Record (artifact writers)
Subclass `qlib.workflow.record_temp.SignalRecord` / a `Record` and log metrics +
artifacts into the MLflow run. There is no shipped example Record in `contrib/`
yet — write one against the pattern in `qlib.workflow.record_temp` when a
workflow needs a bespoke simulator (e.g. beta-neutral 3L/3S) that
`PortAnaRecord` doesn't cover.
## Empirical knobs that moved the numbers (measured on the 50-ETF lake)
All experiments used: 50-ETF universe, train 2015-01-03..2025-09-01 / valid
2025-09-03..2026-01-03 / test 2026-01-04..2026-08-10, benchmark SPY, TopkDropout
or OptimalStopControl, costs open 0.0005 / close 0.0015 / min 5.
> **Rank-dimension reminder**: when the goal is to improve the *ranking* quality of
> a signal (RankIC, long-short spread, top-decile precision), do NOT reinvent the
> stack — use the contrib modules already shipped and verified in this repo:
> `tac_qlib.contrib.model.rank_gbdt.RankICLGBModel` (early-stops training on
> per-day cross-sectional RankIC, `metric='None'` + `first_metric_only`) and
> `tac_qlib.contrib.strategy.optimal_stop.OptimalStopControl` (entry/exit gated by
> signal percentile instead of raw levels). Both are loadable from a workflow YAML
> via `module_path` — see the canonical `tac-qlib/workflows/workflow_lgb_sp5d_rankic.yaml`
> (rank dimension: model) and `workflow_lgb_sp5d_optstop.yaml` (rank dimension:
> portfolio construction). Verified end-to-end on 2026-01-04..2026-08-10:
> RankIC 0.071 / net-of-cost excess +20.7% ann (IR 0.70) vs SPY. Only write a new
> custom Model/Strategy when these proven paths are insufficient.
### Label
- **5-day forward return `Ref($close,-6)/Ref($close,-1)-1` ≫ 2-day.** IC nearly
tripled (0.0207 → 0.0645 standalone; the biggest single lever found). The 2-day
target is too noisy.
### Features
- **Stochastic-process features beat hand-rolled TA.** 55-feature set: OU
(`sp_ou_*`), 2-state HMM (`sp_hmm_*`), jump intensity (`sp_jump_*`, incl.
`sp_max_up`/`sp_max_down`), HARRV vol (`sp_rv*` + `sp_rv_ac1`/`sp_rv_cv_22`),
trend (`sp_trend_slope_*`, `sp_logp`), GARCH (`sp_garch_*`), Hurst
(`sp_hurst_exponent`), path signatures (`sp_sig_*`, lag 1 & 5), entropy
(`sp_ent_*`), realized moments (`sp_rskew_*`/`sp_rkurt_*`/`sp_dsv_*`),
catch22 (`sp_c22_*`). IC 0.036 → 0.047 vs the 19-feature v1.
- **Do NOT add ta-lib indicators on top** (SP+TA, 74 feats): IC dropped
0.047 → 0.031, RankIC 0.047 → 0.020. They're redundant with rv22/hmm/garch/catch22
and dilute CSRankNorm + LGBM.
- **CSRankNorm** (per-day cross-sectional rank) is important for the rank signal.
- Warm-up rows persist as all-NaN feature rows — expected; DropAllNaN/DropnaLabel
handle them.
### Model / training loop
- **LambdaRank / rank_xendcg objectives FAIL here** (RankIC → ~0): with only ~50
"documents" per query the rank gradient is noise.
- **Early-stopping metric beats objective.** MSE objective + early-stop on a
**RankIC feval** (mean per-day Spearman) lifted RankIC 0.047 → 0.075 (standalone).
- **The workflow gap was qlib's training loop**: `lgb.train` default
`first_metric_only=False` + `metric=l2` keeps training while l2 improves after
RankIC peaks. `RankICLGBModel` sets `metric='None'` + `first_metric_only=True`
so early-stopping tracks RankIC only.
- **RankIC-only early stop + bigger/smaller budget is the win**: `num_boost_round
3000`, `learning_rate 0.02`, `early_stopping_rounds 200`, `min_data_in_leaf 20`,
`lambda_l2 0.5` → test excess **+9.1% ann w/o cost (IR 1.03, maxDD −3.8%)** and
**+0.85% ann after costs** — the only config that beat SPY net. Note IC/RankIC
themselves were slightly lower (0.042) than the 500-tree run (0.051); the tuned
budget selects the iteration maximizing *valid* RankIC, converting to realized
excess return.
### Strategy / portfolio construction
- **Long-only construction leaves the edge on the table.** The SP-5d signal has
long-short **+31.6% ann (Sharpe 2.51)**, but TopkDropout long-only ≈ flat vs SPY,
and OptimalStopControl underperformed (valid-window threshold overfit: valid
+7.5% → test −17.7% on one calibration).
- **Costs eat most of the gross edge** (+9.1% → +0.85% net). Reduce turnover or go
long-short to widen the net edge.
- OptimalStopControl thresholds must be calibrated on the *valid* window and are
sensitive to overfit — prefer robust defaults or penalize turnover in selection.
## Gotchas
- **Installed package copy**: `tac_qlib` in the venv is a copy under
`/opt/venv/lib/python3.12/site-packages/tac_qlib/`. After editing any
`tac_qlib/contrib/**` module, `cp` it there or the workflow imports the stale
version. New subpackages need `mkdir -p` first.
- `qlib.backtest` exports `Order` but not `OrderDir`/`Position` at top level —
import `Order` from `qlib.backtest`, `OrderDir`/`TradeDecisionWO` from
`qlib.backtest.decision`, `Position` from `qlib.backtest.position`.
- `qlib.backtest.high_performance_ds` may not export `Order` in this build — don't
import from it.
- HMM / GARCH / catch22 features must not see test data at fit time: fit the HMM
on the train window only (`fit_end=FIT_END`), and compute rolling windows ending
at each day. GARCH/entropy use a stride + forward-fill for speed (~5x).
- `pycatch22`, `arch`, `hurst`, `antropy`, `hmmlearn` are required for the full
feature set; install with `uv pip install --python /app/.venv/bin/python <pkg>`
(a C compiler is needed for `pycatch22`). `duckdb` and `pyarrow` are declared in
`tac-qlib/pyproject.toml`; if a workflow import fails on either, lazy-install with
`uv pip install --python /app/.venv/bin/python duckdb pyarrow`.
- `rd_run_workflow` defaults to `wait=false`: it returns immediately with
`status: started` and the workflow runs in a background thread — poll
`rd_exp_get_run` / `rd_exp_list` for the newest run of the experiment
(status `RUNNING` until it finishes), then reuse its `run_id`. Pass
`wait=true` only for small windows that finish within the MCP call timeout.
- After fixing a YAML model/handler change, remember both `/app/tac-qlib/...` and
the `/opt/venv` copy stay in sync.
## Files this skill is based on
Minimal, runnable examples live next to this skill in `examples/` — they are the
canonical reference for every artifact the skill describes:
- Workflows (full `record` block → MLflow on disk):
- `examples/workflow_minimal.yaml` — the canonical backtest template (req: every
traced backtest runs through a workflow like this via `rd_run_workflow`)
- `examples/workflow_rankic.yaml` — RankIC-early-stop model wired in
- Repo workflows for reference: `tac-qlib/workflows/workflow_lgb_taclake.yaml`,
`tune_run1_wider_5d.yaml`, `tune_run2_regularized.yaml`, `tune_run3_label5d_clean_universe.yaml`,
`tune_run4_fix_universe_longtrain.yaml`, `tune_run5_longtest.yaml`
- Models: `examples/model_rank_gbdt.py` (`RankICLGBModel`: per-day groups +
`feval=rankic` + `metric='None'`). Repo: `tac_qlib/contrib/model/rank_gbdt.py`
- Strategies: `examples/strategy_optimal_stop.py` (`OptimalStopControl`),
`examples/strategy_beta_neutral.py` (doc-only 3L/3S stub — pattern for a
custom strategy + Record; not wired into the package)
- Handler: `examples/handler.py` (how to subclass `TACHandler`); repo:
`tac_qlib/contrib/data/handler.py`; providers: `tac_qlib/data/providers.py`
- Feature engineering: `examples/sp_features.py` (OU + Hurst) and
`examples/persist_sp_features.py` (persist `sp_*` into the lake features parquet)
- Ranking experiments: `examples/run_rank_objectives.py` (mse vs lambdarank vs
rank_xendcg ablation on the lake)
- Optstop calibration: `examples/run_optstop_compare.py` (valid-window grid +
overfit warning)
- Traceability tooling: the `rd_trace_*` MCP tools (tac-qlib-rd,
`tac_qlib/trace.py`) — see the traceability section above
@@ -0,0 +1,70 @@
"""Minimal custom DataHandler — how to extend TACHandler for a new feature family.
`TACHandler(DataHandlerLP)` already routes lake bars + ta-lib features via
`LakeFeatureProvider` (see tac_qlib/contrib/data/handler.py). To add a NEW
feature family (computed once, persisted into the lake features parquet — see
examples/persist_sp_features.py), you only need to:
1. persist extra columns into features/market=US/timeframe=1d/symbol=*.parquet
2. list them in `feature_fields` (they are prefixed with `$` and de-duped)
A subclass is only needed when the feature must be computed *inside* the qlib
pipeline (e.g. as an extra processor). This file sketches that pattern.
Reference handler structure (from tac_qlib/contrib/data/handler.py):
class TACHandler(DataHandlerLP):
def __init__(self, instruments, start_time, end_time, freq,
fit_start_time=None, fit_end_time=None,
feature_fields=None, label=None, lake_root=None, market="US",
infer_processors=None, learn_processors=None, **kwargs):
loader = QlibDataLoader(configured=(feature_fields or self.DEFAULT_FIELDS), freq=freq)
super().__init__(instruments, start_time, end_time, freq=freq,
data_loader=loader,
infer_processors=infer_processors or DEFAULT_INFER_PROCESSORS,
learn_processors=learn_processors or DEFAULT_LEARN_PROCESSORS,
fit_start_time=fit_start_time, fit_end_time=fit_end_time,
process_type=DataHandlerLP.PTYPE_A, **kwargs)
"""
from __future__ import annotations
from typing import Any, List, Optional
from tac_qlib.contrib.data.handler import DEFAULT_INFER_PROCESSORS, DEFAULT_LEARN_PROCESSORS, TACHandler
class CustomFeaturesHandler(TACHandler):
"""TACHandler variant that also loads the lake feature columns passed in.
Usage from YAML — only the handler kwargs change:
handler:
class: CustomFeaturesHandler
module_path: tac_qlib.contrib.data.handler # after adding this class there
kwargs:
instruments: AAPL,MSFT,QQQ
start_time: 2026-03-01
end_time: 2026-08-06
freq: day
lake_root: "{{ LAKE }}"
market: US
feature_fields: "$close,sp_ou_alpha,sp_hurst_exponent"
label: "Ref($close,-6)/Ref($close,-1)-1"
"""
def __init__(
self,
feature_fields: Optional[List[str]] = None,
infer_processors: Optional[List[Any]] = None,
learn_processors: Optional[List[Any]] = None,
**kwargs: Any,
):
# `feature_fields` are passed through with the leading `$` stripped by
# TACHandler; infer/learn default to the lake-tuned processor stacks.
super().__init__(
feature_fields=feature_fields,
infer_processors=infer_processors or DEFAULT_INFER_PROCESSORS,
learn_processors=learn_processors or DEFAULT_LEARN_PROCESSORS,
**kwargs,
)
@@ -0,0 +1,78 @@
"""Minimal RankIC early-stopping LightGBM model (the biggest IC/backtest lever).
Drop-in replacement for `qlib.contrib.model.gbdt.LGBModel` in a workflow YAML:
task.model:
class: RankICLGBModel
module_path: tac_qlib.contrib.model.rank_gbdt
kwargs: { loss: mse, learning_rate: 0.02, num_boost_round: 3000,
early_stopping_rounds: 200, lambda_l2: 0.5 }
What it changes vs stock LGBModel:
* `_prepare_data` builds `lgb.Dataset` with per-day `group` query groups, so
metrics are computed per trading day.
* `fit` injects `feval=rankic_feval` (mean per-day Spearman) into `lgb.train`
and forces `metric='None'` + `first_metric_only=True` so early stopping
tracks RankIC — not l2, which keeps improving after RankIC peaks.
Why: with ~50 instruments per day, ranking objectives (lambda_rank/xendcg)
produce near-zero RankIC; MSE objective + RankIC early-stop is what lifts it.
Install: copy to tac_qlib/contrib/model/rank_gbdt.py AND the /opt/venv copy.
"""
from __future__ import annotations
from typing import Any, Dict
import numpy as np
import pandas as pd
from qlib.contrib.model.gbdt import LGBModel
def rankic_feval(preds: np.ndarray, dataset) -> tuple[str, float, bool]:
"""Mean per-day Spearman rank IC between predictions and the label."""
label = dataset.get_label()
group = dataset.get_group() if hasattr(dataset, "get_group") else None
if group is None:
return "rankic", _spearman(preds, label), False
start = 0
ics = []
for g in group:
sl = slice(start, start + g)
start += g
ics.append(_spearman(preds[sl], label[sl]))
return "rankic", float(np.mean(ics)), False
def _spearman(x: np.ndarray, y: np.ndarray) -> float:
if len(x) < 2:
return 0.0
from scipy.stats import spearmanr
rho, _ = spearmanr(x, y)
return float(rho) if rho == rho else 0.0
class RankICLGBModel(LGBModel):
"""LGBModel with per-day query groups and RankIC-only early stopping."""
def _prepare_data(self, dataset, *args, **kwargs):
"""Attach per-day group sizes to the train/valid lgb.Dataset."""
dtrain, dvalid = super()._prepare_data(dataset, *args, **kwargs)
for d, index in ((dtrain, dataset.get_index_by_segment("train")), (dvalid, dataset.get_index_by_segment("valid"))):
if d is not None and index is not None:
# group by calendar day in order
days = pd.Series([i[0] for i in index])
group = days.value_counts().sort_index().tolist()
d.set_group(np.array(group, dtype=np.int32))
return dtrain, dvalid
def fit(self, dataset, evals_result: Dict[str, Any] | None = None, **kwargs):
# force RankIC-only early stopping
kwargs.setdefault("feval", rankic_feval)
kwargs.setdefault("metric", "None")
kwargs.setdefault("first_metric_only", True)
return super().fit(dataset, evals_result=evals_result, **kwargs)
@@ -0,0 +1,66 @@
"""Minimal persistence of computed SP features into the lake features parquet.
Flow: compute sp_* features per symbol (examples/sp_features.py) and MERGE them
into features/market=US/timeframe=1d/symbol=*.parquet so TACHandler /
LakeFeatureProvider can route `$sp_ou_theta` etc. from the workflow YAML.
Run after backfilling bars; re-run drops stale sp_* columns first (see note).
python examples/persist_sp_features.py --market US --timeframe 1d
"""
from __future__ import annotations
import argparse
import os
import pandas as pd
from tac_qlib.data.config import LakeConfig, NON_FEATURE_COLUMNS
from examples.sp_features import build_sp_features
#: columns owned by this feature family (replaced on re-runs, never duplicated)
SP_PREFIX = "sp_"
def persist_symbol(lake: LakeConfig, timeframe: str, symbol: str) -> None:
bars_path = lake.bar_path(timeframe, symbol)
feats_path = lake.features_path(timeframe, symbol)
if not bars_path.exists():
return
bars = pd.read_parquet(bars_path)
feats = build_sp_features(bars)
# bars have a single 't'/'date' column; align feature rows to it
feats = feats.drop(columns=[c for c in NON_FEATURE_COLUMNS if c in feats.columns], errors="ignore")
feats_path.parent.mkdir(parents=True, exist_ok=True)
if feats_path.exists():
existing = pd.read_parquet(feats_path)
# drop stale sp_* columns before merging (idempotent re-runs)
existing = existing[[c for c in existing.columns if not c.startswith(SP_PREFIX)]]
merged = pd.merge(existing, feats, on="t", how="left", suffixes=("", "_dup"))
merged = merged.loc[:, ~merged.columns.str.endswith("_dup")]
# keep original column order + new sp_* appended
merged.to_parquet(feats_path, index=False)
else:
feats.to_parquet(feats_path, index=False)
print(f"persisted {symbol}: {len(feats.columns) - 1} sp_* features")
def main() -> None:
ap = argparse.ArgumentParser()
ap.add_argument("--market", default="US")
ap.add_argument("--timeframe", default="1d")
ap.add_argument("--symbols", default="", help="comma-separated; default: all lake symbols")
args = ap.parse_args()
lake_root = os.environ.get("TAC_LAKE_DIR")
if not lake_root:
raise SystemExit("TAC_LAKE_DIR is required")
lake = LakeConfig(lake_root, args.market)
symbols = [s.strip().upper() for s in args.symbols.split(",") if s.strip()] or lake.load_symbols()
for symbol in symbols:
persist_symbol(lake, args.timeframe, symbol)
if __name__ == "__main__":
main()
@@ -0,0 +1,69 @@
"""Minimal OptimalStopControl threshold calibration — valid-window grid search.
This repo found OptimalStopControl thresholds overfit the valid window (valid
+7.5% → test −17.7% on one calibration). This script runs a small grid over
(entry_pct, exit_pct, max_hold_days) on the VALID window, reports per-config
excess return + turnover, and warns when the best valid config is a spike.
Reference repo impl: tac-qlib/examples/run_optstop_compare.py.
python examples/run_optstop_compare.py --universe AAPL,MSFT,QQQ
"""
from __future__ import annotations
import argparse
import itertools
import os
import pandas as pd
GRID = {
"entry_pct": [0.7, 0.85, 0.95],
"exit_pct": [0.5, 0.7],
"max_hold_days": [5, 10],
}
def evaluate_config(lake_root: str, universe: list[str], window: tuple, config: dict) -> dict:
"""Simplified stand-in: train the RankIC model, backtest OptimalStopControl
on `window`, return (ann_excess_return, turnover, max_drawdown).
The real repo impl calls qlib.backtest with the strategy and reads
report_normal.csv + risk.csv. Keep the interface here so the grid loop is
reusable.
"""
# placeholder — plug in the real backtest here
return {"ann_excess": 0.0, "turnover": 0.0, "max_dd": 0.0}
def main() -> None:
ap = argparse.ArgumentParser()
ap.add_argument("--lake-root", default=os.environ.get("TAC_LAKE_DIR", ""))
ap.add_argument("--universe", default="AAPL,MSFT,QQQ,IVV,SMH,TLT")
args = ap.parse_args()
universe = [s.strip().upper() for s in args.universe.split(",")]
valid = ("2026-06-01", "2026-06-30")
test = ("2026-07-01", "2026-08-06")
keys = list(GRID)
results = []
for combo in itertools.product(*[GRID[k] for k in keys]):
config = dict(zip(keys, combo))
v = evaluate_config(args.lake_root, universe, valid, config)
t = evaluate_config(args.lake_root, universe, test, config)
results.append({**config, "valid_excess": v["ann_excess"], "test_excess": t["ann_excess"]})
df = pd.DataFrame(results).sort_values("valid_excess", ascending=False)
print(df.head(10).to_string(index=False))
# Overfit check: how far is the best-valid config from the median test config?
med = df["test_excess"].median()
best = df.iloc[0]
print(f"\nmedian test excess: {med:+.3f} | best-valid test excess: {best['test_excess']:+.3f}")
if abs(best["test_excess"] - med) > 0.10:
print("WARNING: best-valid config is an outlier on test — likely overfit, prefer robust defaults")
if __name__ == "__main__":
main()
@@ -0,0 +1,90 @@
"""Minimal ranking-objective ablation loop — why lambda_rank fails here.
This repo found that with only ~50 instruments per day the rank-gradient
objectives (lambdarank / rank_xendcg) produce near-zero RankIC, while MSE
objective + RankIC early-stop is the winner. This script replays that check by
training a few LightGBM variants on the same lake split and printing RankIC.
Reference repo impl: tac-qlib/examples/run_rank_objectives.py.
python examples/run_rank_objectives.py --universe AAPL,MSFT,QQQ
"""
from __future__ import annotations
import argparse
import os
import numpy as np
import pandas as pd
OBJECTIVES = ["mse", "lambdarank", "rank_xendcg"]
def load_frame(lake_root: str, universe: list[str], start: str, end: str) -> pd.DataFrame:
"""Stack lake bars into a qlib-like (datetime, instrument) frame."""
from tac_qlib.data.config import LakeConfig
lake = LakeConfig(lake_root, "US")
frames = []
for sym in universe:
p = lake.bar_path("1d", sym)
if p.exists():
df = pd.read_parquet(p)[["t", "c"]].rename(columns={"t": "datetime", "c": "close"})
df["instrument"] = sym
frames.append(df)
out = pd.concat(frames, ignore_index=True)
out["datetime"] = pd.to_datetime(out["datetime"])
out = out[(out["datetime"] >= start) & (out["datetime"] <= end)]
return out.set_index(["datetime", "instrument"])
def label_5d(frame: pd.DataFrame) -> pd.Series:
close = frame["close"].unstack()
lbl = close.shift(-6) / close.shift(-1) - 1
return lbl.stack().rename("label")
def train_one(lake_root: str, universe: list[str], objective: str, train: tuple, test: tuple):
import lightgbm as lgb
frame = load_frame(lake_root, universe, train[0], test[1])
label = label_5d(frame)
data = pd.concat([frame["close"], label], axis=1).dropna()
tr = data.loc[(data.index.get_level_values(0) >= train[0]) & (data.index.get_level_values(0) <= train[1])]
te = data.loc[(data.index.get_level_values(0) >= test[0]) & (data.index.get_level_values(0) <= test[1])]
dtrain = lgb.Dataset(tr[["close"]], label=tr["label"])
dtest = lgb.Dataset(te[["close"]], label=te["label"])
params = {"objective": objective, "learning_rate": 0.05, "num_leaves": 15, "verbosity": -1}
model = lgb.train(params, dtrain, num_boost_round=100, valid_sets=[dtest])
pred = model.predict(te[["close"]], num_iteration=model.best_iteration)
label_te = te["label"].to_numpy()
# per-day RankIC
days = te.index.get_level_values(0).unique()
ics = []
for d in days:
m = te.index.get_level_values(0) == d
if m.sum() >= 3:
ics.append(pd.Series(pred[m]).rank().corr(pd.Series(label_te[m]).rank()))
return float(np.nanmean(ics))
def main() -> None:
ap = argparse.ArgumentParser()
ap.add_argument("--lake-root", default=os.environ.get("TAC_LAKE_DIR", ""))
ap.add_argument("--universe", default="AAPL,MSFT,QQQ,IVV,SMH,TLT")
args = ap.parse_args()
universe = [s.strip().upper() for s in args.universe.split(",")]
train = ("2026-03-01", "2026-05-31")
test = ("2026-07-01", "2026-08-06")
print(f"{'objective':<14}{'test RankIC':>12}")
for obj in OBJECTIVES:
ic = train_one(args.lake_root, universe, obj, train, test)
print(f"{obj:<14}{ic:>12.4f}")
if __name__ == "__main__":
main()
@@ -0,0 +1,80 @@
"""Minimal stochastic-process feature computation — OU mean-reversion + Hurst.
These are the features that beat hand-rolled TA in this repo's 50-ETF runs.
Compute them per symbol on a rolling window ENDING at each day (never let them
see test data at fit time — see the HMM/GARCH note in SKILL.md).
Reference repo impl: tac-qlib/examples/sp_features.py (full 55-feature set:
OU, HMM, jump, HARRV, trend, GARCH, Hurst, path signatures, entropy, catch22).
"""
from __future__ import annotations
import numpy as np
import pandas as pd
#: rolling window for feature computation (days)
LOOKBACK = 250
def compute_ou_features(close: pd.Series) -> pd.DataFrame:
"""Ornstein-Uhlenbeck fit: theta (reversion speed), sigma (vol), residual z.
OU: dx_t = theta (mu - x_t) dt + sigma dW_t (theta is the mean-reversion
speed; higher = faster reversion = tradable mean-reversion signal).
Rolling OLS of dx on lagged log-price gives theta = -b (reversion speed);
sigma is the residual std. Vectorized via rolling cov/var.
"""
logp = np.log(close)
dx = logp.diff()
x_prev = logp.shift(1)
df = pd.DataFrame({"dx": dx, "x": x_prev})
out = pd.DataFrame(index=close.index, dtype=float)
cov = df["dx"].rolling(LOOKBACK, min_periods=30).cov(df["x"])
var = df["x"].rolling(LOOKBACK, min_periods=30).var()
theta = (-cov / var).rename("sp_ou_theta")
out["sp_ou_theta"] = theta
out["sp_ou_sigma"] = df["dx"].rolling(LOOKBACK, min_periods=30).std()
# standardized residual z = (x - mu) / sigma of the fitted process
mu = df["x"].rolling(LOOKBACK, min_periods=30).mean()
scale = np.sqrt(np.clip(1 / (2 * theta + 1e-9), 0, None))
out["sp_ou_zscore"] = (df["x"] - mu) / (out["sp_ou_sigma"] * scale)
return out
def compute_hurst(close: pd.Series, lookback: int = 100) -> pd.Series:
"""Rolling Hurst exponent via rescaled range (R/S). H>0.5 = trending."""
def _hurst(x: np.ndarray) -> float:
if len(x) < 20:
return np.nan
lags = range(2, min(len(x) // 2, 50))
tau = []
for lag in lags:
diff = x[lag:] - x[:-lag]
tau.append(np.sqrt(np.std(diff)))
tau = np.array(tau)
lags = np.array(lags, dtype=float)
poly = np.polyfit(np.log(lags), np.log(tau), 1)
return float(poly[0])
return close.rolling(lookback, min_periods=20).apply(lambda w: _hurst(w.to_numpy()), raw=False).rename(
"sp_hurst_exponent"
)
def build_sp_features(bars: pd.DataFrame) -> pd.DataFrame:
"""bars: lake 1d bars indexed by (datetime, instrument) or a symbol frame."""
if isinstance(bars.index, pd.MultiIndex):
frames = []
for inst, sub in bars.groupby(level=1):
close = sub.droplevel(1)["close"]
feats = pd.concat([compute_ou_features(close), compute_hurst(close)], axis=1)
feats["instrument"] = inst
frames.append(feats.reset_index())
out = pd.concat(frames).set_index(["datetime", "instrument"])
else:
close = bars["close"]
out = pd.concat([compute_ou_features(close), compute_hurst(close)], axis=1)
return out
@@ -0,0 +1,82 @@
"""Minimal beta-neutral 3L/3S strategy + record — stub of tac_qlib/contrib/strategy/beta_neutral.py.
Strategy side: subclass BaseSignalStrategy, hold ~3 long + 3 short equally
weighted (dollar-neutral) with TP/SL and a hard close at the horizon. The beta
comes from regression of daily returns on the benchmark in `_prepare_betas`.
Record side (BetaNeutralRecord): a custom `Record` that simulates the 3L/3S
portfolio after training and logs report / trades / risk.csv into the MLflow
run — the pattern to follow for any custom Record.
Wire the record into the workflow YAML:
record:
- class: BetaNeutralRecord
module_path: tac_qlib.contrib.strategy.beta_neutral
kwargs: { benchmark: QQQ, n_long: 3, n_short: 3 }
"""
from __future__ import annotations
from typing import Any, Dict, List
import pandas as pd
from qlib.backtest import Order
from qlib.backtest.decision import OrderDir, TradeDecisionWO
from qlib.contrib.strategy.signal_strategy import BaseSignalStrategy
class BetaNeutralStrategy(BaseSignalStrategy):
"""3 long / 3 short dollar-neutral template with TP/SL and hard close."""
def __init__(self, *, n_long: int = 3, n_short: int = 3, tp: float = 0.06, sl: float = -0.05, **kwargs: Any):
super().__init__(**kwargs)
self.n_long = n_long
self.n_short = n_short
self.tp = tp
self.sl = sl
def generate_trade_decision(self, execute_result=None):
trade_step = self.trade_calendar.get_trade_step()
start_time, end_time = self.trade_calendar.get_step_time(trade_step)
pred_start, pred_end = self.trade_calendar.get_step_time(trade_step - 1)
pred = self.signal.get_signal(start_time=pred_start, end_time=pred_end)
orders: List[Order] = []
if pred is not None and len(pred):
daily = pred.groupby(level=0).mean().iloc[-1].dropna().sort_values()
longs = daily.tail(self.n_long).index.tolist()
shorts = daily.head(self.n_short).index.tolist()
for inst in longs:
orders.append(self._order(inst, 1, start_time, end_time))
for inst in shorts:
orders.append(self._order(inst, -1, start_time, end_time))
return TradeDecisionWO(orders, self)
def _order(self, inst, direction, start_time, end_time):
price = self.trade_exchange.get_close(inst, end_time) or 1.0
qty = int(self.trade_exchange.account.cash / (len(self.trade_exchange.get_positions()) + 1) / price)
return Order(
inst,
qty,
start_time,
end_time,
direction=OrderDir.BUY if direction > 0 else OrderDir.SELL,
type="market",
)
class BetaNeutralRecord: # subclass qlib.workflow.record_temp.Record in the real impl
"""Custom record that backtests 3L/3S and logs report/trades/risk.csv."""
def __init__(self, *, benchmark: str = "QQQ", n_long: int = 3, n_short: int = 3, **_: Any):
self.benchmark = benchmark
self.n_long = n_long
self.n_short = n_short
def generate(self, **kwargs):
# Real impl: run qlib.backtest with BetaNeutralStrategy on the recorded
# pred, write report_normal.csv / positions_normal.csv / risk.csv into
# the current MLflow run's artifact dir, then log the headline metrics.
print("BetaNeutralRecord.generate: simulate 3L/3S and log artifacts")
@@ -0,0 +1,77 @@
"""Minimal OptimalStopControl strategy — a stub of tac_qlib/contrib/strategy/optimal_stop.py.
Subclasses qlib's BaseSignalStrategy; override `generate_trade_decision` to build
`qlib.backtest.Order`s and return a `TradeDecisionWO`. The real implementation
gates entry by cross-sectional signal percentile, exits by percentile / time /
stop-loss, and sizes equal-weight with `risk_degree` control.
Wire into a workflow YAML under PortAnaRecord.config.strategy:
strategy:
class: OptimalStopControl
module_path: tac_qlib.contrib.strategy.optimal_stop
kwargs:
signal: "<PRED>"
topk: 10
entry_pct: 0.85
exit_pct: 0.7
max_hold_days: 10
min_hold_days: 2
sl: -0.08
risk_degree: 0.95
"""
from __future__ import annotations
from typing import Any, Dict, List, Optional
import numpy as np
from qlib.backtest import Order
from qlib.backtest.decision import OrderDir, TradeDecisionWO
from qlib.contrib.strategy.signal_strategy import BaseSignalStrategy
class OptimalStopControl(BaseSignalStrategy):
def __init__(
self,
*,
topk: int = 10,
entry_pct: float = 0.85,
exit_pct: float = 0.7,
max_hold_days: int = 10,
min_hold_days: int = 2,
sl: float = -0.08,
risk_degree: float = 0.95,
**kwargs: Any,
):
super().__init__(**kwargs)
self.topk = topk
self.entry_pct = entry_pct
self.exit_pct = exit_pct
self.max_hold_days = max_hold_days
self.min_hold_days = min_hold_days
self.sl = sl
self.risk_degree = risk_degree
def generate_trade_decision(self, execute_result=None):
"""Build orders for one trade step (minimal sketch — see repo impl)."""
trade_step = self.trade_calendar.get_trade_step()
# signal is known at t-1 via shift=-1 in the signal object
start_time, end_time = self.trade_calendar.get_step_time(trade_step)
pred_start, pred_end = self.trade_calendar.get_step_time(trade_step - 1)
pred = self.signal.get_signal(start_time=pred_start, end_time=pred_end)
orders: List[Order] = []
if pred is not None and len(pred):
# take the top-k by cross-sectional percentile, equal-weight size
cross = pred.groupby(level=0).rank(pct=True) # 0..1 per day
keep = pred.index[cross >= 1.0 - self.entry_pct]
for inst, (dt, _instr) in zip(keep, keep):
price = self.trade_exchange.get_close(inst, end_time) or 1.0
qty = int((self.risk_degree * self.trade_exchange.account.cash) / (self.topk * price))
if qty > 0:
orders.append(
Order(inst, qty, start_time, end_time, direction=OrderDir.BUY, type="market")
)
return TradeDecisionWO(orders, self)
@@ -0,0 +1,107 @@
# -----------------------------------------------------------------------------
# MINIMAL workflow — the canonical "run a backtest" template for the skill.
#
# Every traced backtest runs through a workflow YAML like this one via
# rd_run_workflow, so the `record` blocks write MLflow artifacts to disk
# (<lake>/mlruns/<exp_id>/<run_id>). The traced experiment's ref id IS the
# mlflow run id returned by rd_run_workflow.
#
# Trigger:
# rd_run_workflow config_path=examples/workflow_minimal.yaml \
# experiment_name=tac-rd-minimal
# -----------------------------------------------------------------------------
{%- set LAKE = TAC_LAKE_DIR %}
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
expression_cache: null
dataset_cache: null
calendar_provider:
class: tac_qlib.data.providers.LakeCalendarProvider
kwargs: { lake_root: "{{ LAKE }}", market: US }
instrument_provider:
class: tac_qlib.data.providers.LakeInstrumentProvider
kwargs: { lake_root: "{{ LAKE }}", market: US, markets: {} }
feature_provider:
class: tac_qlib.data.providers.LakeFeatureProvider
kwargs: { lake_root: "{{ LAKE }}", market: US }
exp_manager:
class: MLflowExpManager
module_path: qlib.workflow.expm
kwargs: { uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd-minimal" }
task:
model:
class: LGBModel
module_path: qlib.contrib.model.gbdt
kwargs:
loss: mse
learning_rate: 0.05
num_leaves: 15
n_estimators: 200
colsample_bytree: 0.8
subsample: 0.8
subsample_freq: 1
reg_alpha: 0.01
reg_lambda: 0.01
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: AAPL,MSFT,QQQ,IVV,SMH,TLT
start_time: 2026-03-01
end_time: 2026-08-06
fit_start_time: 2026-03-01
fit_end_time: 2026-05-31
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-6)/Ref($close,-1)-1"
segments:
train: [2026-03-01, 2026-05-31]
valid: [2026-06-01, 2026-06-30]
test: [2026-07-01, 2026-08-06]
# Record block — REQUIRED. Each entry writes one artifact family to mlruns:
# SignalRecord pred.pkl + label.pkl
# SigAnaRecord IC / Rank IC series + long-short group returns
# PortAnaRecord backtest report / positions / risk
record:
- { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} }
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs: { ana_long_short: true, ann_scaler: 252 }
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config:
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: "<PRED>"
topk: 2
n_drop: 1
only_tradable: true
risk_degree: 0.95
backtest:
start_time: 2026-07-01
end_time: 2026-08-06
account: 1000000
benchmark: QQQ
exchange_kwargs:
codes: AAPL,MSFT,QQQ,IVV,SMH,TLT
deal_price: $close
freq: day
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5.0
risk_analysis_freq: 1d
@@ -0,0 +1,113 @@
# -----------------------------------------------------------------------------
# RankIC early-stop workflow — minimal example wiring the custom model.
#
# model_rank_gbdt.py must be importable: copy it (or symlink) into
# tac_qlib/contrib/model/ and sync to /opt/venv site-packages (see SKILL.md
# "Installed package copy" gotcha). Then run:
#
# rd_run_workflow config_path=examples/workflow_rankic.yaml \
# experiment_name=tac-rd-rankic
# -----------------------------------------------------------------------------
{%- set LAKE = TAC_LAKE_DIR %}
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
expression_cache: null
dataset_cache: null
calendar_provider:
class: tac_qlib.data.providers.LakeCalendarProvider
kwargs: { lake_root: "{{ LAKE }}", market: US }
instrument_provider:
class: tac_qlib.data.providers.LakeInstrumentProvider
kwargs: { lake_root: "{{ LAKE }}", market: US, markets: {} }
feature_provider:
class: tac_qlib.data.providers.LakeFeatureProvider
kwargs: { lake_root: "{{ LAKE }}", market: US }
exp_manager:
class: MLflowExpManager
module_path: qlib.workflow.expm
kwargs: { uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd-rankic" }
task:
model:
# Custom model — see examples/model_rank_gbdt.py (RankICLGBModel):
# per-day query groups + feval=rankic + metric='None' so early-stopping
# tracks mean per-day Spearman instead of l2.
class: RankICLGBModel
module_path: tac_qlib.contrib.model.rank_gbdt
kwargs:
loss: mse
learning_rate: 0.02
num_leaves: 15
num_boost_round: 3000
early_stopping_rounds: 200
min_data_in_leaf: 20
lambda_l1: 0.0
lambda_l2: 0.5
colsample_bytree: 0.8
subsample: 0.8
subsample_freq: 1
seed: 2026
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: AAPL,MSFT,QQQ,IVV,SMH,TLT
start_time: 2026-03-01
end_time: 2026-08-06
fit_start_time: 2026-03-01
fit_end_time: 2026-05-31
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-6)/Ref($close,-1)-1"
infer_processors:
- { class: DropAllNaN, kwargs: {} }
- { class: ProcessInf, kwargs: {} }
- { class: CSRankNorm, kwargs: {} }
- { class: ZScoreNorm, kwargs: {} }
- { class: Fillna, kwargs: {} }
segments:
train: [2026-03-01, 2026-05-31]
valid: [2026-06-01, 2026-06-30]
test: [2026-07-01, 2026-08-06]
record:
- { class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {} }
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs: { ana_long_short: true, ann_scaler: 252 }
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config:
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: "<PRED>"
topk: 2
n_drop: 1
only_tradable: true
risk_degree: 0.95
backtest:
start_time: 2026-07-01
end_time: 2026-08-06
account: 1000000
benchmark: QQQ
exchange_kwargs:
codes: AAPL,MSFT,QQQ,IVV,SMH,TLT
deal_price: $close
freq: day
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5.0
risk_analysis_freq: 1d
+168
View File
@@ -0,0 +1,168 @@
---
name: tradeac-rd-explain
description: Guide agents to retrieve, visualise and interpret TradeAC R&D workflow data — from the input qrun YAML to final IC / backtest metrics — via the tac-qlib-rd MCP tools (rd_exp_*) and the built-in R&D dashboard (/dashboard/rd). Use when asked about experiments, mlruns runs, workflow inputs, model hyper-parameters, IC/Rank IC evaluation, or backtest results.
---
# tradeac-rd-explain
Every `qrun` workflow run is recorded into **mlflow** in the unified R&D store under the lake
root — sqlite `mlruns.db` + artifact files under `mlruns/<experiment_id>/<run_uuid>/` in
`$TAC_LAKE_DIR`. This skill tells you how to pull that data out with the
`rd_exp_*` MCP tools (from `tac_qlib.rd_server`), how to read the raw files directly, and how to
visualise/interpret everything — either from the built-in TradeAC UI or from the raw data.
Quick map of the R&D data:
| Step | Where it lives | `rd_exp_*` tool |
|------|----------------|-----------------|
| Input config (rendered YAML) | artifact `config` on the run | `rd_exp_input` |
| Runs / experiments list | `mlruns.db` → `experiments`, `runs`, `tags`, `params`, `metrics` | `rd_exp_list`, `rd_exp_get_experiment`, `rd_exp_get_run` |
| Predictions & labels | artifacts `pred.pkl`, `label.pkl` | `rd_exp_result` |
| IC / Rank IC | artifacts `ic.pkl`, `ric.pkl` + metrics `IC`, `ICIR`, `Rank IC`, `Rank ICIR` | `rd_exp_result` |
| Group returns | artifacts `long_short_r.pkl`, `long_avg_r.pkl` | `rd_exp_result` |
| Backtest / risk | artifacts `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl` + `1day.*` metrics | `rd_exp_result` |
| Model + hyper-params | artifact `params.pkl` (qlib model), config `task.model` | `rd_exp_model` |
| Hypothesis / evaluation notes | sidecar `rd-notes.json` | `rd_exp_get_notes` / `rd_exp_set_notes` |
## MCP-first policy
- **Use the `rd_exp_*` MCP tools to read all of the above** — do not reinvent them with `sqlite3`/pickle/pandas scripts. The tools are the canonical, JSON-safe way to pull experiment data (they fall back to the raw files automatically).
- **NEVER script directly against the MCP server** (spawning `tac_qlib.rd_server`, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- The raw-file/sqlite snippets in §3 below are **only** for cases where the MCP surface is unavailable or the user explicitly asks for a direct peek.
- If the venv is missing a runtime dep (`duckdb`, `pyarrow`, `sqlite3`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) rather than working around it.
## 1. Prerequisites
- The tac-qlib-rd MCP server is registered in `opencode.json` (`.venv/bin/python -m tac_qlib.rd_server`).
- The server resolves the unified R&D store from the lake root, so `uri` defaults to `sqlite:///<lake>/mlruns.db` (overridable via `MLRUNS_URI`).
- Runs must exist first: use `rd_run_workflow` (or `rd_train` + records) to create them.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`) in scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## 2. Getting the data
### 2.1 List experiments and their runs
```text
rd_exp_list
# -> [{experiment_id, name, run_count, latest_run: {run_id, status, headline_metrics}}]
rd_exp_get_experiment experiment_id=1
# -> experiment meta + every run: run_id, status, start/end, git, metrics, params, tags, notes, artifacts
```
### 2.2 Input configuration (what went in)
```text
rd_exp_input experiment_id=1 run_id=<run_uuid>
```
Returns the **saved `config` artifact** (the fully-rendered workflow YAML: `qlib_init`,
`task.model.kwargs` hyper-parameters, `task.dataset.kwargs.handler` universe/window/features,
`segments`, `record` list) plus the resolved `universe` and `feature_fields`. If a run has no
`config` artifact (e.g. older `rd_train` runs) the tool falls back to reconstructing from
recorded params/tags and marks `source: "reconstructed"` / `"partial"`.
> Rule of thumb: **the `config` artifact is the most complete input record**; the sqlite
> `params` table alone (only `cmd-sys.argv`) is not enough to reconstruct the input.
### 2.3 Results & evaluation (what came out)
```text
rd_exp_result experiment_id=1 run_id=<run_uuid>
```
Returns: headline `metrics` (IC / ICIR / Rank IC / Rank ICIR, `l2.train`/`l2.valid`, `1day.*`
risk metrics), per-day `ic_series` (`[{date, ic, ric}]`), `pred_stats`, `group_returns`
(`long_short` / `long_avg`), and the `backtest` report (per-day cumulative return vs benchmark)
+ `risk` table.
### 2.4 Model & hyper-parameters
```text
rd_exp_model experiment_id=1 run_id=<run_uuid>
rd_exp_model experiment_id=1 run_id=<run_uuid> tree_id=7
```
Returns `hyperparams` (from config, preferred), `feature_names`, `feature_importances`,
`num_trees`, `best_iteration`, and a **pruned top-layers tree** for LightGBM:
`tree: {nodes: [{id, depth, feature, threshold, gain, leaf_value, node_count, left, right}]}`.
### 2.5 Notes (hypothesis / evaluation)
```text
rd_exp_get_notes experiment_id=1 run_id=<run_uuid>
rd_exp_set_notes experiment_id=1 run_id=<run_uuid> hypothesis="..." evaluation="..."
# persisted to mlruns/<experiment_id>/<run_uuid>/rd-notes.json
```
## 3. Reading the raw files directly
Everything above is a JSON view of these files (all under `$TAC_LAKE_DIR`):
- `mlruns.db` (sqlite) — `experiments`, `runs`, `tags`, `params`, `metrics`, `latest_metrics`.
Quick peek: `sqlite3 $TAC_LAKE_DIR/mlruns.db "SELECT * FROM latest_metrics;"`.
- `mlruns/<experiment_id>/<run_uuid>/artifacts/` — pickle files:
- `config` → the input YAML (dict); carries the resolved `feature_fields`
- `params.pkl` → the trained model (qlib `LGBModel`; `.model` is a `lightgbm.Booster`)
- `pred.pkl`, `label.pkl`, `ic.pkl`, `ric.pkl`, `long_short_r.pkl`, `long_avg_r.pkl`
- `portfolio_analysis/report_normal_1d.pkl`, `port_analysis_1d.pkl`
- `mlruns/<experiment_id>/<run_uuid>/rd-notes.json` — hypothesis/evaluation notes.
In Python:
```python
import os
import pickle
from pathlib import Path
run_dir = Path(os.environ["TAC_LAKE_DIR"]) / "mlruns/1/<run_uuid>"
cfg = pickle.loads((run_dir / "artifacts/config").read_bytes()) # input config dict
ic = pickle.loads((run_dir / "artifacts/ic.pkl").read_bytes()) # per-day IC Series
import lightgbm
model = pickle.loads((run_dir / "artifacts/params.pkl").read_bytes()) # needs qlib import
tree = model.model.dump_model()["tree_info"] # LightGBM trees
```
## 4. Visualising & interpreting
### 4.1 Built-in TradeAC UI
Open the dashboard: `/dashboard/rd` lists experiments + runs with headline metrics and
`Input` / `Result` / `Model` action buttons:
- `/dashboard/rd/input?expId=<id>` — universe, windows, features, model settings (tables).
- `/dashboard/rd/result?expId=<id>` — ECharts IC/Rank IC, cumulative group returns, backtest vs
benchmark, per-day IC table, training-loss curves.
- `/dashboard/rd/model?expId=<id>` — hyper-parameter table, feature importances, LightGBM tree
viewer (pick a tree id).
### 4.2 Interpreting the numbers
- **IC / ICIR**: mean per-day IC (predictive power of the signal); ICIR = mean/std × √252.
|IC| ≥ ~0.02 daily with stable sign is notable for cross-sectional signals; ICIR ≥ 1 is decent,
≥ 2 strong. Rank IC is the Spearman version (more robust to outliers).
- **Training loss (`l2.train`/`l2.valid`)**: watch the gap — widening gap ⇒ overfitting;
valid flat/rising ⇒ underfitting or stale features.
- **Group returns (`long_short_r`)**: cumulative return of top-decile-minus-bottom-decile signal
baskets; steady positive slope = the ranking carries money.
- **Backtest risk** (`annualized_return`, `information_ratio`, `max_drawdown`): IR = excess
return / tracking error; max drawdown shows path risk. Compare against the benchmark column
in the cumulative chart.
- **Tree viewer**: root splits on the strongest features (high gain). Repeated use of a feature
across the top layers ⇒ it dominates; suspicious thresholds near feature extremes often
indicate leakage/sample bias.
## 5. Troubleshooting
| Symptom | Cause / fix |
|---------|-------------|
| `experiment_id` not found | Check `rd_exp_list`; ids are the mlflow `experiment_id`, not the name. |
| `no recorder` / empty input | Run lacks a `config` artifact (pre-fix `rd_train`). Re-run via `rd_run_workflow` or `rd_train` on the fixed server to record config. |
| Pickle errors on `params.pkl` | Ensure qlib + lightgbm importable (server venv). Tool returns a warning and skips the artifact rather than failing. |
| Empty result series | Records were not run (only `rd_train`). Use `rd_run_workflow` or add `SignalRecord`/`SigAnaRecord`/`PortAnaRecord`. |
+296
View File
@@ -0,0 +1,296 @@
---
name: tradeac-rd
description: Guide agents to run quant R&D on the TradeAC data lake with a Qlib-based research server exposed over MCP (tac-qlib-rd). Train GBDT (LightGBM/XGBoost) and Linear/QDA ML models on lake bars + TA features, generate cross-sectional alpha predictions, evaluate IC/Rank IC, run TopkDropout backtests with benchmark comparison, and execute one-shot YAML workflows — all through `tac_qlib.rd_server`, an MCP server in the repo venv.
---
# tradeac-rd
Quant R&D server for the TradeAC data lake. Wraps [Qlib](https://github.com/microsoft/qlib) in a local **MCP server** (`tac_qlib.rd_server` in the repo `.venv`) and uses custom qlib data providers that read directly from the lake (see `tac-engine/skills/tradeac-lake/SKILL.md` for the lake itself, and `tac-qlib/README.md` for the package).
Registered in `opencode.json` as `tac-qlib-rd` — the tools below are available directly once opencode is restarted.
## MCP-first policy
- **Prefer the tac-qlib-rd MCP tools** (`rd_train`, `rd_predict`, `rd_evaluate`, `rd_backtest`, `rd_strategy_targets`, `rd_run_workflow`, `rd_exp_*`, `rd_status`) over writing scripts that reimplement the R&D loop (custom qlib glue, own train/predict/eval/backtest, hand-rolled mlruns readers, own JSON-RPC clients).
- **NEVER script directly against the MCP server** (spawning `python -m tac_qlib.rd_server`, driving it via bash/curl/stdio) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- Data prep (bars/features backfill) is done with the tac-engine lake MCP tools — see `tac-engine/skills/tradeac-lake/SKILL.md`. Inspect runs with `rd_exp_*` instead of reading `mlruns.db`/pickles directly.
- If the venv is missing a runtime dep (e.g. `duckdb`, `pyarrow`), lazy-install it (`uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow`) instead of switching to another tool.
## The R&D loop
| Tool | Purpose |
|------|---------|
| `rd_train` | Fit a model on lake data + TA features, log to MLflow, return run metadata. |
| `rd_predict` | Generate out-of-sample predictions from a trained model (by `model_path` or `run_id`). |
| `rd_evaluate` | IC / Rank IC stats of a `pred.pkl` vs `label.pkl`. |
| `rd_backtest` | TopkDropout backtest of predictions vs a benchmark, with risk metrics + artifacts. |
| `rd_strategy_targets` | Turn a prediction's signal day into a deterministic target buy list (TopkDropout selection + sizing). |
| `rd_run_workflow` | One-shot: run an entire YAML workflow (train → predict → sig-ana → backtest) and return metrics + artifacts. |
| `rd_exp_list` | List MLflow experiments with run ids on the local sqlite store. |
| `rd_exp_get_experiment` | Experiment detail: all runs (meta, metrics, notes, artifact files). |
| `rd_exp_get_run` | Single run meta + latest metrics. |
| `rd_exp_input` | What went into a run: qlib_init, model kwargs, dataset handler kwargs, segments, features, universe, label. |
| `rd_exp_result` | What came out: headline IC/ICIR/Rank IC/Rank ICIR, per-day IC series, group returns, prediction stats, backtest report + risk (benchmark-relative). |
| `rd_exp_model` | Trained model: class, hyperparameters, LightGBM tree/feature importance. |
| `rd_exp_blotter` | Execution log: account P&L summary, daily equity, current positions, trade table, signal blotter. |
| `rd_exp_get_notes` / `rd_exp_set_notes` | Read / write hypothesis + evaluation notes on a run. |
| `rd_exp_delete` | **HARD-delete** an MLflow experiment: all runs (metrics/params/tags), the traced `rd_experiments` rows that reference them (FK is `ON DELETE CASCADE`), and on-disk artifacts. Irreversible — confirm with the user first. |
| `rd_exp_delete_run` | **HARD-delete** a single MLflow run + its traced `rd_experiments` row + artifacts. Irreversible — confirm with the user first. |
| `rd_status` | Lake + qlib readiness: data window, symbols, persisted features, qlib version. |
The standard flow is `rd_train` → `rd_predict` → `rd_evaluate` → `rd_backtest`; `rd_run_workflow` replaces all of it with a YAML config. The `rd_exp_*` inspection tools read saved MLflow artifacts, so every page in the R&D app (`/rd/input`, `/rd/result`, `/rd/model`, `/rd/blotter`) is backed by an MCP call (`rd_exp_input`, `rd_exp_result`, `rd_exp_model`, `rd_exp_blotter`) keyed by `experiment_id` + `run_id`.
## Setup / env
| Var | Default | Purpose |
|-----|---------|---------|
| `TAC_LAKE_DIR` | **required** (no default) | lake root (bars + features + metadata). Local dev: absolute path (e.g. `/home/data/lake`). |
| `TAC_RD_MARKET` | `US` | market partition for lake reads |
| `DATABASE_URL` | – | Postgres tracking store for MLflow (its own `experiments`/`runs`/… tables) when set |
| `MLRUNS_URI` | postgres (`$DATABASE_URL`) or `sqlite:///<lake>/mlruns.db` | MLflow tracking URI override. Artifact files always live under `<lake>/mlruns/<exp_id>/<run_uuid>/` |
The server lives in the repo `.venv`; the MCP config is already registered. Restart opencode after editing `opencode.json`.
## Secrets policy
- NEVER write secrets into files: DB passwords, API keys, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `MLRUNS_URI`, `EMBEDDING_API_KEY`) in workflow YAMLs, scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- Tracking store: use `uri: "sqlite:///mlruns.db"` (relative) in workflows — `rd_run_workflow` normalizes it to Postgres when `$DATABASE_URL` is set, else the lake sqlite. Never hardcode a `postgres://user:pass@…` URI.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## `rd_train`
Args: `universe` (comma-separated), `train_start/valid_end/test_end` (`YYYY-MM-DD`), `experiment_name`, `out_dir`, `model` (`lgb` default, `xgb`, `linear`, `qda`), optional `label` (default `Ref($close,-2)/Ref($close,-1)-1`, the next-day return), `topk`/`n_drop` for later backtests.
- Loads 1d bars + all persisted TA features (`features/` dir) for the universe from the lake.
- Splits into train / valid / test; fits on train with early stopping on valid.
- Logs the run to MLflow (`run_id`), saves `params.pkl` (model) + `pred.pkl` + `label.pkl` to `out_dir`.
- With `record_analysis=true` (default) also runs SignalRecord / SigAnaRecord (`ana_long_short`) / PortAnaRecord inside the run, so the result page gets IC/Rank IC series, long-short group returns, monthly IC and the portfolio backtest. `benchmark`, `topk`, `n_drop`, `account`, `risk_degree`, `open_cost`/`close_cost`/`min_cost` tune that backtest.
- Returns `run_id`, `status`, `fit_seconds`, `feature_fields`, `label`, per-segment rows/date/instrument counts, and artifact paths.
- **`wait=false`** runs the fit in a background thread and returns immediately (`status: started`, `background: true`) — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until its status is `FINISHED`, then use that `run_id`. Use it for slow windows (e.g. a 4y retrain) where a blocking MCP call can time out.
> If the lake lacks bars or features for `universe`, backfill first via the tac-engine `get_lake_bars` / `get_lake_ta` tools, or raise the training start date.
## `rd_predict`
Args: `universe`, `model_path` **or** `run_id`+`experiment_name` (artifact `params.pkl` is loaded from MLflow), same date ranges as `rd_train`, `out_dir`, optional `top` (number of top-scored rows in the `head` list).
- Rebuilds the same feature matrix for `test_start..test_end`, produces scores.
- Writes `pred.pkl` (scores) and `label.pkl` (labels) to `out_dir`.
- Returns paths, `count`, `date_min/max`, instruments, score distribution stats, and a small `head`.
## `rd_evaluate`
Args: `pred_path`, `label_path` (the two pkl files from `rd_train`/`rd_predict`).
- Returns `IC` and `RankIC` tables (`days`, `mean`, `std`, `ann_vol`, `ir`, `skew`, `kurt`, `maxdd`) and a `headline` (`IC`, `ICIR`, `Rank IC`, `Rank ICIR`).
## `rd_backtest`
Args: `pred_path`, `start_time`/`end_time`, `topk`, `n_drop`, `benchmark`, optional `out_dir`.
- TopkDropoutStrategy (topk long, n_drop drop), $100k account, `risk_degree 0.95`, day freq, benchmark comparison.
- Returns `start_time`, `end_time`, `trading_days`, `risk` (`mean`, `std`, `annualized_return`, `information_ratio`, `max_drawdown`), `benchmark`, and artifact paths (`report_normal.csv`, `positions_normal.csv`, `risk.csv`).
## `rd_strategy_targets`
Args: `pred_path`, optional `signal_date` (defaults to the last day in the prediction), `account`, `risk_degree`, `topk`, `n_drop`, optional `prices` (JSON `{symbol: price}`).
- Applies the **exact TopkDropout selection** for one signal day: rank the cross-sectional scores, drop the top `n_drop`, take the next `topk` as buys, sized at `account × risk_degree / topk` per name. Use this to chain a prediction straight into an order list — no manual strategy replication.
- With `prices`, floors each order to whole shares (`qty`) and reports `expected_price` / `invested`.
- Returns `signal_date`, `per_name_notional`, the deterministic `targets` list (`symbol`, `rank`, `score`, `side`, `notional`, `qty`), and the top-20 `ranking` for context. If fewer than `topk + n_drop` names have a score that day it returns empty `targets` with a `reason`.
## `rd_run_workflow`
Args: `config_path` (YAML, see `tac-qlib/workflows/workflow_lgb_taclake.yaml`), `experiment_name`, optional `wait` (default `false`), optional `run_in_new_process` (default `false`).
- Runs the full pipeline (qlib `signal` + `records`), returns `run_id`, `status`, the resolved `qlib_init`/`model`/`dataset`/`records` config, and `metrics` (train/valid loss, IC/ICIR/Rank IC/Rank ICIR, and the `1day.*` backtest metrics).
- `wait=false` (default) returns immediately with `status: started`; the workflow runs in a background thread — poll `rd_exp_get_run` / `rd_exp_list` for the newest run of `experiment_name` until it finishes. `wait=true` blocks until completion (only for small windows that finish inside the MCP call timeout).
- `run_in_new_process=true` runs the workflow in a **separate OS process** instead of a thread. qlib `init` sets process-global state, so this is the safe mode for concurrent or long workflows — it isolates crashes, releases memory on exit, and avoids the thread-safety race. stdout/stderr are redirected to `<lake>/logs/rd-workflow-<exp>-<ts>.log` (returned as `log_path`; the child must never write to the MCP stdio pipe). Polling works identically because the child writes to the same mlflow store. The process `pid` is returned.
## Tracing every run started from a chat (REQUIRED)
**Every experiment you start from this chat must be traced FIRST.** The R&D
lineage (`/rd/lineage`) and the round book build on the `rd_experiments` table —
an experiment created by `rd_run_workflow` / `rd_train` without a
`rd_trace_start` is invisible there (no lineage node, no chat link). So before
triggering any run, use the `rd_trace_*` MCP tools (tac-qlib-rd):
1. Open the trace BEFORE the run (see `tac-qlib/skills/tac-qlib-custom/SKILL.md`,
"Experiment traceability" — the skill that owns the trace flow):
```
rd_trace_start rational="<what this run tests, in one line>" \
details="<universe / features / label / model / strategy sizing>" \
experiment_name=<the experiment you will run into> \
evolved_from=<predecessor traced id or auto> \
session_id="<this chat's opencode session id>"
# -> {"experiment_id": N, "branch": "...", "evolved_from": ..., "base_branch": ...}
```
2. Run the workflow into that **same** `experiment_name`:
```
rd_run_workflow config_path=<yaml> experiment_name=<the experiment name>
```
3. On success, **finish the trace** (links the run, copies metrics/evaluation):
```
rd_trace_finish experiment_id=<N> ref_id=<run_id> \
evaluation="<outcome>" metrics='{...headline...}' \
mlruns_dir=<lake>/mlruns/<experiment_id>/<run_id>
```
If you are NOT tracing (quick throwaway exploration), say so explicitly and note
the run will not appear in the lineage graph. The default for any run started
from a chat is to trace it.
## Building a workflow YAML and triggering a run
A workflow YAML is a qrun config: `qlib_init` (lake providers + MLflow exp manager), `task.model`, `task.dataset`, and `task.record`. Copy `tac-qlib/workflows/workflow_lgb_taclake.yaml` as the template.
```yaml
# jinja is available: {%- set LAKE = TAC_LAKE_DIR %} (TAC_LAKE_DIR is required)
qlib_init:
provider_uri: "{{ LAKE }}"
region: us
expression_cache: null
dataset_cache: null
calendar_provider: {class: tac_qlib.data.providers.LakeCalendarProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
instrument_provider: {class: tac_qlib.data.providers.LakeInstrumentProvider, kwargs: {lake_root: "{{ LAKE }}", market: US, markets: {}}}
feature_provider: {class: tac_qlib.data.providers.LakeFeatureProvider, kwargs: {lake_root: "{{ LAKE }}", market: US}}
exp_manager: {class: MLflowExpManager, module_path: qlib.workflow.expm, kwargs: {uri: "sqlite:///mlruns.db", default_exp_name: "tac-rd"}}
task:
model:
class: LGBModel
module_path: qlib.contrib.model.gbdt
kwargs: {loss: mse, learning_rate: 0.05, num_leaves: 15, n_estimators: 200,
colsample_bytree: 0.8, subsample: 0.8, subsample_freq: 1,
reg_alpha: 0.01, reg_lambda: 0.01}
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: TACHandler
module_path: tac_qlib.contrib.data.handler
kwargs:
instruments: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ # universe
start_time: 2000-01-03 # lake look-back for features
end_time: 2026-08-06
fit_start_time: 2026-03-01 # normalization fit window
fit_end_time: 2026-05-31
freq: day
lake_root: "{{ LAKE }}"
market: US
label: "Ref($close,-2)/Ref($close,-1)-1" # next-day return
segments: # train/valid/test split
train: [2026-03-01, 2026-05-31]
valid: [2026-06-01, 2026-06-30]
test: [2026-07-01, 2026-08-06]
record:
- {class: SignalRecord, module_path: qlib.workflow.record_temp, kwargs: {}}
- {class: SigAnaRecord, module_path: qlib.workflow.record_temp, kwargs: {ana_long_short: true, ann_scaler: 252}}
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config:
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs: {signal: "<PRED>", topk: 2, n_drop: 1, only_tradable: true, risk_degree: 0.95}
backtest:
start_time: 2026-07-01
end_time: 2026-08-06
account: 1000000
benchmark: QQQ # any symbol in the lake; empty = no benchmark
exchange_kwargs:
codes: AAPL,MSFT,TSLA,QQQ,IVV,SMH,TLT,IBIT,MCHI,AIQ
deal_price: $close
freq: day
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5.0
risk_analysis_freq: 1d
```
Then trigger it (each call = one new run in the named experiment):
```text
rd_run_workflow config_path=tac-qlib/workflows/tune_run1_wider_5d.yaml experiment_name=tac-rd-tune
# -> run_id <uuid>; save it, then inspect via rd_exp_*.
```
The saved `config` artifact (same shape as above) is what `rd_exp_input` returns, so runs are reproducible from their YAML.
## Inspecting a run given experiment_id + run_id
The URL on the R&D app is `/rd/input|result|model|blotter?expId=<id>&run=<uuid>`; the underlying MCP calls are:
| You want | MCP call | Args |
|----------|----------|------|
| Full results (IC/ICIR/Rank IC, group returns, backtest risk) | `rd_exp_result` | `experiment_id`, `run_id` |
| Input config (universe, windows, features, label, model) | `rd_exp_input` | `experiment_id`, `run_id` |
| Execution blotter (P&L, positions, trades, signals) | `rd_exp_blotter` | `experiment_id`, `run_id` |
| Model (hyperparams, tree, importances) | `rd_exp_model` | `experiment_id`, `run_id` |
| Run notes | `rd_exp_get_notes` / `rd_exp_set_notes` | `experiment_id`, `run_id` (+ `hypothesis`/`evaluation`) |
Get the run ids first: `rd_exp_list` → pick an experiment → `rd_exp_get_experiment` returns its runs (meta + latest metrics), or `rd_exp_get_run run_id=<uuid>` for one run.
## Evaluating a run and proposing the next one
Treat each run as one hypothesis. To evaluate and iterate:
1. **Read the input** (`rd_exp_input`): universe, train/valid/test windows, label expression, features, model + hyperparams. Note what was held fixed vs changed.
2. **Read the signal metrics** (`rd_exp_result.headline`): IC (predictive power), ICIR (stability — |ICIR| ≥ 0.5 strong, 0.2–0.5 weak but persistent, < 0.2 noise), Rank IC/Rank ICIR. A decent IC with Rank IC ≈ 0 means the ranking is noisy even if the mean cross-section is predictive.
3. **Read the backtest** (`rd_exp_result.backtest`): `annualized_return`, `information_ratio`, `max_drawdown` are **excess vs the benchmark** (qlib mean-daily × 238). Compare against `return_annualized` (raw strategy) and `benchmark_annualized`; check the benchmark is a sensible peer (a single high-flying stock like AAPL is a brutal benchmark for an ETF universe).
4. **Read the blotter** (`rd_exp_blotter.summary`): `n_trades`/`trading_days` reveal turnover; `total_cost` vs account is the cost drag; positions show concentration. High turnover + low topk on correlated names = cost-heavy, undiversified book.
5. **Diagnose** and pick ONE lever for the next run — change one thing, hold the rest fixed so the comparison is clean:
- *Weak/noisy signal* (ICIR < 0.3, Rank IC ≈ 0): longer label horizon (e.g. 5-day `Ref($close,-6)/Ref($close,-1)-1`), stronger regularization (`reg_alpha`/`reg_lambda` up, `subsample`/`colsample` down), or a cleaner universe (drop leveraged/duplicate names).
- *Good signal, bad book* (high IC but poor excess return): raise `topk` for diversification, tune `n_drop` for rotation, reduce turnover, check `total_cost`.
- *Benchmark mismatch*: pick an index ETF (QQQ/IVV) the universe tracks instead of a single stock.
- *Data window*: a 3-month fit window is short; consider rolling/expanding if the lake history allows.
6. **Write the next run as a YAML** (see section above), **trace it first** (`rd_trace_start experiment_name=<exp>`), then trigger with `rd_run_workflow` into that **same new experiment** (e.g. `tac-rd-tune`), and `rd_trace_finish experiment_id=<N> ref_id=<run_id>` when it succeeds. Then `rd_exp_get_experiment` to compare run-to-run. Record the hypothesis/evaluation via `rd_exp_set_notes`.
Example: the baseline `Exp-1 Run-f29f5446` shows IC 0.071 / ICIR 0.17 / Rank IC 0.014 with excess return −0.94 ann (IR −2.23) vs a +89% ann benchmark — the 1-day signal is unstable, the topk=2 book turned 24 trades in 27 days (~1.1% cost drag) on correlated ETFs + leveraged hedges, and AAPL is an unfair benchmark. Two improvement runs are ready in `tac-qlib/workflows/tune_run1_wider_5d.yaml` (5-day label, topk=5, deduped 10-name universe, benchmark QQQ) and `tune_run2_regularized.yaml` (stronger regularization, topk=3/n_drop=2, same-day label) — trigger both into `experiment_name=tac-rd-tune` and compare.
## Example session
```text
# 1) train
rd_train universe=AAPL,MSFT,TSLA,USO,SLV,TLT train_start=2026-03-01 train_end=2026-05-31
valid_start=2026-06-01 valid_end=2026-06-30 test_start=2026-07-01 test_end=2026-08-06
experiment_name=tac-rd-mcp out_dir=/tmp/rd_out
# -> run_id ...
# 2) predict on the test window (from the mlflow run)
rd_predict universe=AAPL,MSFT,TSLA,USO,SLV,TLT run_id=<run_id> experiment_name=tac-rd-mcp
train_start=2026-03-01 train_end=2026-05-31 valid_start=2026-06-01 valid_end=2026-06-30
test_start=2026-07-01 test_end=2026-08-06 out_dir=/tmp/rd_out top=5
# 3) evaluate the alpha
rd_evaluate pred_path=/tmp/rd_out/pred.pkl label_path=/tmp/rd_out/label.pkl
# 4) backtest the signal
rd_backtest pred_path=/tmp/rd_out/pred.pkl start_time=2026-07-01 end_time=2026-08-06 topk=2 n_drop=1 benchmark=AAPL
# 5) one-shot equivalent — trace first, then run, then finish
rd_trace_start --rational "<hypothesis>" --experiment-name tac-rd-one-shot --evolved-from auto
rd_run_workflow config_path=tac-qlib/workflows/workflow_lgb_taclake.yaml experiment_name=tac-rd-one-shot
rd_trace_finish --id <EXPERIMENT_ID> --ref-id <run_id> --evaluation "<outcome>"
# 6) inspect that run later — given experiment_id + run_id (the /rd pages call exactly these)
rd_exp_get_experiment experiment_id=1 # -> runs with meta + latest metrics
rd_exp_input experiment_id=1 run_id=<run_id> # what went in: universe, windows, features, label, model
rd_exp_result experiment_id=1 run_id=<run_id> # what came out: IC/ICIR/Rank IC, backtest risk
rd_exp_blotter experiment_id=1 run_id=<run_id> # execution: P&L, positions, trades, signals
rd_exp_model experiment_id=1 run_id=<run_id> # hyperparameters + tree / importances
rd_exp_set_notes experiment_id=1 run_id=<run_id> hypothesis="5d label + topk5" evaluation="ICIR 0.5, ann +12%"
```
## Notes
- The server reads the lake lazily via the custom `LakeCalendarProvider` / `LakeInstrumentProvider` / `LakeFeatureProvider`; if data is missing the relevant provider raises a clear error — backfill through the tac-engine lake tools first.
- All tools return JSON via stdio (MCP). Diagnostics/logs go to stderr.
- MLflow runs are stored in the tracking store at `$DATABASE_URL` (Postgres) when
set, else the unified lake sqlite `mlruns.db`; artifact files always live under
`<lake>/mlruns/<exp_id>/<run_uuid>/`. Override the tracking URI with `MLRUNS_URI` if needed.
- `rd_run_workflow` / `rd_train` pin each experiment's MLflow `artifact_location` to `<lake>/mlruns` so DB and artifacts stay co-located even when the server process runs from another cwd. Readers (`rd_exp_*`) resolve each run's artifact dir from its recorded `artifact_uri`, falling back to the side-by-side `<lake>/mlruns` layout — so runs whose artifacts were written elsewhere (e.g. `<cwd>/mlruns`) still display.