tac-qlib
Run stock qlib ML workflows (LightGBM → signal → backtest) directly on the
TradeAC parquet lake. No CSV/bin dump, no data conversion: the lake's calendar, instrument master, OHLCV bars and
pre-computed ta-lib features plug into qlib as first-class data providers, and a custom DataHandlerLP
(TACHandler) exposes them through the normal qlib dataset/processor pipeline.
$ qrun workflows/workflow_lgb_taclake.yaml --experiment_name tac-lake-lgb
trains a LightGBM, records predictions/labels, evaluates the signal (IC/RankIC), runs a daily
TopkDropoutStrategy backtest with cost model and risk analysis, and logs everything to mlflow (sqlite).
Layout
tac-qlib/
├── tac_qlib/
│ ├── qlib_init.py # qlib_init() drop-in wired to the lake providers
│ ├── data/
│ │ ├── config.py # LakeConfig: paths + metadata readers, freq↔timeframe map
│ │ └── providers.py # LakeCalendarProvider / LakeInstrumentProvider / LakeFeatureProvider
│ └── contrib/data/
│ └── handler.py # TACHandler (DataHandlerLP) + DropAllNaN processor
├── workflows/
│ └── workflow_lgb_taclake.yaml # qrun workflow: train -> signal -> backtest
├── examples/
│ └── run_backtest.py # same loop as the workflow, plain Python (no yaml)
└── tests/
└── test_lake_providers.py # plain-assert smoke tests
Requirements / install
- Python 3.12,
qlibnightly (0.1.dev2066in the repo venv), pandas/pyarrow, lightgbm. - The TradeAC lake (see below).
TAC_LAKE_DIRis mandatory (no default) — set it to the lake root, or pass thelake_rootkwargs.
The lake (data layout)
$TAC_LAKE_DIR/
├── market=US/
│ └── timeframe=1d/
│ └── symbol=AAPL.parquet # OHLCV bars: t, date, o, h, l, c, v, n, vw
├── features/
│ └── market=US/timeframe=1d/
│ └── symbol=AAPL.parquet # ta-lib indicators, wide format: t, sma_5, rsi_14, ...
├── calendar.parquet # trading days per market
├── coverage.parquet # per (market,timeframe,symbol) loaded windows
└── symbols.parquet # asset master
Field routing (tac_qlib/data/config.py):
$open $high $low $close $volume $vwap→ bar parquet columns;$amount=v * vw,$avg_amount=vw.$factor $change $trade_unit $suspend_flag→ all-NaN (not stored; the backtest Exchange only needs$close).- anything else (e.g.
$rsi_14,$sma_20) → a ta-lib column in the features parquet.
Step 1 — Prepare data
The lake is populated and backfilled with the tac-engine MCP lake tools (see
tac-engine/skills/tradeac-lake). Typical sequence:
- Seed the trading calendar from historical auctions (so the 1d completeness check has an expected day set):
backfill_lake_calendar(symbols="AAPL,MSFT,..."). - Backfill bars:
get_lake_bars(symbols="AAPL,MSFT,...", timeframe="1d", start="2026-02-09")(lazy: missing windows are fetched from Alpaca and persisted;sip/iexauto-fallback on 403). - Persist features:
get_lake_ta(symbol="AAPL", timeframe="1d", indicators="sma_5,sma_20,rsi_14,macd,bb,atr_14", persist=true). Only indicators that exist in every features file are auto-loaded by the handler; add columns per symbol by re-runningget_lake_ta. get_lake_symbols/get_lake_coverageto verify the universe and loaded windows.
TACHandler discovers the feature columns itself (get_common_feature_fields = the intersection of columns
across all features files), so no config change is needed as the lake grows.
Step 2 — Preprocess
Preprocessing happens in TACHandler (a DataHandlerLP), composed from standard qlib processors:
- infer (
DEFAULT_INFER_PROCESSORS), applied to the input features:DropAllNaN— drops columns that are all-NaN over the fit window (fixes the lake's fully-empty ta-lib columns, e.g. astoch_*output that is NaN from the start). The drop set is fixed infit()and applied identically to train/valid/test so feature columns never diverge.ProcessInf,ZScoreNorm(fit on the fit window),Fillna.
- learn (
DEFAULT_LEARN_PROCESSORS), applied to the label:DropnaLabel,CSZScoreNorm.
Handler kwargs (used by both the workflow yaml and the Python API):
| kwarg | default | meaning |
|---|---|---|
instruments |
all |
universe; list, all, or a named pool from markets: |
start_time / end_time |
– | queried window (must be within the lake calendar) |
fit_start_time / fit_end_time |
start/end | window the fit-able processors (ZScoreNorm, DropAllNaN) fit on |
freq |
day |
maps to the lake timeframe (day→1d, 1min→1m, …) |
feature_fields |
auto | raw OHLCV + common ta-lib columns; or an explicit list |
label |
Ref($close,-2)/Ref($close,-1)-1 |
qlib expression for the target |
lake_root / market |
$TAC_LAKE_DIR / US |
lake location (required) / market partition |
Only daily (1d) is currently supported by the calendar provider; intraday freq raises NotImplementedError.
Step 3 — Train
Either write the model task in yaml and run qrun (see Glue with qrun), or train in Python:
from qlib.data.dataset import DatasetH
from tac_qlib.qlib_init import qlib_init
from tac_qlib.contrib.data.handler import TACHandler
from qlib.contrib.model.gbdt import LGBModel
from qlib.workflow import R
qlib_init(provider_uri=os.environ["TAC_LAKE_DIR"], market="US", freq="day")
handler = TACHandler(
instruments="all",
start_time="2026-03-01", end_time="2026-08-06",
fit_start_time="2026-03-01", fit_end_time="2026-05-31",
freq="day", lake_root=os.environ["TAC_LAKE_DIR"], market="US",
)
dataset = DatasetH(handler=handler, segments={
"train": ("2026-03-01", "2026-05-31"),
"valid": ("2026-06-01", "2026-06-30"),
"test": ("2026-07-01", "2026-08-06"),
})
model = LGBModel(n_estimators=200, learning_rate=0.05, num_leaves=15, ...)
with R.start(experiment_name="tac-lake-demo"):
model.fit(dataset) # trains on the train segment
Step 4 — Test / evaluate the signal
model.predict(dataset) returns the prediction on the test segment (a (datetime, instrument) Series).
Evaluate it with qlib's SigAnaRecord / sig_analysis:
from qlib.workflow.record_temp import SigAnaRecord
from qlib.contrib.evaluate import signal_analysis
pred = model.predict(dataset) # "score" column
label = dataset.prepare("test", col_set="label", data_key=DataHandlerLP.DK_I)["LABEL0"]
# per-day + overall IC / ICIR / RankIC / RankICIR
report = signal_analysis(pred, label)
In the workflow this is automatic (SigAnaRecord): the run logs IC 0.0072 / ICIR 0.016 /
RankIC 0.0138 / RankICIR 0.034 for the default split — weak but the plumbing is verified.
Step 5 — Backtesting
from qlib.contrib.evaluate import backtest_daily, risk_analysis
from qlib.contrib.strategy.signal_strategy import TopkDropoutStrategy
strategy = TopkDropoutStrategy(signal=pred, topk=2, n_drop=1, only_tradable=True, risk_degree=0.95)
report_normal, positions_normal = backtest_daily(
start_time="2026-07-01", end_time="2026-08-06",
strategy=strategy, account=1_000_000, benchmark=None, # lake has no index quotes
exchange_kwargs={"codes": universe, "deal_price": "$close", "freq": "day",
"open_cost": 0.0005, "close_cost": 0.0015, "min_cost": 5.0},
)
risk = risk_analysis(report_normal["return"], freq="day")
TopkDropoutStrategyis the default mapping prediction → positions (hold top-k, dropn_dropper day). For other sizing frameworks — equal/score-weighting, softmax, z-score, fractional Kelly, mean-variance — subclassqlib.contrib.strategy.SignalStrategyand implementgenerate_trade_decision(see../.tmp/signalTrade.mdfor the recipe catalogue).- The Exchange needs
$close; other fields the backtest probes ($factor,$trade_unit) are all-NaN and fine. - Benchmark: pick any symbol the lake holds (e.g.
benchmark: AAPL); null benchmark triggers benign "Mean of empty slice" warnings from the risk analysis.
Step 6 — Predict
SignalRecord already saved pred.pkl (test segment) during the qrun run. For predictions on arbitrary data:
pred = model.predict(dataset) # predict on the "test" segment
pred.to_frame("score").to_pickle("pred.pkl") # (datetime, instrument) x ["score"]
To predict a live/rolling window instead of the configured test segment, point a handler's segments["test"]
at the window of interest, or call model.predict(dataset, segment="test") after overriding the segment.
Glue everything with qrun
workflows/workflow_lgb_taclake.yaml wires the whole chain (init → train → signal record → signal analysis →
backtest + risk analysis) into one qrun invocation:
cd tac-qlib
qrun workflows/workflow_lgb_taclake.yaml --experiment_name tac-lake-lgb
# custom lake root:
TAC_LAKE_DIR=/path/to/lake qrun workflows/workflow_lgb_taclake.yaml --experiment_name tac-lake-lgb
YAML anatomy:
qlib_init— pointsprovider_uriat the lake and installs the lake providers by their full class paths (tac_qlib.data.providers.Lake*Provider), plus anexp_managerbacked bysqlite:///<lake>/mlruns.db(avoids mlflow's filesystem-backend maintenance-mode opt-in). The unified R&D store lives under the lake root:mlruns.db+mlruns/<exp>/<run>/. Override the tracking URI withMLRUNS_URI.task.model—LGBModelhyperparameters.task.dataset—DatasetHoverTACHandler;segments.train/valid/testsplit the window;fit_start_time/fit_end_timepin the processor fit window to train.task.record— ordered records:SignalRecord→ writespred.pkl(andlabel.pkl).SigAnaRecord→sig_analysis/{ic,ric}.pkl(IC/ICIR/RankIC/RankICIR).PortAnaRecord→ dailyTopkDropoutStrategybacktest +risk_analysis_freq: 1d→portfolio_analysis/*.pkl(report, positions, indicators, risk metrics, benchmark & cost-adjusted excess returns).
Run artifacts land under the mlflow run: <lake>/mlruns/<exp>/<run>/artifacts/*.pkl (metadata in <lake>/mlruns.db).
Template notes:
- The header uses jinja2 (
{%- set LAKE = TAC_LAKE_DIR %}) —TAC_LAKE_DIRis required and names the lake root. Do not use-%}on the closing tag — it strips the newline and gluesqlib_init:onto the comment line (YAML parse error). qrunisqlib.cli.run:run(fire): positional CONFIG_PATH +--experiment_name/--uri_folder. No--configflag.
Manual (no-yaml) path
examples/run_backtest.py runs the identical loop in plain Python (good for parametrizing universe, features,
label, topk, costs):
.venv/bin/python tac-qlib/examples/run_backtest.py
.venv/bin/python tac-qlib/examples/run_backtest.py --features '$close,$rsi_14,$sma_5,$macd' \
--universe AAPL,MSFT,TSLA,USO,SLV,TLT --topk 2 --n-drop 1 --output ./backtest_out
Writes pred.pkl, report_normal.csv, positions_normal.csv, risk.csv to the output dir.
Reference
tac_qlib/data/providers.py— the three lake providers; they match qlib's provider interface (feature()keyed by calendar position,list_instruments()with listing spans,load_calendar()), so the expression engine,DatasetHand the backtestExchangework unchanged.tac_qlib/contrib/data/handler.py—TACHandler(DataHandlerLP overQlibDataLoader),DropAllNaN,get_common_feature_fields,discover_feature_fields.tac_qlib/data/config.py—LakeConfigpath/reader helpers,FREQ_TO_TIMEFRAME,BAR_FIELD_MAP,resolve_lake_root($TAC_LAKE_DIR, required — fails fast if unset).- Tests (no pytest; plain asserts):
.venv/bin/python tac-qlib/tests/test_lake_providers.py