28 KiB
name, description
| name | description |
|---|---|
| tradeac-lake | Guide agents to build and query the TradeAC parquet+DuckDB data lake on the local filesystem — hive-partitioned bar store (market/timeframe/symbol) plus symbols, watchlist, calendar, features and coverage metadata — with lazy backfill from the tac-engine MCP get_stock_bars tool (tradeac-alpaca skill). |
tradeac-lake
Local-first market data lake: Apache Parquet files on disk, consumed with DuckDB (or Apache Arrow). Bar data is the core payload; the lake also keeps small metadata parquet files (symbols, watchlist, calendar, features, coverage) at the lake root.
Reading is a cache-first pattern: if the requested range is already in the lake, serve it directly from parquet; otherwise lazy-load the missing window via the tac-engine MCP get_stock_bars tool (see tac-engine/skills/tradeac-alpaca/SKILL.md), persist it, update metadata, then return.
MCP-first policy
- Prefer the tac-engine lake MCP tools (
get_lake_bars,get_lake_ta,get_lake_sp,get_lake_features,get_lake_status,get_lake_coverage,get_lake_calendar,backfill_lake_calendar, …) whenever they cover the need. They handle coverage checks, lazy backfill, feed fallback, metadata updates and pagination for you — do not reimplement that in DuckDB/pyarrow scripts. - Direct parquet reads are only for verification (DuckDB CLI / pyarrow snippets below) or when no lake tool covers the query (e.g. an arbitrary ad-hoc SQL join). Keep hand-rolled lake writes off the happy path — the write path is what the MCP tools automate.
- NEVER script directly against the MCP server (spawning the engine binary, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then stop and ask the user to confirm first.
- The engine bundles its own DuckDB; direct verification only needs the
duckdbCLI or a Python venv withduckdb+pyarrow(see dependencies below).
Env / root
| Var | Default | Purpose |
|---|---|---|
TAC_LAKE_DIR |
required (no default) | lake root on the local filesystem. Local dev: absolute path (e.g. /home/data/lake). |
export TAC_LAKE_DIR=/path/to/lake
mkdir -p "$TAC_LAKE_DIR"
Secrets policy
- NEVER write secrets into files: API keys, DB passwords, OAuth tokens, or credential-bearing URLs (
DATABASE_URL,APCA_*) in scripts, configs, notes or committed code. - NEVER read
*.env/.env.*directly (cat/tail/grep/sed/headon.env). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned
.env) and reference it by name ($VAR), never by value. If it's missing, report which variable is required instead of reading it yourself. - If you find a committed secret, flag it, remove it, and replace it with a placeholder.
Lake layout
Hive partition convention, partitioned by market, timeframe, symbol. Metadata parquet files live alongside the partition dirs at the lake root.
$TAC_LAKE_DIR/
├── market=US/
│ └── timeframe=1d/
│ ├── symbol=AAPL.parquet
│ ├── symbol=MSFT.parquet
│ └── ...
│ └── timeframe=10m/
│ └── symbol=AAPL.parquet
├── market=CRYPTO/... # optional: BTC/USD etc.
├── features/ # TA + SP indicators, hive-partitioned with family tier
│ └── market=US/
│ └── timeframe=1d/
│ ├── family=ta/
│ │ └── symbol=AAPL.parquet # TA indicators (sma, rsi, macd, ...)
│ └── family=sp/
│ └── symbol=AAPL.parquet # Stochastic-process features (ou, hmm, har, ...)
├── symbols.parquet # asset master seen/known to the lake
├── watchlist.parquet # watchlists
├── calendar.parquet # trading days per market (coverage ground truth)
├── coverage.parquet # per (market,timeframe,symbol) loaded window
└── manifest.yaml # lake config: feed, adjustment, timezone
Convention: one parquet file per symbol per timeframe per family under the partition dirs. Bar upserts merge by canonical timestamp (read existing file → overlay new bars → write the full set atomically); feature persistence is a fresh write per family (Appender-based, no merge with existing).
Lake MCP tools (tac-engine)
The tac-engine MCP server exposes a lake tools category that wraps the read/write paths below. symbols/market are uppercased, timeframe is normalized to lake spelling, and start/end accept YYYY-MM-DD or RFC-3339 (default end=now, start=end-30d).
| Tool | Purpose |
|---|---|
get_lake_bars |
Cache-first bars: {market?, symbols, timeframe, start?, end?, feed?, adjustment?, lazy?, quiet?}. Feed defaults to iex; SIP is never used (requires license). For daily bars, Yahoo Finance fills gaps before Alpaca's earliest available date. lazy=true (default) backfills missing windows via Alpaca and persists; lazy=false reads the lake only. Returns {request, source, bars: {SYM: [{t,o,h,l,c,v,n,vw}]}}; source is lake (complete hit), partial (present but missing windows and lazy=false), or fetched (gaps backfilled). With quiet=true returns {request, source, summary: {SYM: {count, first_t, last_t}}} instead of the bar rows — use for backfill-to-lake jobs. Bar writes merge by timestamp (read existing file, overlay new bars, write the full set atomically) — safe for both tail appends and leading-gap backfills. Coverage/symbols/calendar metadata are reconciled against the actual file on each write. |
get_lake_ta |
Compute + optionally persist indicators: {market?, symbol, timeframe, start?, end?, indicators?, persist?, quiet?}. indicators is comma-separated, default all: sma_5,sma_20,ema_12,ema_26,rsi_14,macd,bb,atr_14,adx_14. Lookback is pulled automatically; returned rows cover [start,end]. When persist=true, features are written to features/market=*/timeframe=*/family=ta/symbol=*.parquet. With quiet=true returns {count, columns, persisted} instead of the feature rows — use when the goal is persisting indicators. |
get_lake_sp |
Compute + optionally persist stochastic-process features (sp_* columns) from lake bars: {market?, symbol, timeframe, start?, end?, fit_end?, families?, persist?, quiet?}. Rust port of sp_features.py on the stochastic-rs stack. families is comma-separated, default all: ou,hmm,jump,har,trend,hurst,signature,moments. In addition to sp_rv*/sp_vol_ratio_* the har family also emits sp_rv_ac1 (RV lag-1 autocorr) + sp_rv_cv_22 (RV coefficient of variation); jump also emits sp_max_up/sp_max_down (signed max-move asymmetry); signature also emits the lag-5 level-2 cross terms sp_sig_level2_{lead_lag,lag_lead}_5; and moments emits the scale-free realized skewness/kurtosis (sp_rskew_{5,22}, sp_rkurt_{5,22}; a 1-day window is undefined) and downside semi-variance (sp_dsv_{1,5,22}, sp_dsv_ratio_{1,5,22}) via stochastic-rs realized. fit_end limits the 2-state Gaussian-HMM fit window (no lookahead; posteriors still cover the whole window). On persist=true, features are written to features/market=*/timeframe=*/family=sp/symbol=*.parquet. With quiet=true returns {count, sp_columns, persisted} instead of the feature rows. Deferred (not in stochastic-rs): garch, entropy, catch22. |
get_lake_features |
Read persisted TA + SP features (hive-partitioned features/ dir, families ta and sp merged by timestamp): {market?, symbol?, timeframe?, start?, end?, quiet?}. With quiet=true returns {count, columns} instead of the full feature rows. |
get_lake_symbols |
Read symbols.parquet asset master; optional {symbol?} filter. |
get_lake_watchlist |
Read watchlist.parquet. |
get_lake_calendar |
Read calendar.parquet trading days: {market?, start?, end?}. |
get_lake_coverage |
Read coverage.parquet cache index: {market?, timeframe?, symbol?}. |
get_lake_status |
Lake root, manifest.yaml, bar partition inventory and metadata file sizes. |
rebuild_lake_symbol |
Delete bar + feature parquet files and re-fetch from TAC_LAKE_START_DATE (default 2000-01-03) for a single symbol: {market?, symbol, timeframe, feed?, adjustment?}. Resets coverage so the next get_lake_bars call re-downloads the full history. Use after changing TAC_LAKE_START_DATE or to fix stale/corrupt data. |
load_lake_symbols |
Bulk-load + persist bars + TA + SP features for a comma-separated list of symbols. Runs in a background thread and returns immediately with a job_id: {market?, symbols, timeframe, start?, end?, feed?, adjustment?, indicators?, families?}. Per-symbol start is computed automatically from lake coverage: if the lake has no data or first_t > TAC_LAKE_START_DATE, fetches from TAC_LAKE_START_DATE (default 2000-01-03); if first_t <= TAC_LAKE_START_DATE, fetches only from last_t (tail refresh). TA/SP features are always computed over the full TAC_LAKE_START_DATE to end range. Poll load_lake_status with the returned job_id to track progress. |
load_lake_status |
Query the status of a background bulk-load job: {job_id}. Returns {job_id, status, total_symbols, processed, results, error, started_at, completed_at} where status is running, completed, or failed, and results contains per-symbol bar counts, TA/SP column counts, and any errors. |
backfill_lake_calendar |
Gap-fill tool: seed/enrich calendar.parquet from Alpaca historical auctions (feed=iex; records exist only on trading days): {market?, symbols, start?, end?}. Returns {market, symbols, start, end, calendar_days_added}. Call this before lazy bar loads so the 1d completeness check knows the expected trading-day set. |
validate_lake_dataset |
Pre-workflow quality gate: {market?, timeframe?, symbols?, start?, end?}. Scans every symbol in coverage (or a comma-separated symbols subset) and reports verdict: OK/WARNINGS/ERRORS plus per-symbol issues. Catches the failure modes qlib silently tolerates: all-NaN feature columns (would be dropped by DropAllNaN — the model trains on fewer features without notice), missing TA/SP feature files, hollow coverage / stale date ranges (coverage claims a wide span but the bar file is empty/truncated/sparse), stale coverage (first/last/num_bars vs the actual file), and partition misalignment (flat-layout feature orphans the family=ta |
Example:
{"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true}
→ {"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "bars": {"AAPL": [{...}], "MSFT": [{...}]}}
Quiet mode
get_lake_bars, get_lake_ta, get_lake_sp and get_lake_features accept "quiet": true. When the point of the call is writing to the lake (backfill/fetch bars, compute + persist indicators or sp_* features), use quiet: true — the tool still performs the full backfill / computation / persist, but returns a summary instead of echoing back the potentially huge payload (thousands of bar rows / feature rows). Full-row output (bars / features) is the default, so requests that need the data to read it must leave quiet unset/false.
| Tool | quiet: true response |
|---|---|
get_lake_bars |
{request, source: {SYM: lake|partial|fetched}, summary: {SYM: {count, first_t, last_t}}} |
get_lake_ta |
{market, symbol, timeframe, start, end, count, columns, persisted} |
get_lake_sp |
{market, symbol, timeframe, start, end, fit_end, count, sp_columns, persisted} |
get_lake_features |
{count, columns} |
Backfill-to-lake job (no payload echoed):
{"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true, "quiet": true}
→ {"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "summary": {"AAPL": {"count": 44, "first_t": "2026-06-06T04:00:00Z", "last_t": "2026-08-05T04:00:00Z"}, "MSFT": {...}}}
Persist indicators to the lake (summary only):
{"symbol": "AAPL", "timeframe": "1d", "indicators": "sma_5,sma_20,rsi_14", "persist": true, "quiet": true}
→ {"market": "US", "symbol": "AAPL", "timeframe": "1d", "count": 44, "columns": ["sma_5","sma_20","rsi_14"], "persisted": true}
{"symbols": "AAPL,MSFT", "start": "2026-06-06"}
→ {"market": "US", "symbols": ["AAPL","MSFT"], "start": ..., "end": ..., "calendar_days_added": 44}
Conventions
market:US(equities),CRYPTO,FOREX. Uppercase.timeframe: normalized lake name — lowercase,1m 5m 10m 15m 30m 1h 2h 4h 1d 1w 1M. The MCP tool spells them differently; always map:Lake MCP timeframeLake MCP timeframe1m1Min2h2Hour5m5Min4h4Hour10m10Min1d1Day15m15Min1w1Week30m30Min1M1Month1h1Hoursymbol: uppercase, e.g.AAPL. Hyphens/.in special symbols (e.g.BRK-B,SPY) are valid filenames; avoid/and spaces.- All timestamps stored as UTC instants (
TIMESTAMPTZ). Alpaca returns RFC-3339 UTC; normalize on write. 1dbars:tis the session date at04:00Z(midnight ET — Alpaca stamps daily bars at04:00:00Z); also store adatecolumn (CAST(t AS DATE), UTC) for calendar joins. A date-onlyend(e.g.2026-08-05) is treated as inclusive of the whole end day, so the end-day bar is not dropped.
Bar parquet schema (market=…/timeframe=…/symbol=….parquet)
| col | type | source field |
|---|---|---|
t |
TIMESTAMPTZ | bar t (UTC) |
o |
DOUBLE | o |
h |
DOUBLE | h |
l |
DOUBLE | l |
c |
DOUBLE | c |
v |
BIGINT | v |
n |
BIGINT | n |
vw |
DOUBLE | vw |
Partition columns market/timeframe/symbol are derived from the path; DuckDB exposes them automatically when reading a hive glob.
Metadata parquet files
All written with DuckDB COPY … (FORMAT PARQUET) from in-memory SELECT, or pyarrow.parquet.
symbols.parquet
| col | type | notes |
|---|---|---|
symbol |
VARCHAR (pk) | |
name |
VARCHAR | |
asset_class |
VARCHAR | |
exchange |
VARCHAR | |
tradable |
BOOLEAN | |
status |
VARCHAR | |
first_seen |
TIMESTAMPTZ | the symbol's earliest bar in the lake (its first trading date), not the load timestamp |
updated_at |
TIMESTAMPTZ |
watchlist.parquet
| col | type |
|---|---|
watchlist_id |
VARCHAR |
name |
VARCHAR |
symbol |
VARCHAR |
added_at |
TIMESTAMPTZ |
updated_at |
TIMESTAMPTZ |
calendar.parquet — the trading-day ground truth per market (see “Calendar gap” below). Bars only seed which dates are trading days; per-symbol prices/session times are NOT attributed by the bars path (no symbol column, 1d bars all share t=04:00).
| col | type | notes |
|---|---|---|
market |
VARCHAR | pk + date |
date |
DATE | a trading day (UTC) |
session_open |
TIMESTAMPTZ | from auctions o[0].t only (nullable; not set by bars) |
session_close |
TIMESTAMPTZ | from auctions c[0].t only (nullable; not set by bars) |
open_price |
DOUBLE | opening auction price (nullable) |
close_price |
DOUBLE | closing auction price (nullable) |
source |
VARCHAR | auctions | bars | manual |
updated_at |
TIMESTAMPTZ |
features/ — TA + stochastic-process indicators, wide format, hive-partitioned with a family tier: features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet and family=sp/symbol=AAPL.parquet. Each row is one t, with one column per indicator. The partition columns (market/symbol/timeframe) come from the directory structure; the file itself stores t + indicator columns (e.g. sma_5, sma_20, ema_12, ema_26, rsi_14 for family=ta; sp_ou_halflife, sp_hmm_regime, sp_har_rv_5 for family=sp), all DOUBLE. Writes are Appender-based fresh writes per family (no read-merge-write cycle).
| col | type |
|---|---|
t |
TIMESTAMPTZ |
sma_5, sma_20, ema_12, ema_26 |
DOUBLE |
rsi_14 |
DOUBLE |
macd, macd_signal, macd_hist |
DOUBLE |
bb_upper, bb_middle, bb_lower |
DOUBLE |
atr_14, adx_14, stoch_k, stoch_d |
DOUBLE |
_feature_<name> |
DOUBLE |
coverage.parquet — the cache index: the exact loaded window per bar set. This is what makes direct hits fast.
| col | type |
|---|---|
market |
VARCHAR |
timeframe |
VARCHAR |
symbol |
VARCHAR |
first_t |
TIMESTAMPTZ |
last_t |
TIMESTAMPTZ |
num_bars |
BIGINT |
feed |
VARCHAR |
adjustment |
VARCHAR |
loaded_at |
TIMESTAMPTZ |
updated_at |
TIMESTAMPTZ |
manifest.yaml (plain text, not parquet) — lake config so reads/writes stay consistent:
lake_version: 1
default_market: US
default_feed: iex # iex is the default; SIP is never used (requires license)
default_adjustment: raw # raw|split|dividend|all — pick once per lake
timezone: UTC
features_lib: ta-lib
Read path (cache-first)
DuckDB
duckdb :memory:
-- hive glob adds market/timeframe/symbol columns automatically
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=*/timeframe=*/symbol=*.parquet');
Canonical queries:
-- past 2 months, 1d bars
SELECT symbol, date, o, h, l, c, v, n, vw
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet')
WHERE symbol = 'AAPL'
AND t >= now() - INTERVAL 2 MONTH
ORDER BY t;
-- past 2 days, 10m bars
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=10m/symbol=*.parquet')
WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 DAY ORDER BY t;
-- past 2 hours, 1m bars
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1m/symbol=*.parquet')
WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 HOUR ORDER BY t;
Join with features (hive-partitioned, family=ta):
SELECT b.t, b.c, f.sma_20, f.rsi_14
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b
LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet') f
ON f.t=b.t
WHERE b.t >= now() - INTERVAL 2 MONTH;
Join with SP features (family=sp):
SELECT b.t, b.c, sp.sp_ou_halflife, sp.sp_hmm_regime
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b
LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=sp/symbol=AAPL.parquet') sp
ON sp.t=b.t
WHERE b.t >= now() - INTERVAL 2 MONTH;
Apache Arrow / Python
import pyarrow.parquet as pq
t = pq.read_table(
"$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet",
filters=[("symbol", "==", "AAPL")],
)
df = t.to_pandas()
Verify lake data (duckdb CLI)
Any parquet file in the lake can be inspected directly with the DuckDB CLI — no MCP call needed. Handy for confirming a get_lake_bars/backfill_lake_calendar write landed:
Dependencies (duckdb + apache arrow)
duckdb and pyarrow are declared in tac-qlib/pyproject.toml (installed into the repo .venv by uv). If the runtime venv lacks them, lazy-install rather than falling back to another SQL tool:
uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow # or: uv pip install -e ./tac-qlib
Then re-check with python -c "import duckdb, pyarrow". Only use the DuckDB CLI / pyarrow path when the MCP lake tools can't answer (see MCP-first policy above).
duckdb :memory: "SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;"
Or interactively:
duckdb :memory:
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;
Quick checks:
- Bars written:
SELECT count(*), min(t), max(t) FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet'); - Coverage index:
SELECT * FROM read_parquet('$TAC_LAKE_DIR/coverage.parquet') LIMIT 10; - Metadata:
SELECT * FROM read_parquet('$TAC_LAKE_DIR/symbols.parquet') LIMIT 10; - Calendar:
SELECT * FROM read_parquet('$TAC_LAKE_DIR/calendar.parquet') LIMIT 10;
Note:
TAC_LAKE_DIRis mandatory and must be an absolute path — do not use~or$HOME(no fallback/expansion logic exists; a literal~is not expanded by shells/duckdb inside an env var).
Lazy-load write path
The core procedure when the requested range is not fully covered. Steps 1–9 are automated by the get_lake_bars lake tool (lazy=true) — the manual walk-through below documents what it does under the hood, and is the pattern to follow if writing the lake directly (DuckDB/pyarrow):
-
Normalize the request.
market, laketimeframe(map back to MCP spelling),symbols,start,end. Decidefeedandadjustmentfrommanifest.yaml(or request overrides). Keep them fixed per lake — mixing feeds/adjustments corrupts history. -
Check coverage (
coverage.parquet). See decision table below. -
Compute the missing window(s). e.g. request
[S,E], lake has[S,M]→ fetch(M,E]; no row → fetch[S,E]. -
Call MCP
get_stock_bars(multi-symbol variant; comma-separatedsymbols):{"symbols":"AAPL,MSFT","timeframe":"1Day","start":"2026-06-06T00:00:00Z","end":"2026-08-06T00:00:00Z","feed":"iex","adjustment":"raw","limit":10000}Response:
{"bars": {"AAPL": [{t,o,h,l,c,v,n,vw}, …], …}, "next_page_token": "…"}. Bars are sorted symbol-first, so a page may contain only some symbols — loop withnext_page_tokenuntilnull. -
Parse + normalize. Keep
t,o,h,l,c,v,n,vw; converttto UTCTIMESTAMPTZ; adddatefor1d. -
Merge into the partition file
$TAC_LAKE_DIR/market=<m>/timeframe=<tf>/symbol=<s>.parquet: read the existing file, overlay the fetched bars keyed by canonical timestamp (new wins on duplicatet), write the full merged set to a tmp file, then atomically rename over the old one. This is safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history). -
Update
coverage.parquet: recomputefirst_t/last_t/num_barsfrom the actual file contents (not the fetched range) — a fetch that landed nothing must not widen the span into a hollow coverage. -
Update metadata: upsert
symbols.parquet(first_seen= the symbol's earliest bar in the lake) andcalendar.parquet(distinctdates observed in bars,source='bars'; the bars path records only trading days — no session/prices). -
Return the requested range from the lake (the read path above).
Coverage decision table
For a request (market, timeframe, symbol, S, E) against coverage.parquet:
| coverage row | action |
|---|---|
| missing | backfill whole [S,E] |
first_t <= S and last_t >= E |
direct hit — read from lake, no fetch |
first_t > S |
fetch [S, first_t) prefix, merge |
last_t < E |
fetch (last_t, E] suffix, merge |
(with calendar) for 1d: expected trading days ∈ [S,E] == bars present |
consider complete |
Use calendar.parquet for the 1d completeness check — a weekend/holiday gap is normal, so “no bar on Saturday” must not trigger a refetch. Also treat the in-progress current session carefully: an intraday end=now should not trigger a refetch loop on the forming bar.
Bulk loading multiple symbols
For loading bars + TA + SP features for many symbols at once, use load_lake_symbols. It runs in a background thread and returns immediately with a job_id:
{"symbols":"AAPL,MSFT,GOOGL,AMZN","timeframe":"1d","start":"2020-01-03","end":"2026-08-15"}
Response:
{"job_id":"load-20260815-143022","status":"started","symbols":["AAPL","MSFT","GOOGL","AMZN"],"note":"load running in background -- poll load_lake_status with this job_id to track progress"}
Poll progress with load_lake_status:
{"job_id":"load-20260815-143022"}
Response (while running):
{"job_id":"load-20260815-143022","status":"running","total_symbols":4,"processed":2,"results":[...]}
Response (when done):
{"job_id":"load-20260815-143022","status":"completed","total_symbols":4,"processed":4,"results":[...],"completed_at":"2026-08-15T14:35:00Z"}
Each entry in results contains per-symbol fetch_start (the date the load started from), bars_count, bars_source, ta_count, ta_columns, sp_count, sp_columns, and any *_error fields.
Calendar gap — why get_stock_auctions
The MCP tool surface has no calendar endpoint, but the coverage check needs to know which days are trading days before bars exist. The auctions tool fills this gap:
get_stock_auctionsonly acceptsfeed: "sip"(SIP is the only valid feed for auctions).- Auction records exist only on trading days → the set of distinct dates
dacross symbols is the trading-day set. - Response shape:
{"auctions": {"AAPL": [{"d":"2026-06-09","o":[{t,x,p,c}…],"c":[{t,x,p,c}…]}, …]}, "next_page_token": "…"}—o= opening auctions,c= closing auctions.
{"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"}
Usage: to seed/enrich calendar.parquet for a range before loading bars, call the backfill_lake_calendar lake tool (it loops get_stock_auctions internally across symbols and pages, inserts one row per distinct d with source='auctions', session_open/open_price from o[0], session_close/close_price from c[0], and reports calendar_days_added). Direct call equivalent:
{"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"}
Cheap single-day confirmation for “was this a trading day?” and first pass of daily open/close. Intraday bars and per-symbol coverage still come from get_stock_bars.
Operations notes
- Rate limits: Alpaca data API ~200 req/min. On
429back off (exponential, start 1s) and retry. Batch symbols in one call, but page throughnext_page_token. - Atomicity: write parquet to a
.<name>.tmpthenrename(); readers never see partial files. Apply the same pattern to metadata upserts. - Consistency: one
feed+ oneadjustmentper lake (record inmanifest.yaml). Refetching a window with a different feed/adjustment would silently corrupt merged history. - Feed / history limits (Alpaca): SIP is never used (requires license). IEX goes back to 2020-07-27 for daily bars. For earlier data, Yahoo Finance fills gaps automatically (daily bars only). The
TAC_LAKE_START_DATE(default2000-01-03) controls the earliest date requested; Yahoo provides data back to ~1970 for most symbols. - Dedup / merge: bar upserts read the existing file, merge by canonical timestamp (new wins on duplicate
t), and write the full set atomically — safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history). Feature persistence is a fresh write per family, so no dedup needed. - Features are derived from the lake bars (compute after bars are persisted, keyed
(market, symbol, timeframe, t)), so indicator history stays aligned with bar history.
End-to-end example (1d, 2 months, AAPL)
coverage.parquethas no(US,1d,AAPL)row → backfill.- Seed calendar:
backfill_lake_calendar{"symbols":"AAPL","start":…,"end":…}(loopsget_stock_auctions) →calendar.parquettrading days. get_lake_bars{"symbols":"AAPL","timeframe":"1d","start":…,"end":…,"lazy":true,"quiet":true}→ auto backfills the window, persists bars, updates coverage/symbols/calendar, returns a{count, first_t, last_t}summary instead of the bar rows.- Answer: DuckDB
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') WHERE t >= now() - INTERVAL 2 MONTH(orget_lake_barsagain). - Next identical request is a direct hit (
source:"lake") from step 1’s decision table — no Alpaca fetch.