Files
book-tac/tac-engine/skills/tradeac-lake/SKILL.md
T

418 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: tradeac-lake
description: Guide agents to build and query the TradeAC parquet+DuckDB data lake on the local filesystem — hive-partitioned bar store (market/timeframe/symbol) plus symbols, watchlist, calendar, features and coverage metadata — with lazy backfill from the tac-engine MCP get_stock_bars tool (tradeac-alpaca skill).
---
# tradeac-lake
Local-first market data lake: **Apache Parquet** files on disk, consumed with **DuckDB** (or Apache Arrow). Bar data is the core payload; the lake also keeps small metadata parquet files (symbols, watchlist, calendar, features, coverage) at the lake root.
Reading is a **cache-first** pattern: if the requested range is already in the lake, serve it directly from parquet; otherwise **lazy-load** the missing window via the tac-engine MCP `get_stock_bars` tool (see `tac-engine/skills/tradeac-alpaca/SKILL.md`), persist it, update metadata, then return.
## MCP-first policy
- **Prefer the tac-engine lake MCP tools** (`get_lake_bars`, `get_lake_ta`, `get_lake_sp`, `get_lake_features`, `get_lake_status`, `get_lake_coverage`, `get_lake_calendar`, `backfill_lake_calendar`, …) whenever they cover the need. They handle coverage checks, lazy backfill, feed fallback, metadata updates and pagination for you — do not reimplement that in DuckDB/pyarrow scripts.
- **Direct parquet reads are only for verification** (DuckDB CLI / pyarrow snippets below) or when no lake tool covers the query (e.g. an arbitrary ad-hoc SQL join). Keep hand-rolled lake *writes* off the happy path — the write path is what the MCP tools automate.
- **NEVER script directly against the MCP server** (spawning the engine binary, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- The engine bundles its own DuckDB; direct verification only needs the `duckdb` CLI or a Python venv with `duckdb` + `pyarrow` (see dependencies below).
## Env / root
| Var | Default | Purpose |
|-----|---------|---------|
| `TAC_LAKE_DIR` | **required** (no default) | lake root on the local filesystem. Local dev: absolute path (e.g. `/home/data/lake`). |
```bash
export TAC_LAKE_DIR=/path/to/lake
mkdir -p "$TAC_LAKE_DIR"
```
## Secrets policy
- NEVER write secrets into files: API keys, DB passwords, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `APCA_*`) in scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## Lake layout
Hive partition convention, partitioned by `market`, `timeframe`, `symbol`. Metadata parquet files live alongside the partition dirs at the lake root.
```
$TAC_LAKE_DIR/
├── market=US/
│ └── timeframe=1d/
│ ├── symbol=AAPL.parquet
│ ├── symbol=MSFT.parquet
│ └── ...
│ └── timeframe=10m/
│ └── symbol=AAPL.parquet
├── market=CRYPTO/... # optional: BTC/USD etc.
├── features/ # TA + SP indicators, hive-partitioned with family tier
│ └── market=US/
│ └── timeframe=1d/
│ ├── family=ta/
│ │ └── symbol=AAPL.parquet # TA indicators (sma, rsi, macd, ...)
│ └── family=sp/
│ └── symbol=AAPL.parquet # Stochastic-process features (ou, hmm, har, ...)
├── symbols.parquet # asset master seen/known to the lake
├── watchlist.parquet # watchlists
├── calendar.parquet # trading days per market (coverage ground truth)
├── coverage.parquet # per (market,timeframe,symbol) loaded window
└── manifest.yaml # lake config: feed, adjustment, timezone
```
Convention: **one parquet file per symbol per timeframe per family** under the partition dirs. Bar upserts **merge by canonical timestamp** (read existing file → overlay new bars → write the full set atomically); feature persistence is a **fresh write** per family (Appender-based, no merge with existing).
## Lake MCP tools (tac-engine)
The tac-engine MCP server exposes a **lake tools** category that wraps the read/write paths below. `symbols`/`market` are uppercased, `timeframe` is normalized to lake spelling, and `start`/`end` accept `YYYY-MM-DD` or RFC-3339 (default `end=now`, `start=end-30d`).
| Tool | Purpose |
|------|---------|
| `get_lake_bars` | Cache-first bars: `{market?, symbols, timeframe, start?, end?, feed?, adjustment?, lazy?, quiet?}`. Feed defaults to `iex`; SIP is never used (requires license). For daily bars, Yahoo Finance fills gaps before Alpaca's earliest available date. `lazy=true` (default) backfills missing windows via Alpaca and persists; `lazy=false` reads the lake only. Returns `{request, source, bars: {SYM: [{t,o,h,l,c,v,n,vw}]}}`; `source` is `lake` (complete hit), `partial` (present but missing windows and `lazy=false`), or `fetched` (gaps backfilled). With `quiet=true` returns `{request, source, summary: {SYM: {count, first_t, last_t}}}` instead of the bar rows — use for backfill-to-lake jobs. Bar writes **merge by timestamp** (read existing file, overlay new bars, write the full set atomically) — safe for both tail appends and leading-gap backfills. Coverage/symbols/calendar metadata are reconciled against the actual file on each write. |
| `get_lake_ta` | Compute + optionally persist indicators: `{market?, symbol, timeframe, start?, end?, indicators?, persist?, quiet?}`. `indicators` is comma-separated, default all: `sma_5,sma_20,ema_12,ema_26,rsi_14,macd,bb,atr_14,adx_14`. Lookback is pulled automatically; returned rows cover `[start,end]`. When `persist=true`, features are written to `features/market=*/timeframe=*/family=ta/symbol=*.parquet`. With `quiet=true` returns `{count, columns, persisted}` instead of the feature rows — use when the goal is persisting indicators. |
| `get_lake_sp` | Compute + optionally persist **stochastic-process features** (`sp_*` columns) from lake bars: `{market?, symbol, timeframe, start?, end?, fit_end?, families?, persist?, quiet?}`. Rust port of `sp_features.py` on the stochastic-rs stack. `families` is comma-separated, default all: `ou,hmm,jump,har,trend,hurst,signature,moments`. In addition to `sp_rv*`/`sp_vol_ratio_*` the `har` family also emits `sp_rv_ac1` (RV lag-1 autocorr) + `sp_rv_cv_22` (RV coefficient of variation); `jump` also emits `sp_max_up`/`sp_max_down` (signed max-move asymmetry); `signature` also emits the lag-5 level-2 cross terms `sp_sig_level2_{lead_lag,lag_lead}_5`; and `moments` emits the scale-free realized skewness/kurtosis (`sp_rskew_{5,22}`, `sp_rkurt_{5,22}`; a 1-day window is undefined) and downside semi-variance (`sp_dsv_{1,5,22}`, `sp_dsv_ratio_{1,5,22}`) via stochastic-rs `realized`. `fit_end` limits the 2-state Gaussian-HMM fit window (no lookahead; posteriors still cover the whole window). On `persist=true`, features are written to `features/market=*/timeframe=*/family=sp/symbol=*.parquet`. With `quiet=true` returns `{count, sp_columns, persisted}` instead of the feature rows. Deferred (not in stochastic-rs): `garch`, `entropy`, `catch22`. |
| `get_lake_features` | Read persisted TA + SP features (hive-partitioned `features/` dir, families `ta` and `sp` merged by timestamp): `{market?, symbol?, timeframe?, start?, end?, quiet?}`. With `quiet=true` returns `{count, columns}` instead of the full feature rows. |
| `get_lake_symbols` | Read `symbols.parquet` asset master; optional `{symbol?}` filter. |
| `get_lake_watchlist` | Read `watchlist.parquet`. |
| `get_lake_calendar` | Read `calendar.parquet` trading days: `{market?, start?, end?}`. |
| `get_lake_coverage` | Read `coverage.parquet` cache index: `{market?, timeframe?, symbol?}`. |
| `get_lake_status` | Lake root, `manifest.yaml`, bar partition inventory and metadata file sizes. |
| `rebuild_lake_symbol` | Delete bar + feature parquet files and re-fetch from `TAC_LAKE_START_DATE` (default `2000-01-03`) for a single symbol: `{market?, symbol, timeframe, feed?, adjustment?}`. Resets coverage so the next `get_lake_bars` call re-downloads the full history. Use after changing `TAC_LAKE_START_DATE` or to fix stale/corrupt data. |
| `load_lake_symbols` | **Bulk-load + persist** bars + TA + SP features for a comma-separated list of symbols. Runs in a background thread and returns immediately with a `job_id`: `{market?, symbols, timeframe, start?, end?, feed?, adjustment?, indicators?, families?}`. Per-symbol start is computed automatically from lake coverage: if the lake has no data or `first_t > TAC_LAKE_START_DATE`, fetches from `TAC_LAKE_START_DATE` (default 2000-01-03); if `first_t <= TAC_LAKE_START_DATE`, fetches only from `last_t` (tail refresh). TA/SP features are always computed over the full `TAC_LAKE_START_DATE` to `end` range. Poll `load_lake_status` with the returned `job_id` to track progress. |
| `load_lake_status` | Query the status of a background bulk-load job: `{job_id}`. Returns `{job_id, status, total_symbols, processed, results, error, started_at, completed_at}` where `status` is `running`, `completed`, or `failed`, and `results` contains per-symbol bar counts, TA/SP column counts, and any errors. |
| `backfill_lake_calendar` | **Gap-fill tool**: seed/enrich `calendar.parquet` from Alpaca historical auctions (feed=iex; records exist only on trading days): `{market?, symbols, start?, end?}`. Returns `{market, symbols, start, end, calendar_days_added}`. Call this before lazy bar loads so the `1d` completeness check knows the expected trading-day set. |
| `validate_lake_dataset` | **Pre-workflow quality gate**: `{market?, timeframe?, symbols?, start?, end?}`. Scans every symbol in coverage (or a comma-separated `symbols` subset) and reports `verdict: OK/WARNINGS/ERRORS` plus per-symbol issues. Catches the failure modes qlib silently tolerates: **all-NaN feature columns** (would be dropped by `DropAllNaN` — the model trains on fewer features without notice), **missing TA/SP feature files**, **hollow coverage / stale date ranges** (coverage claims a wide span but the bar file is empty/truncated/sparse), **stale coverage** (first/last/num_bars vs the actual file), and **partition misalignment** (flat-layout feature orphans the family=ta|sp consumers can't see). Pass `start`/`end` to also check feature-vs-bar row alignment and per-column all-NaN status in that window. Call before `rd_run_workflow` / `rd_train` to fail fast instead of training on silent data holes. |
Example:
```json
{"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true}
```
→ `{"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "bars": {"AAPL": [{...}], "MSFT": [{...}]}}`
## Quiet mode
`get_lake_bars`, `get_lake_ta`, `get_lake_sp` and `get_lake_features` accept `"quiet": true`. When the point of the call is **writing to the lake** (backfill/fetch bars, compute + persist indicators or `sp_*` features), use `quiet: true` — the tool still performs the full backfill / computation / persist, but returns a **summary** instead of echoing back the potentially huge payload (thousands of bar rows / feature rows). Full-row output (`bars` / `features`) is the default, so requests that *need* the data to read it must leave `quiet` unset/false.
| Tool | `quiet: true` response |
|------|------------------------|
| `get_lake_bars` | `{request, source: {SYM: lake\|partial\|fetched}, summary: {SYM: {count, first_t, last_t}}}` |
| `get_lake_ta` | `{market, symbol, timeframe, start, end, count, columns, persisted}` |
| `get_lake_sp` | `{market, symbol, timeframe, start, end, fit_end, count, sp_columns, persisted}` |
| `get_lake_features` | `{count, columns}` |
Backfill-to-lake job (no payload echoed):
```json
{"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true, "quiet": true}
```
→ `{"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "summary": {"AAPL": {"count": 44, "first_t": "2026-06-06T04:00:00Z", "last_t": "2026-08-05T04:00:00Z"}, "MSFT": {...}}}`
Persist indicators to the lake (summary only):
```json
{"symbol": "AAPL", "timeframe": "1d", "indicators": "sma_5,sma_20,rsi_14", "persist": true, "quiet": true}
```
→ `{"market": "US", "symbol": "AAPL", "timeframe": "1d", "count": 44, "columns": ["sma_5","sma_20","rsi_14"], "persisted": true}`
```json
{"symbols": "AAPL,MSFT", "start": "2026-06-06"}
```
→ `{"market": "US", "symbols": ["AAPL","MSFT"], "start": ..., "end": ..., "calendar_days_added": 44}`
## Conventions
- `market`: `US` (equities), `CRYPTO`, `FOREX`. Uppercase.
- `timeframe`: normalized lake name — lowercase, `1m 5m 10m 15m 30m 1h 2h 4h 1d 1w 1M`. The MCP tool spells them differently; always map:
| Lake | MCP `timeframe` | Lake | MCP `timeframe` |
|------|-----------------|------|-----------------|
| `1m` | `1Min` | `2h` | `2Hour` |
| `5m` | `5Min` | `4h` | `4Hour` |
| `10m` | `10Min` | `1d` | `1Day` |
| `15m` | `15Min` | `1w` | `1Week` |
| `30m` | `30Min` | `1M` | `1Month` |
| `1h` | `1Hour` | | |
- `symbol`: uppercase, e.g. `AAPL`. Hyphens/`.` in special symbols (e.g. `BRK-B`, `SPY`) are valid filenames; avoid `/` and spaces.
- All timestamps stored as **UTC** instants (`TIMESTAMPTZ`). Alpaca returns RFC-3339 UTC; normalize on write.
- `1d` bars: `t` is the session date at `04:00Z` (midnight ET — Alpaca stamps daily bars at `04:00:00Z`); also store a `date` column (`CAST(t AS DATE)`, UTC) for calendar joins. A date-only `end` (e.g. `2026-08-05`) is treated as **inclusive of the whole end day**, so the end-day bar is not dropped.
## Bar parquet schema (`market=…/timeframe=…/symbol=….parquet`)
| col | type | source field |
|-----|------|--------------|
| `t` | TIMESTAMPTZ | bar `t` (UTC) |
| `o` | DOUBLE | `o` |
| `h` | DOUBLE | `h` |
| `l` | DOUBLE | `l` |
| `c` | DOUBLE | `c` |
| `v` | BIGINT | `v` |
| `n` | BIGINT | `n` |
| `vw` | DOUBLE | `vw` |
Partition columns `market`/`timeframe`/`symbol` are derived from the path; DuckDB exposes them automatically when reading a hive glob.
## Metadata parquet files
All written with DuckDB `COPY … (FORMAT PARQUET)` from in-memory `SELECT`, or `pyarrow.parquet`.
`symbols.parquet`
| col | type | notes |
|-----|------|-------|
| `symbol` | VARCHAR (pk) |
| `name` | VARCHAR |
| `asset_class` | VARCHAR |
| `exchange` | VARCHAR |
| `tradable` | BOOLEAN |
| `status` | VARCHAR |
| `first_seen` | TIMESTAMPTZ | the symbol's earliest bar in the lake (its first trading date), not the load timestamp |
| `updated_at` | TIMESTAMPTZ | |
`watchlist.parquet`
| col | type |
|-----|------|
| `watchlist_id` | VARCHAR |
| `name` | VARCHAR |
| `symbol` | VARCHAR |
| `added_at` | TIMESTAMPTZ |
| `updated_at` | TIMESTAMPTZ |
`calendar.parquet` — the trading-day ground truth per market (see “Calendar gap” below). Bars only seed which dates are trading days; per-symbol prices/session times are NOT attributed by the bars path (no symbol column, 1d bars all share `t=04:00`).
| col | type | notes |
|-----|------|-------|
| `market` | VARCHAR | pk + `date` |
| `date` | DATE | a trading day (UTC) |
| `session_open` | TIMESTAMPTZ | from auctions `o[0].t` only (nullable; not set by bars) |
| `session_close` | TIMESTAMPTZ | from auctions `c[0].t` only (nullable; not set by bars) |
| `open_price` | DOUBLE | opening auction price (nullable) |
| `close_price` | DOUBLE | closing auction price (nullable) |
| `source` | VARCHAR | `auctions` \| `bars` \| `manual` |
| `updated_at` | TIMESTAMPTZ | |
`features/` — TA + stochastic-process indicators, **wide** format, hive-partitioned with a `family` tier: `features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet` and `family=sp/symbol=AAPL.parquet`. Each row is one `t`, with one column per indicator. The partition columns (market/symbol/timeframe) come from the directory structure; the file itself stores `t` + indicator columns (e.g. `sma_5`, `sma_20`, `ema_12`, `ema_26`, `rsi_14` for `family=ta`; `sp_ou_halflife`, `sp_hmm_regime`, `sp_har_rv_5` for `family=sp`), all DOUBLE. Writes are Appender-based fresh writes per family (no read-merge-write cycle).
| col | type |
|-----|------|
| `t` | TIMESTAMPTZ |
| `sma_5`, `sma_20`, `ema_12`, `ema_26` | DOUBLE |
| `rsi_14` | DOUBLE |
| `macd`, `macd_signal`, `macd_hist` | DOUBLE |
| `bb_upper`, `bb_middle`, `bb_lower` | DOUBLE |
| `atr_14`, `adx_14`, `stoch_k`, `stoch_d` | DOUBLE |
| `_feature_<name>` | DOUBLE |
`coverage.parquet` — **the cache index**: the exact loaded window per bar set. This is what makes direct hits fast.
| col | type |
|-----|------|
| `market` | VARCHAR |
| `timeframe` | VARCHAR |
| `symbol` | VARCHAR |
| `first_t` | TIMESTAMPTZ |
| `last_t` | TIMESTAMPTZ |
| `num_bars` | BIGINT |
| `feed` | VARCHAR |
| `adjustment` | VARCHAR |
| `loaded_at` | TIMESTAMPTZ |
| `updated_at` | TIMESTAMPTZ |
`manifest.yaml` (plain text, not parquet) — lake config so reads/writes stay consistent:
```yaml
lake_version: 1
default_market: US
default_feed: iex # iex is the default; SIP is never used (requires license)
default_adjustment: raw # raw|split|dividend|all — pick once per lake
timezone: UTC
features_lib: ta-lib
```
## Read path (cache-first)
### DuckDB
```bash
duckdb :memory:
```
```sql
-- hive glob adds market/timeframe/symbol columns automatically
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=*/timeframe=*/symbol=*.parquet');
```
Canonical queries:
```sql
-- past 2 months, 1d bars
SELECT symbol, date, o, h, l, c, v, n, vw
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet')
WHERE symbol = 'AAPL'
AND t >= now() - INTERVAL 2 MONTH
ORDER BY t;
-- past 2 days, 10m bars
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=10m/symbol=*.parquet')
WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 DAY ORDER BY t;
-- past 2 hours, 1m bars
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1m/symbol=*.parquet')
WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 HOUR ORDER BY t;
```
Join with features (hive-partitioned, family=ta):
```sql
SELECT b.t, b.c, f.sma_20, f.rsi_14
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b
LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet') f
ON f.t=b.t
WHERE b.t >= now() - INTERVAL 2 MONTH;
```
Join with SP features (family=sp):
```sql
SELECT b.t, b.c, sp.sp_ou_halflife, sp.sp_hmm_regime
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b
LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=sp/symbol=AAPL.parquet') sp
ON sp.t=b.t
WHERE b.t >= now() - INTERVAL 2 MONTH;
```
### Apache Arrow / Python
```python
import pyarrow.parquet as pq
t = pq.read_table(
"$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet",
filters=[("symbol", "==", "AAPL")],
)
df = t.to_pandas()
```
## Verify lake data (duckdb CLI)
Any parquet file in the lake can be inspected directly with the **DuckDB CLI** — no MCP call needed. Handy for confirming a `get_lake_bars`/`backfill_lake_calendar` write landed:
### Dependencies (duckdb + apache arrow)
`duckdb` and `pyarrow` are declared in `tac-qlib/pyproject.toml` (installed into the repo `.venv` by `uv`). If the runtime venv lacks them, **lazy-install** rather than falling back to another SQL tool:
```bash
uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow # or: uv pip install -e ./tac-qlib
```
Then re-check with `python -c "import duckdb, pyarrow"`. Only use the DuckDB CLI / pyarrow path when the MCP lake tools can't answer (see MCP-first policy above).
```bash
duckdb :memory: "SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;"
```
Or interactively:
```bash
duckdb :memory:
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;
```
Quick checks:
- **Bars written:** `SELECT count(*), min(t), max(t) FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet');`
- **Coverage index:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/coverage.parquet') LIMIT 10;`
- **Metadata:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/symbols.parquet') LIMIT 10;`
- **Calendar:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/calendar.parquet') LIMIT 10;`
> Note: `TAC_LAKE_DIR` is **mandatory** and must be an absolute path — do **not** use `~` or `$HOME`
> (no fallback/expansion logic exists; a literal `~` is not expanded by shells/duckdb inside an env var).
## Lazy-load write path
The core procedure when the requested range is **not** fully covered. Steps 1–9 are automated by the **`get_lake_bars`** lake tool (`lazy=true`) — the manual walk-through below documents what it does under the hood, and is the pattern to follow if writing the lake directly (DuckDB/pyarrow):
1. **Normalize the request.** `market`, lake `timeframe` (map back to MCP spelling), `symbols`, `start`, `end`. Decide `feed` and `adjustment` from `manifest.yaml` (or request overrides). Keep them fixed per lake — mixing feeds/adjustments corrupts history.
2. **Check coverage** (`coverage.parquet`). See decision table below.
3. **Compute the missing window(s).** e.g. request `[S,E]`, lake has `[S,M]` → fetch `(M,E]`; no row → fetch `[S,E]`.
4. **Call MCP `get_stock_bars`** (multi-symbol variant; comma-separated `symbols`):
```json
{"symbols":"AAPL,MSFT","timeframe":"1Day","start":"2026-06-06T00:00:00Z","end":"2026-08-06T00:00:00Z","feed":"iex","adjustment":"raw","limit":10000}
```
Response: `{"bars": {"AAPL": [{t,o,h,l,c,v,n,vw}, …], …}, "next_page_token": "…"}`. Bars are sorted symbol-first, so a page may contain only some symbols — **loop with `next_page_token`** until `null`.
5. **Parse + normalize.** Keep `t,o,h,l,c,v,n,vw`; convert `t` to UTC `TIMESTAMPTZ`; add `date` for `1d`.
6. **Merge into the partition file** `$TAC_LAKE_DIR/market=<m>/timeframe=<tf>/symbol=<s>.parquet`: read the existing file, overlay the fetched bars keyed by canonical timestamp (new wins on duplicate `t`), write the full merged set to a tmp file, then atomically rename over the old one. This is safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history).
7. **Update `coverage.parquet`**: recompute `first_t`/`last_t`/`num_bars` from the **actual file contents** (not the fetched range) — a fetch that landed nothing must not widen the span into a hollow coverage.
8. **Update metadata**: upsert `symbols.parquet` (`first_seen` = the symbol's earliest bar in the lake) and `calendar.parquet` (distinct `date`s observed in bars, `source='bars'`; the bars path records only trading days — no session/prices).
9. **Return the requested range** from the lake (the read path above).
### Coverage decision table
For a request `(market, timeframe, symbol, S, E)` against `coverage.parquet`:
| coverage row | action |
|--------------|--------|
| missing | backfill whole `[S,E]` |
| `first_t <= S` and `last_t >= E` | **direct hit** — read from lake, no fetch |
| `first_t > S` | fetch `[S, first_t)` prefix, merge |
| `last_t < E` | fetch `(last_t, E]` suffix, merge |
| (with `calendar`) for `1d`: expected trading days `∈ [S,E]` == bars present | consider complete |
Use `calendar.parquet` for the `1d` completeness check — a weekend/holiday gap is normal, so “no bar on Saturday” must **not** trigger a refetch. Also treat the in-progress current session carefully: an intraday `end=now` should not trigger a refetch loop on the forming bar.
### Bulk loading multiple symbols
For loading bars + TA + SP features for many symbols at once, use **`load_lake_symbols`**. It runs in a background thread and returns immediately with a `job_id`:
```json
{"symbols":"AAPL,MSFT,GOOGL,AMZN","timeframe":"1d","start":"2020-01-03","end":"2026-08-15"}
```
Response:
```json
{"job_id":"load-20260815-143022","status":"started","symbols":["AAPL","MSFT","GOOGL","AMZN"],"note":"load running in background -- poll load_lake_status with this job_id to track progress"}
```
Poll progress with **`load_lake_status`**:
```json
{"job_id":"load-20260815-143022"}
```
Response (while running):
```json
{"job_id":"load-20260815-143022","status":"running","total_symbols":4,"processed":2,"results":[...]}
```
Response (when done):
```json
{"job_id":"load-20260815-143022","status":"completed","total_symbols":4,"processed":4,"results":[...],"completed_at":"2026-08-15T14:35:00Z"}
```
Each entry in `results` contains per-symbol `fetch_start` (the date the load started from), `bars_count`, `bars_source`, `ta_count`, `ta_columns`, `sp_count`, `sp_columns`, and any `*_error` fields.
### Calendar gap — why `get_stock_auctions`
The MCP tool surface has **no calendar endpoint**, but the coverage check needs to know which days are trading days before bars exist. The **auctions** tool fills this gap:
- `get_stock_auctions` only accepts `feed: "sip"` (SIP is the only valid feed for auctions).
- Auction records exist **only on trading days** → the set of distinct dates `d` across symbols is the trading-day set.
- Response shape: `{"auctions": {"AAPL": [{"d":"2026-06-09","o":[{t,x,p,c}…],"c":[{t,x,p,c}…]}, …]}, "next_page_token": "…"}` — `o` = opening auctions, `c` = closing auctions.
```json
{"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"}
```
Usage: to seed/enrich `calendar.parquet` for a range **before** loading bars, call the **`backfill_lake_calendar`** lake tool (it loops `get_stock_auctions` internally across symbols and pages, inserts one row per distinct `d` with `source='auctions'`, `session_open`/`open_price` from `o[0]`, `session_close`/`close_price` from `c[0]`, and reports `calendar_days_added`). Direct call equivalent:
```json
{"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"}
```
Cheap single-day confirmation for “was this a trading day?” and first pass of daily open/close. Intraday bars and per-symbol coverage still come from `get_stock_bars`.
## Operations notes
- **Rate limits:** Alpaca data API ~200 req/min. On `429` back off (exponential, start 1s) and retry. Batch symbols in one call, but page through `next_page_token`.
- **Atomicity:** write parquet to a `.<name>.tmp` then `rename()`; readers never see partial files. Apply the same pattern to metadata upserts.
- **Consistency:** one `feed` + one `adjustment` per lake (record in `manifest.yaml`). Refetching a window with a different feed/adjustment would silently corrupt merged history.
- **Feed / history limits (Alpaca):** SIP is never used (requires license). IEX goes back to **2020-07-27** for daily bars. For earlier data, Yahoo Finance fills gaps automatically (daily bars only). The `TAC_LAKE_START_DATE` (default `2000-01-03`) controls the earliest date requested; Yahoo provides data back to ~1970 for most symbols.
- **Dedup / merge:** bar upserts read the existing file, merge by canonical timestamp (new wins on duplicate `t`), and write the full set atomically — safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history). Feature persistence is a fresh write per family, so no dedup needed.
- **Features** are derived from the lake bars (compute after bars are persisted, keyed `(market, symbol, timeframe, t)`), so indicator history stays aligned with bar history.
## End-to-end example (1d, 2 months, AAPL)
1. `coverage.parquet` has no `(US,1d,AAPL)` row → backfill.
2. Seed calendar: `backfill_lake_calendar` `{"symbols":"AAPL","start":…,"end":…}` (loops `get_stock_auctions`) → `calendar.parquet` trading days.
3. `get_lake_bars` `{"symbols":"AAPL","timeframe":"1d","start":…,"end":…,"lazy":true,"quiet":true}` → auto backfills the window, persists bars, updates coverage/symbols/calendar, returns a `{count, first_t, last_t}` summary instead of the bar rows.
4. Answer: DuckDB `SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') WHERE t >= now() - INTERVAL 2 MONTH` (or `get_lake_bars` again).
5. **Next identical request is a direct hit** (`source:"lake"`) from step 1’s decision table — no Alpaca fetch.