--- name: tradeac-lake description: Guide agents to build and query the TradeAC parquet+DuckDB data lake on the local filesystem — hive-partitioned bar store (market/timeframe/symbol) plus symbols, watchlist, calendar, features and coverage metadata — with lazy backfill from the tac-engine MCP get_stock_bars tool (tradeac-alpaca skill). --- # tradeac-lake Local-first market data lake: **Apache Parquet** files on disk, consumed with **DuckDB** (or Apache Arrow). Bar data is the core payload; the lake also keeps small metadata parquet files (symbols, watchlist, calendar, features, coverage) at the lake root. Reading is a **cache-first** pattern: if the requested range is already in the lake, serve it directly from parquet; otherwise **lazy-load** the missing window via the tac-engine MCP `get_stock_bars` tool (see `tac-engine/skills/tradeac-alpaca/SKILL.md`), persist it, update metadata, then return. ## MCP-first policy - **Prefer the tac-engine lake MCP tools** (`get_lake_bars`, `get_lake_ta`, `get_lake_sp`, `get_lake_features`, `get_lake_status`, `get_lake_coverage`, `get_lake_calendar`, `backfill_lake_calendar`, …) whenever they cover the need. They handle coverage checks, lazy backfill, feed fallback, metadata updates and pagination for you — do not reimplement that in DuckDB/pyarrow scripts. - **Direct parquet reads are only for verification** (DuckDB CLI / pyarrow snippets below) or when no lake tool covers the query (e.g. an arbitrary ad-hoc SQL join). Keep hand-rolled lake *writes* off the happy path — the write path is what the MCP tools automate. - **NEVER script directly against the MCP server** (spawning the engine binary, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**. - The engine bundles its own DuckDB; direct verification only needs the `duckdb` CLI or a Python venv with `duckdb` + `pyarrow` (see dependencies below). ## Env / root | Var | Default | Purpose | |-----|---------|---------| | `TAC_LAKE_DIR` | **required** (no default) | lake root on the local filesystem. Local dev: absolute path (e.g. `/home/data/lake`). | ```bash export TAC_LAKE_DIR=/path/to/lake mkdir -p "$TAC_LAKE_DIR" ``` ## Secrets policy - NEVER write secrets into files: API keys, DB passwords, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `APCA_*`) in scripts, configs, notes or committed code. - NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it. - When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself. - If you find a committed secret, flag it, remove it, and replace it with a placeholder. ## Lake layout Hive partition convention, partitioned by `market`, `timeframe`, `symbol`. Metadata parquet files live alongside the partition dirs at the lake root. ``` $TAC_LAKE_DIR/ ├── market=US/ │ └── timeframe=1d/ │ ├── symbol=AAPL.parquet │ ├── symbol=MSFT.parquet │ └── ... │ └── timeframe=10m/ │ └── symbol=AAPL.parquet ├── market=CRYPTO/... # optional: BTC/USD etc. ├── features/ # TA + SP indicators, hive-partitioned with family tier │ └── market=US/ │ └── timeframe=1d/ │ ├── family=ta/ │ │ └── symbol=AAPL.parquet # TA indicators (sma, rsi, macd, ...) │ └── family=sp/ │ └── symbol=AAPL.parquet # Stochastic-process features (ou, hmm, har, ...) ├── symbols.parquet # asset master seen/known to the lake ├── watchlist.parquet # watchlists ├── calendar.parquet # trading days per market (coverage ground truth) ├── coverage.parquet # per (market,timeframe,symbol) loaded window └── manifest.yaml # lake config: feed, adjustment, timezone ``` Convention: **one parquet file per symbol per timeframe per family** under the partition dirs. Bar upserts **merge by canonical timestamp** (read existing file → overlay new bars → write the full set atomically); feature persistence is a **fresh write** per family (Appender-based, no merge with existing). ## Lake MCP tools (tac-engine) The tac-engine MCP server exposes a **lake tools** category that wraps the read/write paths below. `symbols`/`market` are uppercased, `timeframe` is normalized to lake spelling, and `start`/`end` accept `YYYY-MM-DD` or RFC-3339 (default `end=now`, `start=end-30d`). | Tool | Purpose | |------|---------| | `get_lake_bars` | Cache-first bars: `{market?, symbols, timeframe, start?, end?, feed?, adjustment?, lazy?, quiet?}`. Feed defaults to `iex`; SIP is never used (requires license). For daily bars, Yahoo Finance fills gaps before Alpaca's earliest available date. `lazy=true` (default) backfills missing windows via Alpaca and persists; `lazy=false` reads the lake only. Returns `{request, source, bars: {SYM: [{t,o,h,l,c,v,n,vw}]}}`; `source` is `lake` (complete hit), `partial` (present but missing windows and `lazy=false`), or `fetched` (gaps backfilled). With `quiet=true` returns `{request, source, summary: {SYM: {count, first_t, last_t}}}` instead of the bar rows — use for backfill-to-lake jobs. Bar writes **merge by timestamp** (read existing file, overlay new bars, write the full set atomically) — safe for both tail appends and leading-gap backfills. Coverage/symbols/calendar metadata are reconciled against the actual file on each write. | | `get_lake_ta` | Compute + optionally persist indicators: `{market?, symbol, timeframe, start?, end?, indicators?, persist?, quiet?}`. `indicators` is comma-separated, default all: `sma_5,sma_20,ema_12,ema_26,rsi_14,macd,bb,atr_14,adx_14`. Lookback is pulled automatically; returned rows cover `[start,end]`. When `persist=true`, features are written to `features/market=*/timeframe=*/family=ta/symbol=*.parquet`. With `quiet=true` returns `{count, columns, persisted}` instead of the feature rows — use when the goal is persisting indicators. | | `get_lake_sp` | Compute + optionally persist **stochastic-process features** (`sp_*` columns) from lake bars: `{market?, symbol, timeframe, start?, end?, fit_end?, families?, persist?, quiet?}`. Rust port of `sp_features.py` on the stochastic-rs stack. `families` is comma-separated, default all: `ou,hmm,jump,har,trend,hurst,signature,moments`. In addition to `sp_rv*`/`sp_vol_ratio_*` the `har` family also emits `sp_rv_ac1` (RV lag-1 autocorr) + `sp_rv_cv_22` (RV coefficient of variation); `jump` also emits `sp_max_up`/`sp_max_down` (signed max-move asymmetry); `signature` also emits the lag-5 level-2 cross terms `sp_sig_level2_{lead_lag,lag_lead}_5`; and `moments` emits the scale-free realized skewness/kurtosis (`sp_rskew_{5,22}`, `sp_rkurt_{5,22}`; a 1-day window is undefined) and downside semi-variance (`sp_dsv_{1,5,22}`, `sp_dsv_ratio_{1,5,22}`) via stochastic-rs `realized`. `fit_end` limits the 2-state Gaussian-HMM fit window (no lookahead; posteriors still cover the whole window). On `persist=true`, features are written to `features/market=*/timeframe=*/family=sp/symbol=*.parquet`. With `quiet=true` returns `{count, sp_columns, persisted}` instead of the feature rows. Deferred (not in stochastic-rs): `garch`, `entropy`, `catch22`. | | `get_lake_features` | Read persisted TA + SP features (hive-partitioned `features/` dir, families `ta` and `sp` merged by timestamp): `{market?, symbol?, timeframe?, start?, end?, quiet?}`. With `quiet=true` returns `{count, columns}` instead of the full feature rows. | | `get_lake_symbols` | Read `symbols.parquet` asset master; optional `{symbol?}` filter. | | `get_lake_watchlist` | Read `watchlist.parquet`. | | `get_lake_calendar` | Read `calendar.parquet` trading days: `{market?, start?, end?}`. | | `get_lake_coverage` | Read `coverage.parquet` cache index: `{market?, timeframe?, symbol?}`. | | `get_lake_status` | Lake root, `manifest.yaml`, bar partition inventory and metadata file sizes. | | `rebuild_lake_symbol` | Delete bar + feature parquet files and re-fetch from `TAC_LAKE_START_DATE` (default `2000-01-03`) for a single symbol: `{market?, symbol, timeframe, feed?, adjustment?}`. Resets coverage so the next `get_lake_bars` call re-downloads the full history. Use after changing `TAC_LAKE_START_DATE` or to fix stale/corrupt data. | | `load_lake_symbols` | **Bulk-load + persist** bars + TA + SP features for a comma-separated list of symbols. Runs in a background thread and returns immediately with a `job_id`: `{market?, symbols, timeframe, start?, end?, feed?, adjustment?, indicators?, families?}`. Per-symbol start is computed automatically from lake coverage: if the lake has no data or `first_t > TAC_LAKE_START_DATE`, fetches from `TAC_LAKE_START_DATE` (default 2000-01-03); if `first_t <= TAC_LAKE_START_DATE`, fetches only from `last_t` (tail refresh). TA/SP features are always computed over the full `TAC_LAKE_START_DATE` to `end` range. Poll `load_lake_status` with the returned `job_id` to track progress. | | `load_lake_status` | Query the status of a background bulk-load job: `{job_id}`. Returns `{job_id, status, total_symbols, processed, results, error, started_at, completed_at}` where `status` is `running`, `completed`, or `failed`, and `results` contains per-symbol bar counts, TA/SP column counts, and any errors. | | `backfill_lake_calendar` | **Gap-fill tool**: seed/enrich `calendar.parquet` from Alpaca historical auctions (feed=iex; records exist only on trading days): `{market?, symbols, start?, end?}`. Returns `{market, symbols, start, end, calendar_days_added}`. Call this before lazy bar loads so the `1d` completeness check knows the expected trading-day set. | | `validate_lake_dataset` | **Pre-workflow quality gate**: `{market?, timeframe?, symbols?, start?, end?}`. Scans every symbol in coverage (or a comma-separated `symbols` subset) and reports `verdict: OK/WARNINGS/ERRORS` plus per-symbol issues. Catches the failure modes qlib silently tolerates: **all-NaN feature columns** (would be dropped by `DropAllNaN` — the model trains on fewer features without notice), **missing TA/SP feature files**, **hollow coverage / stale date ranges** (coverage claims a wide span but the bar file is empty/truncated/sparse), **stale coverage** (first/last/num_bars vs the actual file), and **partition misalignment** (flat-layout feature orphans the family=ta|sp consumers can't see). Pass `start`/`end` to also check feature-vs-bar row alignment and per-column all-NaN status in that window. Call before `rd_run_workflow` / `rd_train` to fail fast instead of training on silent data holes. | Example: ```json {"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true} ``` → `{"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "bars": {"AAPL": [{...}], "MSFT": [{...}]}}` ## Quiet mode `get_lake_bars`, `get_lake_ta`, `get_lake_sp` and `get_lake_features` accept `"quiet": true`. When the point of the call is **writing to the lake** (backfill/fetch bars, compute + persist indicators or `sp_*` features), use `quiet: true` — the tool still performs the full backfill / computation / persist, but returns a **summary** instead of echoing back the potentially huge payload (thousands of bar rows / feature rows). Full-row output (`bars` / `features`) is the default, so requests that *need* the data to read it must leave `quiet` unset/false. | Tool | `quiet: true` response | |------|------------------------| | `get_lake_bars` | `{request, source: {SYM: lake\|partial\|fetched}, summary: {SYM: {count, first_t, last_t}}}` | | `get_lake_ta` | `{market, symbol, timeframe, start, end, count, columns, persisted}` | | `get_lake_sp` | `{market, symbol, timeframe, start, end, fit_end, count, sp_columns, persisted}` | | `get_lake_features` | `{count, columns}` | Backfill-to-lake job (no payload echoed): ```json {"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true, "quiet": true} ``` → `{"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "summary": {"AAPL": {"count": 44, "first_t": "2026-06-06T04:00:00Z", "last_t": "2026-08-05T04:00:00Z"}, "MSFT": {...}}}` Persist indicators to the lake (summary only): ```json {"symbol": "AAPL", "timeframe": "1d", "indicators": "sma_5,sma_20,rsi_14", "persist": true, "quiet": true} ``` → `{"market": "US", "symbol": "AAPL", "timeframe": "1d", "count": 44, "columns": ["sma_5","sma_20","rsi_14"], "persisted": true}` ```json {"symbols": "AAPL,MSFT", "start": "2026-06-06"} ``` → `{"market": "US", "symbols": ["AAPL","MSFT"], "start": ..., "end": ..., "calendar_days_added": 44}` ## Conventions - `market`: `US` (equities), `CRYPTO`, `FOREX`. Uppercase. - `timeframe`: normalized lake name — lowercase, `1m 5m 10m 15m 30m 1h 2h 4h 1d 1w 1M`. The MCP tool spells them differently; always map: | Lake | MCP `timeframe` | Lake | MCP `timeframe` | |------|-----------------|------|-----------------| | `1m` | `1Min` | `2h` | `2Hour` | | `5m` | `5Min` | `4h` | `4Hour` | | `10m` | `10Min` | `1d` | `1Day` | | `15m` | `15Min` | `1w` | `1Week` | | `30m` | `30Min` | `1M` | `1Month` | | `1h` | `1Hour` | | | - `symbol`: uppercase, e.g. `AAPL`. Hyphens/`.` in special symbols (e.g. `BRK-B`, `SPY`) are valid filenames; avoid `/` and spaces. - All timestamps stored as **UTC** instants (`TIMESTAMPTZ`). Alpaca returns RFC-3339 UTC; normalize on write. - `1d` bars: `t` is the session date at `04:00Z` (midnight ET — Alpaca stamps daily bars at `04:00:00Z`); also store a `date` column (`CAST(t AS DATE)`, UTC) for calendar joins. A date-only `end` (e.g. `2026-08-05`) is treated as **inclusive of the whole end day**, so the end-day bar is not dropped. ## Bar parquet schema (`market=…/timeframe=…/symbol=….parquet`) | col | type | source field | |-----|------|--------------| | `t` | TIMESTAMPTZ | bar `t` (UTC) | | `o` | DOUBLE | `o` | | `h` | DOUBLE | `h` | | `l` | DOUBLE | `l` | | `c` | DOUBLE | `c` | | `v` | BIGINT | `v` | | `n` | BIGINT | `n` | | `vw` | DOUBLE | `vw` | Partition columns `market`/`timeframe`/`symbol` are derived from the path; DuckDB exposes them automatically when reading a hive glob. ## Metadata parquet files All written with DuckDB `COPY … (FORMAT PARQUET)` from in-memory `SELECT`, or `pyarrow.parquet`. `symbols.parquet` | col | type | notes | |-----|------|-------| | `symbol` | VARCHAR (pk) | | `name` | VARCHAR | | `asset_class` | VARCHAR | | `exchange` | VARCHAR | | `tradable` | BOOLEAN | | `status` | VARCHAR | | `first_seen` | TIMESTAMPTZ | the symbol's earliest bar in the lake (its first trading date), not the load timestamp | | `updated_at` | TIMESTAMPTZ | | `watchlist.parquet` | col | type | |-----|------| | `watchlist_id` | VARCHAR | | `name` | VARCHAR | | `symbol` | VARCHAR | | `added_at` | TIMESTAMPTZ | | `updated_at` | TIMESTAMPTZ | `calendar.parquet` — the trading-day ground truth per market (see “Calendar gap” below). Bars only seed which dates are trading days; per-symbol prices/session times are NOT attributed by the bars path (no symbol column, 1d bars all share `t=04:00`). | col | type | notes | |-----|------|-------| | `market` | VARCHAR | pk + `date` | | `date` | DATE | a trading day (UTC) | | `session_open` | TIMESTAMPTZ | from auctions `o[0].t` only (nullable; not set by bars) | | `session_close` | TIMESTAMPTZ | from auctions `c[0].t` only (nullable; not set by bars) | | `open_price` | DOUBLE | opening auction price (nullable) | | `close_price` | DOUBLE | closing auction price (nullable) | | `source` | VARCHAR | `auctions` \| `bars` \| `manual` | | `updated_at` | TIMESTAMPTZ | | `features/` — TA + stochastic-process indicators, **wide** format, hive-partitioned with a `family` tier: `features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet` and `family=sp/symbol=AAPL.parquet`. Each row is one `t`, with one column per indicator. The partition columns (market/symbol/timeframe) come from the directory structure; the file itself stores `t` + indicator columns (e.g. `sma_5`, `sma_20`, `ema_12`, `ema_26`, `rsi_14` for `family=ta`; `sp_ou_halflife`, `sp_hmm_regime`, `sp_har_rv_5` for `family=sp`), all DOUBLE. Writes are Appender-based fresh writes per family (no read-merge-write cycle). | col | type | |-----|------| | `t` | TIMESTAMPTZ | | `sma_5`, `sma_20`, `ema_12`, `ema_26` | DOUBLE | | `rsi_14` | DOUBLE | | `macd`, `macd_signal`, `macd_hist` | DOUBLE | | `bb_upper`, `bb_middle`, `bb_lower` | DOUBLE | | `atr_14`, `adx_14`, `stoch_k`, `stoch_d` | DOUBLE | | `_feature_` | DOUBLE | `coverage.parquet` — **the cache index**: the exact loaded window per bar set. This is what makes direct hits fast. | col | type | |-----|------| | `market` | VARCHAR | | `timeframe` | VARCHAR | | `symbol` | VARCHAR | | `first_t` | TIMESTAMPTZ | | `last_t` | TIMESTAMPTZ | | `num_bars` | BIGINT | | `feed` | VARCHAR | | `adjustment` | VARCHAR | | `loaded_at` | TIMESTAMPTZ | | `updated_at` | TIMESTAMPTZ | `manifest.yaml` (plain text, not parquet) — lake config so reads/writes stay consistent: ```yaml lake_version: 1 default_market: US default_feed: iex # iex is the default; SIP is never used (requires license) default_adjustment: raw # raw|split|dividend|all — pick once per lake timezone: UTC features_lib: ta-lib ``` ## Read path (cache-first) ### DuckDB ```bash duckdb :memory: ``` ```sql -- hive glob adds market/timeframe/symbol columns automatically SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=*/timeframe=*/symbol=*.parquet'); ``` Canonical queries: ```sql -- past 2 months, 1d bars SELECT symbol, date, o, h, l, c, v, n, vw FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet') WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 MONTH ORDER BY t; -- past 2 days, 10m bars SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=10m/symbol=*.parquet') WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 DAY ORDER BY t; -- past 2 hours, 1m bars SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1m/symbol=*.parquet') WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 HOUR ORDER BY t; ``` Join with features (hive-partitioned, family=ta): ```sql SELECT b.t, b.c, f.sma_20, f.rsi_14 FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet') f ON f.t=b.t WHERE b.t >= now() - INTERVAL 2 MONTH; ``` Join with SP features (family=sp): ```sql SELECT b.t, b.c, sp.sp_ou_halflife, sp.sp_hmm_regime FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=sp/symbol=AAPL.parquet') sp ON sp.t=b.t WHERE b.t >= now() - INTERVAL 2 MONTH; ``` ### Apache Arrow / Python ```python import pyarrow.parquet as pq t = pq.read_table( "$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet", filters=[("symbol", "==", "AAPL")], ) df = t.to_pandas() ``` ## Verify lake data (duckdb CLI) Any parquet file in the lake can be inspected directly with the **DuckDB CLI** — no MCP call needed. Handy for confirming a `get_lake_bars`/`backfill_lake_calendar` write landed: ### Dependencies (duckdb + apache arrow) `duckdb` and `pyarrow` are declared in `tac-qlib/pyproject.toml` (installed into the repo `.venv` by `uv`). If the runtime venv lacks them, **lazy-install** rather than falling back to another SQL tool: ```bash uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow # or: uv pip install -e ./tac-qlib ``` Then re-check with `python -c "import duckdb, pyarrow"`. Only use the DuckDB CLI / pyarrow path when the MCP lake tools can't answer (see MCP-first policy above). ```bash duckdb :memory: "SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;" ``` Or interactively: ```bash duckdb :memory: SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10; ``` Quick checks: - **Bars written:** `SELECT count(*), min(t), max(t) FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet');` - **Coverage index:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/coverage.parquet') LIMIT 10;` - **Metadata:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/symbols.parquet') LIMIT 10;` - **Calendar:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/calendar.parquet') LIMIT 10;` > Note: `TAC_LAKE_DIR` is **mandatory** and must be an absolute path — do **not** use `~` or `$HOME` > (no fallback/expansion logic exists; a literal `~` is not expanded by shells/duckdb inside an env var). ## Lazy-load write path The core procedure when the requested range is **not** fully covered. Steps 1–9 are automated by the **`get_lake_bars`** lake tool (`lazy=true`) — the manual walk-through below documents what it does under the hood, and is the pattern to follow if writing the lake directly (DuckDB/pyarrow): 1. **Normalize the request.** `market`, lake `timeframe` (map back to MCP spelling), `symbols`, `start`, `end`. Decide `feed` and `adjustment` from `manifest.yaml` (or request overrides). Keep them fixed per lake — mixing feeds/adjustments corrupts history. 2. **Check coverage** (`coverage.parquet`). See decision table below. 3. **Compute the missing window(s).** e.g. request `[S,E]`, lake has `[S,M]` → fetch `(M,E]`; no row → fetch `[S,E]`. 4. **Call MCP `get_stock_bars`** (multi-symbol variant; comma-separated `symbols`): ```json {"symbols":"AAPL,MSFT","timeframe":"1Day","start":"2026-06-06T00:00:00Z","end":"2026-08-06T00:00:00Z","feed":"iex","adjustment":"raw","limit":10000} ``` Response: `{"bars": {"AAPL": [{t,o,h,l,c,v,n,vw}, …], …}, "next_page_token": "…"}`. Bars are sorted symbol-first, so a page may contain only some symbols — **loop with `next_page_token`** until `null`. 5. **Parse + normalize.** Keep `t,o,h,l,c,v,n,vw`; convert `t` to UTC `TIMESTAMPTZ`; add `date` for `1d`. 6. **Merge into the partition file** `$TAC_LAKE_DIR/market=/timeframe=/symbol=.parquet`: read the existing file, overlay the fetched bars keyed by canonical timestamp (new wins on duplicate `t`), write the full merged set to a tmp file, then atomically rename over the old one. This is safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history). 7. **Update `coverage.parquet`**: recompute `first_t`/`last_t`/`num_bars` from the **actual file contents** (not the fetched range) — a fetch that landed nothing must not widen the span into a hollow coverage. 8. **Update metadata**: upsert `symbols.parquet` (`first_seen` = the symbol's earliest bar in the lake) and `calendar.parquet` (distinct `date`s observed in bars, `source='bars'`; the bars path records only trading days — no session/prices). 9. **Return the requested range** from the lake (the read path above). ### Coverage decision table For a request `(market, timeframe, symbol, S, E)` against `coverage.parquet`: | coverage row | action | |--------------|--------| | missing | backfill whole `[S,E]` | | `first_t <= S` and `last_t >= E` | **direct hit** — read from lake, no fetch | | `first_t > S` | fetch `[S, first_t)` prefix, merge | | `last_t < E` | fetch `(last_t, E]` suffix, merge | | (with `calendar`) for `1d`: expected trading days `∈ [S,E]` == bars present | consider complete | Use `calendar.parquet` for the `1d` completeness check — a weekend/holiday gap is normal, so “no bar on Saturday” must **not** trigger a refetch. Also treat the in-progress current session carefully: an intraday `end=now` should not trigger a refetch loop on the forming bar. ### Bulk loading multiple symbols For loading bars + TA + SP features for many symbols at once, use **`load_lake_symbols`**. It runs in a background thread and returns immediately with a `job_id`: ```json {"symbols":"AAPL,MSFT,GOOGL,AMZN","timeframe":"1d","start":"2020-01-03","end":"2026-08-15"} ``` Response: ```json {"job_id":"load-20260815-143022","status":"started","symbols":["AAPL","MSFT","GOOGL","AMZN"],"note":"load running in background -- poll load_lake_status with this job_id to track progress"} ``` Poll progress with **`load_lake_status`**: ```json {"job_id":"load-20260815-143022"} ``` Response (while running): ```json {"job_id":"load-20260815-143022","status":"running","total_symbols":4,"processed":2,"results":[...]} ``` Response (when done): ```json {"job_id":"load-20260815-143022","status":"completed","total_symbols":4,"processed":4,"results":[...],"completed_at":"2026-08-15T14:35:00Z"} ``` Each entry in `results` contains per-symbol `fetch_start` (the date the load started from), `bars_count`, `bars_source`, `ta_count`, `ta_columns`, `sp_count`, `sp_columns`, and any `*_error` fields. ### Calendar gap — why `get_stock_auctions` The MCP tool surface has **no calendar endpoint**, but the coverage check needs to know which days are trading days before bars exist. The **auctions** tool fills this gap: - `get_stock_auctions` only accepts `feed: "sip"` (SIP is the only valid feed for auctions). - Auction records exist **only on trading days** → the set of distinct dates `d` across symbols is the trading-day set. - Response shape: `{"auctions": {"AAPL": [{"d":"2026-06-09","o":[{t,x,p,c}…],"c":[{t,x,p,c}…]}, …]}, "next_page_token": "…"}` — `o` = opening auctions, `c` = closing auctions. ```json {"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"} ``` Usage: to seed/enrich `calendar.parquet` for a range **before** loading bars, call the **`backfill_lake_calendar`** lake tool (it loops `get_stock_auctions` internally across symbols and pages, inserts one row per distinct `d` with `source='auctions'`, `session_open`/`open_price` from `o[0]`, `session_close`/`close_price` from `c[0]`, and reports `calendar_days_added`). Direct call equivalent: ```json {"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"} ``` Cheap single-day confirmation for “was this a trading day?” and first pass of daily open/close. Intraday bars and per-symbol coverage still come from `get_stock_bars`. ## Operations notes - **Rate limits:** Alpaca data API ~200 req/min. On `429` back off (exponential, start 1s) and retry. Batch symbols in one call, but page through `next_page_token`. - **Atomicity:** write parquet to a `..tmp` then `rename()`; readers never see partial files. Apply the same pattern to metadata upserts. - **Consistency:** one `feed` + one `adjustment` per lake (record in `manifest.yaml`). Refetching a window with a different feed/adjustment would silently corrupt merged history. - **Feed / history limits (Alpaca):** SIP is never used (requires license). IEX goes back to **2020-07-27** for daily bars. For earlier data, Yahoo Finance fills gaps automatically (daily bars only). The `TAC_LAKE_START_DATE` (default `2000-01-03`) controls the earliest date requested; Yahoo provides data back to ~1970 for most symbols. - **Dedup / merge:** bar upserts read the existing file, merge by canonical timestamp (new wins on duplicate `t`), and write the full set atomically — safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history). Feature persistence is a fresh write per family, so no dedup needed. - **Features** are derived from the lake bars (compute after bars are persisted, keyed `(market, symbol, timeframe, t)`), so indicator history stays aligned with bar history. ## End-to-end example (1d, 2 months, AAPL) 1. `coverage.parquet` has no `(US,1d,AAPL)` row → backfill. 2. Seed calendar: `backfill_lake_calendar` `{"symbols":"AAPL","start":…,"end":…}` (loops `get_stock_auctions`) → `calendar.parquet` trading days. 3. `get_lake_bars` `{"symbols":"AAPL","timeframe":"1d","start":…,"end":…,"lazy":true,"quiet":true}` → auto backfills the window, persists bars, updates coverage/symbols/calendar, returns a `{count, first_t, last_t}` summary instead of the bar rows. 4. Answer: DuckDB `SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') WHERE t >= now() - INTERVAL 2 MONTH` (or `get_lake_bars` again). 5. **Next identical request is a direct hit** (`source:"lake"`) from step 1’s decision table — no Alpaca fetch.