book: scaffold + ch00 (execution trail as spine) — evidence exp 8-31, round 3

This commit is contained in:
TradeAC Book Agent
2026-08-18 22:35:23 +00:00
commit c93424e76c
83 changed files with 17676 additions and 0 deletions
+161
View File
@@ -0,0 +1,161 @@
---
name: tradeac-alpaca
description: Guide agents to call tac-engine MCP tools for Alpaca trading and market data (news, corporate actions, screener, FX, stocks, options, realtime streams) via MCP Inspector, Cursor, or other clients.
---
# tradeac-alpaca
Use **tac-engine** MCP tools — not raw Alpaca REST — for brokerage ops and market data. Same tool surface will back TradeAC’s Next.js UI later.
## MCP-first policy
- **Prefer the MCP tools registered in this session** (`tac-engine` server, tools listed below) over writing scripts that reimplement them. If a tool exists, call it directly — do not reinvent it with curl/bash/python (raw Alpaca REST, hand-rolled pagination, own JSON-RPC clients).
- **NEVER script directly against the MCP server** (spawning the binary, talking stdio JSON-RPC, or driving it via bash/curl) unless the MCP tool surface genuinely can't do the job — and in that case **stop and ask the user to confirm first** before writing the script.
- If a direct Alpaca call is needed (e.g. an endpoint with no tool), say so and let the user confirm the approach; otherwise keep everything on the MCP surface.
## Hosts / env
| Purpose | Env | Default |
|---------|-----|---------|
| Trading REST | `APCA_BASE_URL` | `https://paper-api.alpaca.markets` |
| Market data REST | `APCA_DATA_BASE_URL` | `https://data.alpaca.markets` |
| Market data WS | `APCA_STREAM_BASE_URL` | `wss://stream.data.alpaca.markets` |
| Auth | `APCA_API_KEY_ID`, `APCA_API_SECRET_KEY` | required |
```bash
cargo build --release
./target/release/tac-engine
```
## Secrets policy
- NEVER write secrets into files: API keys (`APCA_API_KEY_ID`/`APCA_API_SECRET_KEY`), DB passwords, OAuth tokens, or credential-bearing URLs in scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## MCP clients
### MCP Inspector
```bash
npx @modelcontextprotocol/inspector /absolute/path/to/tac-engine/target/release/tac-engine
```
Connect → `tools/list` → `tools/call` with JSON `arguments`.
### Cursor
```json
{
"mcpServers": {
"tac-engine": {
"command": "/absolute/path/to/tac-engine/target/release/tac-engine",
"env": {
"APCA_API_KEY_ID": "${APCA_API_KEY_ID}",
"APCA_API_SECRET_KEY": "${APCA_API_SECRET_KEY}"
}
}
}
}
```
stdio is NDJSON JSON-RPC; logs on stderr.
## Safety
- Paper trading by default (`APCA_BASE_URL`).
- Mutating trading tools: `place_order`, `close_*`, `cancel_*`, watchlist writes.
- `subscribe_market_stream` opens a **short-lived** WebSocket (samples then disconnects). Most Alpaca plans allow **one** concurrent stream — close other clients first.
## Tool catalog
Lake tools (`get_lake_bars`, `get_lake_ta`, `get_lake_sp`, `get_lake_features`, `backfill_lake_calendar`, …) live under **`tradeac-lake`** (see `tac-engine/skills/tradeac-lake/SKILL.md`). When a lake call's purpose is **backfilling/persisting** (not reading the payload), pass `"quiet": true` so the tool returns a summary (`count`/`first_t`/`last_t`/`source`/`columns`) instead of echoing back the full bar/feature rows.
### Trading (brokerage account)
Account: `get_account`, `get_portfolio_history`, `list_account_activities`, `get_account_activities_by_type`
Assets master: `list_assets`, `get_asset`
Watchlists / positions / orders: `list_*`, `get_*`, `create_*`, `place_order`, `close_*`, `cancel_*`
### Market data — news & corporate actions
| Tool | Notes |
|------|------|
| `get_news` | optional `symbols`, `start`/`end`, `limit`, `include_content` |
| `get_corporate_actions` | optional `symbols`, `types`, `start`/`end`, `data_quality` |
### Screener
| Tool | Notes |
|------|------|
| `get_most_actives` | optional `by`=`volume`\|`trades`, `top` |
| `get_market_movers` | required `market_type`=`stocks`\|`crypto`, optional `top` |
### FX
| Tool | Notes |
|------|------|
| `get_forex_latest_rates` | required `currency_pairs` e.g. `USDJPY,EURUSD` |
| `get_forex_rates` | historical; optional `timeframe`, `start`, `end` |
### Stocks
| Tool | Notes |
|------|------|
| `get_stock_bars` / `get_stock_bars_single` | historical; needs `timeframe` |
| `get_stock_latest_bars` | latest minute bars |
| `get_stock_quotes` / `get_stock_latest_quotes` | quotes |
| `get_stock_trades` / `get_stock_latest_trades` | trades |
| `get_stock_snapshots` / `get_stock_snapshot` | trade+quote+bars |
| `get_stock_auctions` | auctions |
Multi-symbol tools take comma-separated `symbols`. Optional `feed` (`iex`/`sip`), `limit`, `page_token`, …
### Options
| Tool | Notes |
|------|------|
| `get_option_bars` | historical bars for contract symbols |
| `get_option_latest_quotes` / `get_option_latest_trades` | latest |
| `get_option_trades` | historical trades |
| `get_option_snapshots` | contracts |
| `get_option_chain` | underlying + filters (`type`, strikes, expiration) |
| `get_option_meta_conditions` / `get_option_meta_exchanges` | code maps |
### Realtime stream sampling
| Tool | Notes |
|------|------|
| `subscribe_market_stream` | `stream`=`stocks`\|`options`\|`news`\|`test`; optional `feed`; channels `trades`/`quotes`/`bars`/`news` as CSV symbols; `duration_secs` (1–30), `max_messages` (1–200) |
Examples:
```json
{"stream":"test","duration_secs":5,"max_messages":20}
```
```json
{"stream":"stocks","feed":"iex","quotes":"AAPL,MSFT","duration_secs":5}
```
```json
{"stream":"news","news":"*","duration_secs":8,"max_messages":30}
```
```json
{"stream":"options","feed":"indicative","quotes":"AAPL250117C00200000","duration_secs":5}
```
## Example workflows
1. **Dashboard:** `get_account` → `list_positions` → `get_stock_snapshots` (`symbols` from positions)
2. **Research:** `get_news` → `get_corporate_actions` → `get_stock_bars`
3. **Screener → trade (paper):** `get_most_actives` → `get_stock_snapshot` → `place_order`
4. **Options:** `get_option_chain` (`underlying_symbol=AAPL`) → `get_option_latest_quotes`
5. **Live sample:** `subscribe_market_stream` with `stream=test` first, then stocks/news
## Protocol
- rmcp **3.1** / MCP **2026-07-28**, stdio
- Prefer these MCP tool names/args over calling Alpaca hosts directly from agents/UI
+417
View File
@@ -0,0 +1,417 @@
---
name: tradeac-lake
description: Guide agents to build and query the TradeAC parquet+DuckDB data lake on the local filesystem — hive-partitioned bar store (market/timeframe/symbol) plus symbols, watchlist, calendar, features and coverage metadata — with lazy backfill from the tac-engine MCP get_stock_bars tool (tradeac-alpaca skill).
---
# tradeac-lake
Local-first market data lake: **Apache Parquet** files on disk, consumed with **DuckDB** (or Apache Arrow). Bar data is the core payload; the lake also keeps small metadata parquet files (symbols, watchlist, calendar, features, coverage) at the lake root.
Reading is a **cache-first** pattern: if the requested range is already in the lake, serve it directly from parquet; otherwise **lazy-load** the missing window via the tac-engine MCP `get_stock_bars` tool (see `tac-engine/skills/tradeac-alpaca/SKILL.md`), persist it, update metadata, then return.
## MCP-first policy
- **Prefer the tac-engine lake MCP tools** (`get_lake_bars`, `get_lake_ta`, `get_lake_sp`, `get_lake_features`, `get_lake_status`, `get_lake_coverage`, `get_lake_calendar`, `backfill_lake_calendar`, …) whenever they cover the need. They handle coverage checks, lazy backfill, feed fallback, metadata updates and pagination for you — do not reimplement that in DuckDB/pyarrow scripts.
- **Direct parquet reads are only for verification** (DuckDB CLI / pyarrow snippets below) or when no lake tool covers the query (e.g. an arbitrary ad-hoc SQL join). Keep hand-rolled lake *writes* off the happy path — the write path is what the MCP tools automate.
- **NEVER script directly against the MCP server** (spawning the engine binary, stdio JSON-RPC, bash/curl) unless a tool genuinely can't do the job — then **stop and ask the user to confirm first**.
- The engine bundles its own DuckDB; direct verification only needs the `duckdb` CLI or a Python venv with `duckdb` + `pyarrow` (see dependencies below).
## Env / root
| Var | Default | Purpose |
|-----|---------|---------|
| `TAC_LAKE_DIR` | **required** (no default) | lake root on the local filesystem. Local dev: absolute path (e.g. `/home/data/lake`). |
```bash
export TAC_LAKE_DIR=/path/to/lake
mkdir -p "$TAC_LAKE_DIR"
```
## Secrets policy
- NEVER write secrets into files: API keys, DB passwords, OAuth tokens, or credential-bearing URLs (`DATABASE_URL`, `APCA_*`) in scripts, configs, notes or committed code.
- NEVER read `*.env` / `.env.*` directly (`cat`/`tail`/`grep`/`sed`/`head` on `.env`). That pulls secrets into this session and leaks them to any agent sharing it.
- When a tool or command needs an env var, ASK the user to set it in the environment (shell/container env, or the user-owned `.env`) and reference it by name (`$VAR`), never by value. If it's missing, report which variable is required instead of reading it yourself.
- If you find a committed secret, flag it, remove it, and replace it with a placeholder.
## Lake layout
Hive partition convention, partitioned by `market`, `timeframe`, `symbol`. Metadata parquet files live alongside the partition dirs at the lake root.
```
$TAC_LAKE_DIR/
├── market=US/
│ └── timeframe=1d/
│ ├── symbol=AAPL.parquet
│ ├── symbol=MSFT.parquet
│ └── ...
│ └── timeframe=10m/
│ └── symbol=AAPL.parquet
├── market=CRYPTO/... # optional: BTC/USD etc.
├── features/ # TA + SP indicators, hive-partitioned with family tier
│ └── market=US/
│ └── timeframe=1d/
│ ├── family=ta/
│ │ └── symbol=AAPL.parquet # TA indicators (sma, rsi, macd, ...)
│ └── family=sp/
│ └── symbol=AAPL.parquet # Stochastic-process features (ou, hmm, har, ...)
├── symbols.parquet # asset master seen/known to the lake
├── watchlist.parquet # watchlists
├── calendar.parquet # trading days per market (coverage ground truth)
├── coverage.parquet # per (market,timeframe,symbol) loaded window
└── manifest.yaml # lake config: feed, adjustment, timezone
```
Convention: **one parquet file per symbol per timeframe per family** under the partition dirs. Bar upserts **merge by canonical timestamp** (read existing file → overlay new bars → write the full set atomically); feature persistence is a **fresh write** per family (Appender-based, no merge with existing).
## Lake MCP tools (tac-engine)
The tac-engine MCP server exposes a **lake tools** category that wraps the read/write paths below. `symbols`/`market` are uppercased, `timeframe` is normalized to lake spelling, and `start`/`end` accept `YYYY-MM-DD` or RFC-3339 (default `end=now`, `start=end-30d`).
| Tool | Purpose |
|------|---------|
| `get_lake_bars` | Cache-first bars: `{market?, symbols, timeframe, start?, end?, feed?, adjustment?, lazy?, quiet?}`. Feed defaults to `iex`; SIP is never used (requires license). For daily bars, Yahoo Finance fills gaps before Alpaca's earliest available date. `lazy=true` (default) backfills missing windows via Alpaca and persists; `lazy=false` reads the lake only. Returns `{request, source, bars: {SYM: [{t,o,h,l,c,v,n,vw}]}}`; `source` is `lake` (complete hit), `partial` (present but missing windows and `lazy=false`), or `fetched` (gaps backfilled). With `quiet=true` returns `{request, source, summary: {SYM: {count, first_t, last_t}}}` instead of the bar rows — use for backfill-to-lake jobs. Bar writes **merge by timestamp** (read existing file, overlay new bars, write the full set atomically) — safe for both tail appends and leading-gap backfills. Coverage/symbols/calendar metadata are reconciled against the actual file on each write. |
| `get_lake_ta` | Compute + optionally persist indicators: `{market?, symbol, timeframe, start?, end?, indicators?, persist?, quiet?}`. `indicators` is comma-separated, default all: `sma_5,sma_20,ema_12,ema_26,rsi_14,macd,bb,atr_14,adx_14`. Lookback is pulled automatically; returned rows cover `[start,end]`. When `persist=true`, features are written to `features/market=*/timeframe=*/family=ta/symbol=*.parquet`. With `quiet=true` returns `{count, columns, persisted}` instead of the feature rows — use when the goal is persisting indicators. |
| `get_lake_sp` | Compute + optionally persist **stochastic-process features** (`sp_*` columns) from lake bars: `{market?, symbol, timeframe, start?, end?, fit_end?, families?, persist?, quiet?}`. Rust port of `sp_features.py` on the stochastic-rs stack. `families` is comma-separated, default all: `ou,hmm,jump,har,trend,hurst,signature,moments`. In addition to `sp_rv*`/`sp_vol_ratio_*` the `har` family also emits `sp_rv_ac1` (RV lag-1 autocorr) + `sp_rv_cv_22` (RV coefficient of variation); `jump` also emits `sp_max_up`/`sp_max_down` (signed max-move asymmetry); `signature` also emits the lag-5 level-2 cross terms `sp_sig_level2_{lead_lag,lag_lead}_5`; and `moments` emits the scale-free realized skewness/kurtosis (`sp_rskew_{5,22}`, `sp_rkurt_{5,22}`; a 1-day window is undefined) and downside semi-variance (`sp_dsv_{1,5,22}`, `sp_dsv_ratio_{1,5,22}`) via stochastic-rs `realized`. `fit_end` limits the 2-state Gaussian-HMM fit window (no lookahead; posteriors still cover the whole window). On `persist=true`, features are written to `features/market=*/timeframe=*/family=sp/symbol=*.parquet`. With `quiet=true` returns `{count, sp_columns, persisted}` instead of the feature rows. Deferred (not in stochastic-rs): `garch`, `entropy`, `catch22`. |
| `get_lake_features` | Read persisted TA + SP features (hive-partitioned `features/` dir, families `ta` and `sp` merged by timestamp): `{market?, symbol?, timeframe?, start?, end?, quiet?}`. With `quiet=true` returns `{count, columns}` instead of the full feature rows. |
| `get_lake_symbols` | Read `symbols.parquet` asset master; optional `{symbol?}` filter. |
| `get_lake_watchlist` | Read `watchlist.parquet`. |
| `get_lake_calendar` | Read `calendar.parquet` trading days: `{market?, start?, end?}`. |
| `get_lake_coverage` | Read `coverage.parquet` cache index: `{market?, timeframe?, symbol?}`. |
| `get_lake_status` | Lake root, `manifest.yaml`, bar partition inventory and metadata file sizes. |
| `rebuild_lake_symbol` | Delete bar + feature parquet files and re-fetch from `TAC_LAKE_START_DATE` (default `2000-01-03`) for a single symbol: `{market?, symbol, timeframe, feed?, adjustment?}`. Resets coverage so the next `get_lake_bars` call re-downloads the full history. Use after changing `TAC_LAKE_START_DATE` or to fix stale/corrupt data. |
| `load_lake_symbols` | **Bulk-load + persist** bars + TA + SP features for a comma-separated list of symbols. Runs in a background thread and returns immediately with a `job_id`: `{market?, symbols, timeframe, start?, end?, feed?, adjustment?, indicators?, families?}`. Per-symbol start is computed automatically from lake coverage: if the lake has no data or `first_t > TAC_LAKE_START_DATE`, fetches from `TAC_LAKE_START_DATE` (default 2000-01-03); if `first_t <= TAC_LAKE_START_DATE`, fetches only from `last_t` (tail refresh). TA/SP features are always computed over the full `TAC_LAKE_START_DATE` to `end` range. Poll `load_lake_status` with the returned `job_id` to track progress. |
| `load_lake_status` | Query the status of a background bulk-load job: `{job_id}`. Returns `{job_id, status, total_symbols, processed, results, error, started_at, completed_at}` where `status` is `running`, `completed`, or `failed`, and `results` contains per-symbol bar counts, TA/SP column counts, and any errors. |
| `backfill_lake_calendar` | **Gap-fill tool**: seed/enrich `calendar.parquet` from Alpaca historical auctions (feed=iex; records exist only on trading days): `{market?, symbols, start?, end?}`. Returns `{market, symbols, start, end, calendar_days_added}`. Call this before lazy bar loads so the `1d` completeness check knows the expected trading-day set. |
| `validate_lake_dataset` | **Pre-workflow quality gate**: `{market?, timeframe?, symbols?, start?, end?}`. Scans every symbol in coverage (or a comma-separated `symbols` subset) and reports `verdict: OK/WARNINGS/ERRORS` plus per-symbol issues. Catches the failure modes qlib silently tolerates: **all-NaN feature columns** (would be dropped by `DropAllNaN` — the model trains on fewer features without notice), **missing TA/SP feature files**, **hollow coverage / stale date ranges** (coverage claims a wide span but the bar file is empty/truncated/sparse), **stale coverage** (first/last/num_bars vs the actual file), and **partition misalignment** (flat-layout feature orphans the family=ta|sp consumers can't see). Pass `start`/`end` to also check feature-vs-bar row alignment and per-column all-NaN status in that window. Call before `rd_run_workflow` / `rd_train` to fail fast instead of training on silent data holes. |
Example:
```json
{"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true}
```
→ `{"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "bars": {"AAPL": [{...}], "MSFT": [{...}]}}`
## Quiet mode
`get_lake_bars`, `get_lake_ta`, `get_lake_sp` and `get_lake_features` accept `"quiet": true`. When the point of the call is **writing to the lake** (backfill/fetch bars, compute + persist indicators or `sp_*` features), use `quiet: true` — the tool still performs the full backfill / computation / persist, but returns a **summary** instead of echoing back the potentially huge payload (thousands of bar rows / feature rows). Full-row output (`bars` / `features`) is the default, so requests that *need* the data to read it must leave `quiet` unset/false.
| Tool | `quiet: true` response |
|------|------------------------|
| `get_lake_bars` | `{request, source: {SYM: lake\|partial\|fetched}, summary: {SYM: {count, first_t, last_t}}}` |
| `get_lake_ta` | `{market, symbol, timeframe, start, end, count, columns, persisted}` |
| `get_lake_sp` | `{market, symbol, timeframe, start, end, fit_end, count, sp_columns, persisted}` |
| `get_lake_features` | `{count, columns}` |
Backfill-to-lake job (no payload echoed):
```json
{"symbols": "AAPL,MSFT", "timeframe": "1d", "start": "2026-06-06", "lazy": true, "quiet": true}
```
→ `{"request": {...}, "source": {"AAPL": "fetched", "MSFT": "lake"}, "summary": {"AAPL": {"count": 44, "first_t": "2026-06-06T04:00:00Z", "last_t": "2026-08-05T04:00:00Z"}, "MSFT": {...}}}`
Persist indicators to the lake (summary only):
```json
{"symbol": "AAPL", "timeframe": "1d", "indicators": "sma_5,sma_20,rsi_14", "persist": true, "quiet": true}
```
→ `{"market": "US", "symbol": "AAPL", "timeframe": "1d", "count": 44, "columns": ["sma_5","sma_20","rsi_14"], "persisted": true}`
```json
{"symbols": "AAPL,MSFT", "start": "2026-06-06"}
```
→ `{"market": "US", "symbols": ["AAPL","MSFT"], "start": ..., "end": ..., "calendar_days_added": 44}`
## Conventions
- `market`: `US` (equities), `CRYPTO`, `FOREX`. Uppercase.
- `timeframe`: normalized lake name — lowercase, `1m 5m 10m 15m 30m 1h 2h 4h 1d 1w 1M`. The MCP tool spells them differently; always map:
| Lake | MCP `timeframe` | Lake | MCP `timeframe` |
|------|-----------------|------|-----------------|
| `1m` | `1Min` | `2h` | `2Hour` |
| `5m` | `5Min` | `4h` | `4Hour` |
| `10m` | `10Min` | `1d` | `1Day` |
| `15m` | `15Min` | `1w` | `1Week` |
| `30m` | `30Min` | `1M` | `1Month` |
| `1h` | `1Hour` | | |
- `symbol`: uppercase, e.g. `AAPL`. Hyphens/`.` in special symbols (e.g. `BRK-B`, `SPY`) are valid filenames; avoid `/` and spaces.
- All timestamps stored as **UTC** instants (`TIMESTAMPTZ`). Alpaca returns RFC-3339 UTC; normalize on write.
- `1d` bars: `t` is the session date at `04:00Z` (midnight ET — Alpaca stamps daily bars at `04:00:00Z`); also store a `date` column (`CAST(t AS DATE)`, UTC) for calendar joins. A date-only `end` (e.g. `2026-08-05`) is treated as **inclusive of the whole end day**, so the end-day bar is not dropped.
## Bar parquet schema (`market=…/timeframe=…/symbol=….parquet`)
| col | type | source field |
|-----|------|--------------|
| `t` | TIMESTAMPTZ | bar `t` (UTC) |
| `o` | DOUBLE | `o` |
| `h` | DOUBLE | `h` |
| `l` | DOUBLE | `l` |
| `c` | DOUBLE | `c` |
| `v` | BIGINT | `v` |
| `n` | BIGINT | `n` |
| `vw` | DOUBLE | `vw` |
Partition columns `market`/`timeframe`/`symbol` are derived from the path; DuckDB exposes them automatically when reading a hive glob.
## Metadata parquet files
All written with DuckDB `COPY … (FORMAT PARQUET)` from in-memory `SELECT`, or `pyarrow.parquet`.
`symbols.parquet`
| col | type | notes |
|-----|------|-------|
| `symbol` | VARCHAR (pk) |
| `name` | VARCHAR |
| `asset_class` | VARCHAR |
| `exchange` | VARCHAR |
| `tradable` | BOOLEAN |
| `status` | VARCHAR |
| `first_seen` | TIMESTAMPTZ | the symbol's earliest bar in the lake (its first trading date), not the load timestamp |
| `updated_at` | TIMESTAMPTZ | |
`watchlist.parquet`
| col | type |
|-----|------|
| `watchlist_id` | VARCHAR |
| `name` | VARCHAR |
| `symbol` | VARCHAR |
| `added_at` | TIMESTAMPTZ |
| `updated_at` | TIMESTAMPTZ |
`calendar.parquet` — the trading-day ground truth per market (see “Calendar gap” below). Bars only seed which dates are trading days; per-symbol prices/session times are NOT attributed by the bars path (no symbol column, 1d bars all share `t=04:00`).
| col | type | notes |
|-----|------|-------|
| `market` | VARCHAR | pk + `date` |
| `date` | DATE | a trading day (UTC) |
| `session_open` | TIMESTAMPTZ | from auctions `o[0].t` only (nullable; not set by bars) |
| `session_close` | TIMESTAMPTZ | from auctions `c[0].t` only (nullable; not set by bars) |
| `open_price` | DOUBLE | opening auction price (nullable) |
| `close_price` | DOUBLE | closing auction price (nullable) |
| `source` | VARCHAR | `auctions` \| `bars` \| `manual` |
| `updated_at` | TIMESTAMPTZ | |
`features/` — TA + stochastic-process indicators, **wide** format, hive-partitioned with a `family` tier: `features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet` and `family=sp/symbol=AAPL.parquet`. Each row is one `t`, with one column per indicator. The partition columns (market/symbol/timeframe) come from the directory structure; the file itself stores `t` + indicator columns (e.g. `sma_5`, `sma_20`, `ema_12`, `ema_26`, `rsi_14` for `family=ta`; `sp_ou_halflife`, `sp_hmm_regime`, `sp_har_rv_5` for `family=sp`), all DOUBLE. Writes are Appender-based fresh writes per family (no read-merge-write cycle).
| col | type |
|-----|------|
| `t` | TIMESTAMPTZ |
| `sma_5`, `sma_20`, `ema_12`, `ema_26` | DOUBLE |
| `rsi_14` | DOUBLE |
| `macd`, `macd_signal`, `macd_hist` | DOUBLE |
| `bb_upper`, `bb_middle`, `bb_lower` | DOUBLE |
| `atr_14`, `adx_14`, `stoch_k`, `stoch_d` | DOUBLE |
| `_feature_<name>` | DOUBLE |
`coverage.parquet` — **the cache index**: the exact loaded window per bar set. This is what makes direct hits fast.
| col | type |
|-----|------|
| `market` | VARCHAR |
| `timeframe` | VARCHAR |
| `symbol` | VARCHAR |
| `first_t` | TIMESTAMPTZ |
| `last_t` | TIMESTAMPTZ |
| `num_bars` | BIGINT |
| `feed` | VARCHAR |
| `adjustment` | VARCHAR |
| `loaded_at` | TIMESTAMPTZ |
| `updated_at` | TIMESTAMPTZ |
`manifest.yaml` (plain text, not parquet) — lake config so reads/writes stay consistent:
```yaml
lake_version: 1
default_market: US
default_feed: iex # iex is the default; SIP is never used (requires license)
default_adjustment: raw # raw|split|dividend|all — pick once per lake
timezone: UTC
features_lib: ta-lib
```
## Read path (cache-first)
### DuckDB
```bash
duckdb :memory:
```
```sql
-- hive glob adds market/timeframe/symbol columns automatically
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=*/timeframe=*/symbol=*.parquet');
```
Canonical queries:
```sql
-- past 2 months, 1d bars
SELECT symbol, date, o, h, l, c, v, n, vw
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet')
WHERE symbol = 'AAPL'
AND t >= now() - INTERVAL 2 MONTH
ORDER BY t;
-- past 2 days, 10m bars
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=10m/symbol=*.parquet')
WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 DAY ORDER BY t;
-- past 2 hours, 1m bars
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1m/symbol=*.parquet')
WHERE symbol = 'AAPL' AND t >= now() - INTERVAL 2 HOUR ORDER BY t;
```
Join with features (hive-partitioned, family=ta):
```sql
SELECT b.t, b.c, f.sma_20, f.rsi_14
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b
LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=ta/symbol=AAPL.parquet') f
ON f.t=b.t
WHERE b.t >= now() - INTERVAL 2 MONTH;
```
Join with SP features (family=sp):
```sql
SELECT b.t, b.c, sp.sp_ou_halflife, sp.sp_hmm_regime
FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') b
LEFT JOIN read_parquet('$TAC_LAKE_DIR/features/market=US/timeframe=1d/family=sp/symbol=AAPL.parquet') sp
ON sp.t=b.t
WHERE b.t >= now() - INTERVAL 2 MONTH;
```
### Apache Arrow / Python
```python
import pyarrow.parquet as pq
t = pq.read_table(
"$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=*.parquet",
filters=[("symbol", "==", "AAPL")],
)
df = t.to_pandas()
```
## Verify lake data (duckdb CLI)
Any parquet file in the lake can be inspected directly with the **DuckDB CLI** — no MCP call needed. Handy for confirming a `get_lake_bars`/`backfill_lake_calendar` write landed:
### Dependencies (duckdb + apache arrow)
`duckdb` and `pyarrow` are declared in `tac-qlib/pyproject.toml` (installed into the repo `.venv` by `uv`). If the runtime venv lacks them, **lazy-install** rather than falling back to another SQL tool:
```bash
uv pip install --python $VIRTUAL_ENV/bin/python duckdb pyarrow # or: uv pip install -e ./tac-qlib
```
Then re-check with `python -c "import duckdb, pyarrow"`. Only use the DuckDB CLI / pyarrow path when the MCP lake tools can't answer (see MCP-first policy above).
```bash
duckdb :memory: "SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;"
```
Or interactively:
```bash
duckdb :memory:
SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') LIMIT 10;
```
Quick checks:
- **Bars written:** `SELECT count(*), min(t), max(t) FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet');`
- **Coverage index:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/coverage.parquet') LIMIT 10;`
- **Metadata:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/symbols.parquet') LIMIT 10;`
- **Calendar:** `SELECT * FROM read_parquet('$TAC_LAKE_DIR/calendar.parquet') LIMIT 10;`
> Note: `TAC_LAKE_DIR` is **mandatory** and must be an absolute path — do **not** use `~` or `$HOME`
> (no fallback/expansion logic exists; a literal `~` is not expanded by shells/duckdb inside an env var).
## Lazy-load write path
The core procedure when the requested range is **not** fully covered. Steps 1–9 are automated by the **`get_lake_bars`** lake tool (`lazy=true`) — the manual walk-through below documents what it does under the hood, and is the pattern to follow if writing the lake directly (DuckDB/pyarrow):
1. **Normalize the request.** `market`, lake `timeframe` (map back to MCP spelling), `symbols`, `start`, `end`. Decide `feed` and `adjustment` from `manifest.yaml` (or request overrides). Keep them fixed per lake — mixing feeds/adjustments corrupts history.
2. **Check coverage** (`coverage.parquet`). See decision table below.
3. **Compute the missing window(s).** e.g. request `[S,E]`, lake has `[S,M]` → fetch `(M,E]`; no row → fetch `[S,E]`.
4. **Call MCP `get_stock_bars`** (multi-symbol variant; comma-separated `symbols`):
```json
{"symbols":"AAPL,MSFT","timeframe":"1Day","start":"2026-06-06T00:00:00Z","end":"2026-08-06T00:00:00Z","feed":"iex","adjustment":"raw","limit":10000}
```
Response: `{"bars": {"AAPL": [{t,o,h,l,c,v,n,vw}, …], …}, "next_page_token": "…"}`. Bars are sorted symbol-first, so a page may contain only some symbols — **loop with `next_page_token`** until `null`.
5. **Parse + normalize.** Keep `t,o,h,l,c,v,n,vw`; convert `t` to UTC `TIMESTAMPTZ`; add `date` for `1d`.
6. **Merge into the partition file** `$TAC_LAKE_DIR/market=<m>/timeframe=<tf>/symbol=<s>.parquet`: read the existing file, overlay the fetched bars keyed by canonical timestamp (new wins on duplicate `t`), write the full merged set to a tmp file, then atomically rename over the old one. This is safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history).
7. **Update `coverage.parquet`**: recompute `first_t`/`last_t`/`num_bars` from the **actual file contents** (not the fetched range) — a fetch that landed nothing must not widen the span into a hollow coverage.
8. **Update metadata**: upsert `symbols.parquet` (`first_seen` = the symbol's earliest bar in the lake) and `calendar.parquet` (distinct `date`s observed in bars, `source='bars'`; the bars path records only trading days — no session/prices).
9. **Return the requested range** from the lake (the read path above).
### Coverage decision table
For a request `(market, timeframe, symbol, S, E)` against `coverage.parquet`:
| coverage row | action |
|--------------|--------|
| missing | backfill whole `[S,E]` |
| `first_t <= S` and `last_t >= E` | **direct hit** — read from lake, no fetch |
| `first_t > S` | fetch `[S, first_t)` prefix, merge |
| `last_t < E` | fetch `(last_t, E]` suffix, merge |
| (with `calendar`) for `1d`: expected trading days `∈ [S,E]` == bars present | consider complete |
Use `calendar.parquet` for the `1d` completeness check — a weekend/holiday gap is normal, so “no bar on Saturday” must **not** trigger a refetch. Also treat the in-progress current session carefully: an intraday `end=now` should not trigger a refetch loop on the forming bar.
### Bulk loading multiple symbols
For loading bars + TA + SP features for many symbols at once, use **`load_lake_symbols`**. It runs in a background thread and returns immediately with a `job_id`:
```json
{"symbols":"AAPL,MSFT,GOOGL,AMZN","timeframe":"1d","start":"2020-01-03","end":"2026-08-15"}
```
Response:
```json
{"job_id":"load-20260815-143022","status":"started","symbols":["AAPL","MSFT","GOOGL","AMZN"],"note":"load running in background -- poll load_lake_status with this job_id to track progress"}
```
Poll progress with **`load_lake_status`**:
```json
{"job_id":"load-20260815-143022"}
```
Response (while running):
```json
{"job_id":"load-20260815-143022","status":"running","total_symbols":4,"processed":2,"results":[...]}
```
Response (when done):
```json
{"job_id":"load-20260815-143022","status":"completed","total_symbols":4,"processed":4,"results":[...],"completed_at":"2026-08-15T14:35:00Z"}
```
Each entry in `results` contains per-symbol `fetch_start` (the date the load started from), `bars_count`, `bars_source`, `ta_count`, `ta_columns`, `sp_count`, `sp_columns`, and any `*_error` fields.
### Calendar gap — why `get_stock_auctions`
The MCP tool surface has **no calendar endpoint**, but the coverage check needs to know which days are trading days before bars exist. The **auctions** tool fills this gap:
- `get_stock_auctions` only accepts `feed: "sip"` (SIP is the only valid feed for auctions).
- Auction records exist **only on trading days** → the set of distinct dates `d` across symbols is the trading-day set.
- Response shape: `{"auctions": {"AAPL": [{"d":"2026-06-09","o":[{t,x,p,c}…],"c":[{t,x,p,c}…]}, …]}, "next_page_token": "…"}` — `o` = opening auctions, `c` = closing auctions.
```json
{"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"}
```
Usage: to seed/enrich `calendar.parquet` for a range **before** loading bars, call the **`backfill_lake_calendar`** lake tool (it loops `get_stock_auctions` internally across symbols and pages, inserts one row per distinct `d` with `source='auctions'`, `session_open`/`open_price` from `o[0]`, `session_close`/`close_price` from `c[0]`, and reports `calendar_days_added`). Direct call equivalent:
```json
{"symbols":"AAPL","feed":"sip","start":"2026-06-06","end":"2026-08-06"}
```
Cheap single-day confirmation for “was this a trading day?” and first pass of daily open/close. Intraday bars and per-symbol coverage still come from `get_stock_bars`.
## Operations notes
- **Rate limits:** Alpaca data API ~200 req/min. On `429` back off (exponential, start 1s) and retry. Batch symbols in one call, but page through `next_page_token`.
- **Atomicity:** write parquet to a `.<name>.tmp` then `rename()`; readers never see partial files. Apply the same pattern to metadata upserts.
- **Consistency:** one `feed` + one `adjustment` per lake (record in `manifest.yaml`). Refetching a window with a different feed/adjustment would silently corrupt merged history.
- **Feed / history limits (Alpaca):** SIP is never used (requires license). IEX goes back to **2020-07-27** for daily bars. For earlier data, Yahoo Finance fills gaps automatically (daily bars only). The `TAC_LAKE_START_DATE` (default `2000-01-03`) controls the earliest date requested; Yahoo provides data back to ~1970 for most symbols.
- **Dedup / merge:** bar upserts read the existing file, merge by canonical timestamp (new wins on duplicate `t`), and write the full set atomically — safe for both tail appends and leading-gap backfills (a file that already holds the newest bar still accepts older fetched history). Feature persistence is a fresh write per family, so no dedup needed.
- **Features** are derived from the lake bars (compute after bars are persisted, keyed `(market, symbol, timeframe, t)`), so indicator history stays aligned with bar history.
## End-to-end example (1d, 2 months, AAPL)
1. `coverage.parquet` has no `(US,1d,AAPL)` row → backfill.
2. Seed calendar: `backfill_lake_calendar` `{"symbols":"AAPL","start":…,"end":…}` (loops `get_stock_auctions`) → `calendar.parquet` trading days.
3. `get_lake_bars` `{"symbols":"AAPL","timeframe":"1d","start":…,"end":…,"lazy":true,"quiet":true}` → auto backfills the window, persists bars, updates coverage/symbols/calendar, returns a `{count, first_t, last_t}` summary instead of the bar rows.
4. Answer: DuckDB `SELECT * FROM read_parquet('$TAC_LAKE_DIR/market=US/timeframe=1d/symbol=AAPL.parquet') WHERE t >= now() - INTERVAL 2 MONTH` (or `get_lake_bars` again).
5. **Next identical request is a direct hit** (`source:"lake"`) from step 1’s decision table — no Alpaca fetch.