5.2 KiB
Chapter 02 — Baseline and the Cost Reality
Status: drafting. Claim inventory: see README.md ch. 02.
This chapter answers the question every quant desk must answer before the first dollar is deployed: what does the raw signal have to be worth, and what survives the cost of trading it?
The honest answer on the TradeAC stack, measured on the clean lake, is that the signal itself was modest — and the cost of expressing it was nearly its entire gross value. The order of magnitude is the lesson.
The noise floor first
Before quoting a single IC, establish what noise looks like. For a cross-section of N independent names, daily RankIC under the null has standard deviation roughly 1/√(N−1). On the 50-ETF panel that is ≈ 0.143 per day. A signal whose daily RankIC mean is a small fraction of that standard deviation is statistically indistinguishable from noise day-to-day, however it may look averaged.
The post-reset clean-lake reference signal (exp 24, the compact stochastic set) reports mean RankIC ≈ 0.066 and RankICIR ≈ 0.255 — positive and above the informal 0.2 RankICIR "noise threshold" used on this desk, but far from overwhelming: PROVEN — EVIDENCE#013 → exp 24. HYPOTHESIS (chat-derived null calibration: clean-lake mean RankIC of earlier runs sat ≈ 0.18–0.4σ of the null per-day distribution — book/data/chat_mining/exp-polluted-lake.txt). Treat the statistical significance of a 7-month, 50-name cross-section as fragile, not robust.
The cost model that decides everything
The backtest and live sizing on this stack use a fixed cost model:
- open cost 0.0005 (5 bp), close cost 0.0015 (15 bp), minimum $5 per side;
- fills assumed at the close (
deal_price = $close), benchmark SPY, $1M starting account.
PROVEN — strategy config of exp 21–31. These are round-trip costs of ~20 bp, which is ordinary for liquid US ETFs at retail/PT sizes but not free. At ~20% of the book traded daily (topk=10, n_drop=2), the annualized cost drag is enormous relative to a signal worth single-digit annual excess.
The gross → net collapse on clean data
The clean-lake sequence shows the pattern with the same signal, same costs, varying only the feature set and turnover:
| Run | Signal (IC / RankICIR) | Gross excess vs SPY | Net excess vs SPY | Net IR |
|---|---|---|---|---|
| exp 22 (full TA+SP) | 0.0486 / 0.243 | +0.12% | −9.09% | −0.80 |
| exp 23 (general sp only) | 0.0728 / 0.206 | +6.73% | −2.39% | −0.22 |
| exp 24 (compact sp) | 0.0511 / 0.255 | +5.99% | −3.21% | −0.32 |
| exp 26 (compact, n_drop=1) | 0.0511 / 0.255 | +7.02% | +2.13% | +0.21 |
PROVEN — EVIDENCE#011/012/013/015 → exp 22/23/24/26. Read the columns, not the rows: even the best clean-lake signal, at the default construction, lost roughly nine to ten percentage points of annualized excess to costs (exp 24: +5.99% gross → −3.21% net). The signal that produced a high long-short Sharpe (L/S ann Sharpe 4.54) could not survive daily rebalancing at 20 bp round trips.
This is the single most important number in the early book: at this turnover, cost is not a haircut, it is the strategy's budget. PROVEN — EVIDENCE#015 → exp 26 (identical IC/RankIC across n_drop 2 and 1; the entire net difference is trading behavior, not signal). The pre-reset campaign observed the same shape historically (baseline +6.2% gross → +1.6% net), which is idea material, not evidence: HYPOTHESIS (idea: pre-clean-lake, EVIDENCE#002 → exp 8).
What fixed it, and what it implies
The only construction change that flipped net from negative to positive was reducing daily forced replacements from n_drop=2 to n_drop=1 — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. PROVEN — EVIDENCE#015 → exp 26.
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are cost reduction (turnover, spread costs, size class) and edge preservation, not adding features. The feature-isolation campaign (ch. 06) then confirmed that most candidate additions reduced the edge anyway.
Desk rules distilled from this chapter
- Establish the null noise floor before believing any IC/RankIC mean on a small cross-section.
- Report gross and net excess side by side, always with universe + window + cost model.
- Treat net-IR-of-signal as the bar for any construction change; signal metrics alone are not a strategy claim.
- When net is negative and gross is positive by ~10pp, attack turnover before features.
TODO(evidence-needed: realized-cost comparison of round 3 vs the 5bp/15bp/$5 model once the position window closes).
Evidence cited in this chapter
| Tag | Source |
|---|---|
EVIDENCE#013 |
exp 24, run fe469a19…, branch exp/24-run-the-rankic-ensemble-in-mlflow-experi |
EVIDENCE#011/012 |
exp 22/23, runs 18db5bc1… / be5cd314… |
EVIDENCE#015 |
exp 26, run 21afc6af…, branch exp/26-test-whether-reducing-topkdropout-daily |
EVIDENCE#002 |
exp 8 (pre-clean-lake, idea only) |
| chat mining | book/data/chat_mining/exp-polluted-lake.txt (null calibration, idea only) |