book: insert ch01 metrics vocabulary + ch02 research loop; renumber 03/06 — evidence exp 21-31, martingale study
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
# Chapter 03 — Baseline and the Cost Reality
|
||||
|
||||
Status: drafting. Claim inventory: see `README.md` ch. 03.
|
||||
|
||||
This chapter answers the question every quant desk must answer before the first dollar is deployed: **what does the raw signal have to be worth, and what survives the cost of trading it?**
|
||||
|
||||
The honest answer on the TradeAC stack, measured on the clean lake, is that the signal itself was modest — and the cost of expressing it was nearly its entire gross value. The order of magnitude is the lesson.
|
||||
|
||||
## The noise floor first
|
||||
|
||||
Before quoting a single IC, establish what noise looks like. For a cross-section of `N` independent names, daily RankIC under the null has standard deviation roughly `1/√(N−1)`. On the 50-ETF panel that is ≈ 0.143 per day. A signal whose daily RankIC mean is a small fraction of that standard deviation is statistically indistinguishable from noise day-to-day, however it may look averaged.
|
||||
|
||||
The post-reset clean-lake reference signal (exp 24, the compact stochastic set) reports mean RankIC ≈ 0.066 and RankICIR ≈ 0.255 — positive and above the informal 0.2 RankICIR "noise threshold" used on this desk, but far from overwhelming: `PROVEN — EVIDENCE#013 → exp 24`. `HYPOTHESIS (chat-derived null calibration: clean-lake mean RankIC of earlier runs sat ≈ 0.18–0.4σ of the null per-day distribution — book/data/chat_mining/exp-polluted-lake.txt)`. Treat the statistical significance of a 7-month, 50-name cross-section as fragile, not robust.
|
||||
|
||||
## The cost model that decides everything
|
||||
|
||||
The backtest and live sizing on this stack use a fixed cost model:
|
||||
|
||||
- open cost 0.0005 (5 bp), close cost 0.0015 (15 bp), minimum $5 per side;
|
||||
- fills assumed at the close (`deal_price = $close`), benchmark SPY, $1M starting account.
|
||||
|
||||
`PROVEN — strategy config of exp 21–31`. These are round-trip costs of ~20 bp, which is ordinary for liquid US ETFs at retail/PT sizes but not free. At ~20% of the book traded daily (topk=10, n_drop=2), the annualized cost drag is enormous relative to a signal worth single-digit annual excess.
|
||||
|
||||
## The gross → net collapse on clean data
|
||||
|
||||
The clean-lake sequence shows the pattern with the same signal, same costs, varying only the feature set and turnover:
|
||||
|
||||
| Run | Signal (IC / RankICIR) | Gross excess vs SPY | Net excess vs SPY | Net IR |
|
||||
|-----|------------------------|---------------------|-------------------|--------|
|
||||
| exp 22 (full TA+SP) | 0.0486 / 0.243 | +0.12% | −9.09% | −0.80 |
|
||||
| exp 23 (general sp only) | 0.0728 / 0.206 | +6.73% | −2.39% | −0.22 |
|
||||
| exp 24 (compact sp) | 0.0511 / 0.255 | +5.99% | −3.21% | −0.32 |
|
||||
| exp 26 (compact, n_drop=1) | 0.0511 / 0.255 | +7.02% | +2.13% | +0.21 |
|
||||
|
||||
`PROVEN — EVIDENCE#011/012/013/015 → exp 22/23/24/26`. Read the columns, not the rows: even the *best* clean-lake signal, at the default construction, lost roughly **nine to ten percentage points of annualized excess to costs** (exp 24: +5.99% gross → −3.21% net). The signal that produced a high long-short Sharpe (L/S ann Sharpe 4.54) could not survive daily rebalancing at 20 bp round trips.
|
||||
|
||||
This is the single most important number in the early book: **at this turnover, cost is not a haircut, it is the strategy's budget.** `PROVEN — EVIDENCE#015 → exp 26 (identical IC/RankIC across n_drop 2 and 1; the entire net difference is trading behavior, not signal)`. The pre-reset campaign observed the same shape historically (baseline +6.2% gross → +1.6% net), which is idea material, not evidence: `HYPOTHESIS (idea: pre-clean-lake, EVIDENCE#002 → exp 8)`.
|
||||
|
||||
## What fixed it, and what it implies
|
||||
|
||||
The only construction change that flipped net from negative to positive was reducing daily forced replacements from `n_drop=2` to `n_drop=1` — holding the previously-dropped name instead of trading around it (exp 26). Signal metrics were byte-identical to exp 24. The gain was pure cost relief. `PROVEN — EVIDENCE#015 → exp 26`.
|
||||
|
||||
Methodological reading: when the gross edge is ~7% and the cost drag ~9–10%, the two levers with the largest expected payoffs are *cost reduction* (turnover, spread costs, size class) and *edge preservation*, not adding features. The feature-isolation campaign (ch. 07) then confirmed that most candidate additions *reduced* the edge anyway.
|
||||
|
||||
## Desk rules distilled from this chapter
|
||||
|
||||
1. Establish the null noise floor before believing any IC/RankIC mean on a small cross-section.
|
||||
2. Report gross and net excess side by side, always with universe + window + cost model.
|
||||
3. Treat net-IR-of-signal as the bar for any construction change; signal metrics alone are not a strategy claim.
|
||||
4. When net is negative and gross is positive by ~10pp, attack turnover before features.
|
||||
5. `TODO(evidence-needed: realized-cost comparison of round 3 vs the 5bp/15bp/$5 model once the position window closes)`.
|
||||
|
||||
## Evidence cited in this chapter
|
||||
|
||||
| Tag | Source |
|
||||
|-----|--------|
|
||||
| `EVIDENCE#013` | exp 24, run `fe469a19…`, branch `exp/24-run-the-rankic-ensemble-in-mlflow-experi` |
|
||||
| `EVIDENCE#011/012` | exp 22/23, runs `18db5bc1…` / `be5cd314…` |
|
||||
| `EVIDENCE#015` | exp 26, run `21afc6af…`, branch `exp/26-test-whether-reducing-topkdropout-daily` |
|
||||
| `EVIDENCE#002` | exp 8 (pre-clean-lake, idea only) |
|
||||
| chat mining | book/data/chat_mining/exp-polluted-lake.txt (null calibration, idea only) |
|
||||
Reference in New Issue
Block a user