Files
book-tac/book/chapters/08-portfolio-construction.md
T

39 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chapter 08 — Portfolio Construction: Dropout vs Optimal Stop
Status: drafting. Claim inventory: see `README.md` ch. 08.
This chapter compares the two portfolio constructions the campaign actually ran — TopkDropout (the rank-based, turnover-conscious book that became the reference and the live book) and stochastic-control OptimalStopControl (entry/exit/stop parametrized allocation). The honest status is that the comparison is **pre-reset idea material**: both constructions were tested on the dirty lake and the alternates were never re-run on the clean lake. What is PROVEN on the clean lake is that the reference book is TopkDropout and that it executes (ch. 11); what the alternates would do on clean data is unknown.
## The two constructions
- **TopkDropout** (qlib TopkDropoutStrategy): each day, rank the cross-sectional scores, apply the dropout rule, and hold the selected top-k names equally weighted. The campaign's `n_drop` parameter controls which names the strategy refuses to chase, and with it the book's turnover (ch. 09).
- **OptimalStopControl** (stochastic control): allocate toward a target portfolio with entry/exit thresholds, holding-period and stop-loss parameters. The campaign tried the baseline (entry 0.85 / exit 0.7 / hold 10 / stop −0.08) and a V2 with turnover bands, cooldown and a cap.
## The pre-reset comparison (idea material)
On the pre-reset lake, TopkDropout beat both stochastic-control variants: OptimalStopControl net excess −2.7% (IR −0.31) versus TopkDropout +7.8% `(idea → exp 13, pre-clean-lake; EVIDENCE#006)`, and the V2 also refuted (net −6.9%, IR −0.72) `(idea → exp 14, pre-clean-lake; EVIDENCE#007)`. The attributed mechanism was **turnover**: the stop-control constructions churned the book and bled ~11.3pp of cost drag `(idea → exp 13, EVIDENCE#006)`. That mechanism is plausible — it is the same cost drag that proved binding on the clean lake (exp 26, ch. 09) — but the numbers themselves are not usable (dirty lake, EVIDENCE#010).
`TODO(evidence-needed: OptimalStopControl vs TopkDropout A/B on the exp-26 reference and its n_drop=1 book — the clean-lake rerun of this comparison)`
## What is PROVEN on the clean lake
- The reference book is TopkDropout with `n_drop=1`, equal weight × risk_degree, and it is the campaign's best net result `(PROVEN → exp 26, EVIDENCE#015)`.
- The same construction is the live book of round 3: 10 targets, 9 fills, realized slippage 4.54 bps `(PROVEN → round 3, EVIDENCE#020)`.
- Construction is not a substitute for signal or cost work: the n_drop change moved net return by ~5.3pp with *identical* signal metrics `(PROVEN → exp 26, EVIDENCE#015)` — construction is where the cost edge is won or lost, and cost is the binding constraint (ch. 09).
## Sizing
Sizing in the campaign is equal weight × `risk_degree` (0.95) — a fixed fraction of account per name, floored to whole shares at execution `(PROVEN → the sizing used in exp 26 and round 3; tac-rd-book intents)`.
- Fractional-Kelly sizing (exp 15) was never verified — the run never finished `(HYPOTHESIS; TODO(evidence-needed: exp 15 Kelly re-run on the clean lake))`.
- The hypothesis that equal-weight × risk_degree throws away edge-magnitude information (a Kelly-style rule would size by score spread) is untested `(HYPOTHESIS → chat-ideas.md)`.
## Practice note
The book's working rule: prefer the construction that minimizes turnover at a fixed topk (TopkDropout with controlled `n_drop` over parametrized stop-control), because cost is the binding constraint on this signal. That rule is a hypothesis until the clean-lake A/B lands.
## Open questions
- `TODO(evidence-needed: OptimalStopControl vs TopkDropout on clean data)`
- `TODO(evidence-needed: Kelly-style sizing vs equal-weight × risk_degree on the exp-26 book)`
- `TODO(evidence-needed: lower topk vs higher topk on the clean-lake reference — concentration vs diversification)`