Files
tac-exp-dev/queue/designs/q20_effective_names.md
T
zhaoli 7124ef5d8e queue: pre-register Series 2 (Q12-Q20) targeting unproven book hypotheses
Series 1 (Q01-Q11, exp 33-43) executed and folded into book/CLAIMS.md.
Series 2 covers the remaining HYPOTHESIS rows and open questions:
  Q12 22d label + weekly recompute (untested combo)
  Q13 weekly rebalance reproduction on a 2nd OOS window
  Q14 out-of-universe validation (single-stock panel, needs lake backfill)
  Q15 5-seed vs single-seed clean A/B
  Q16 hmm family as features
  Q17 realized-moments family as features
  Q18 OptimalStopControl clean re-test
  Q19/Q20 martingale-VR + effective-names scripted studies

Each workflow pins one-variable change vs exp-26 reference and acceptance.
2026-08-20 03:51:58 +00:00

22 lines
1.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# QUEUE-20 — Effective independent names in the 50-ETF book (no qrun)
**Status:** QUEUED · **Priority:** P2 · **Effort:** ad-hoc script under `book/data/`
## Hypothesis (settle)
The 50-ETF book has only ~4 effective independent names (CLAIMS.md HYPOTHESIS,
chat-derived eigenvalue analysis, pre-reset). This is a concentration/diversification
claim with direct sizing relevance; verify it on the clean lake.
## Method (persist everything under `book/data/evidence/q20-effective-names/`)
1. Load the 50-ETF panel 1d returns from the lake for the test window 2026-01-04..2026-08-10.
2. Standardize returns; compute the correlation matrix and its eigendecomposition.
3. Count eigenvalues above the Marchenko–Pastur bound (N=50, T≈150) and report the
cumulative-variance share of the top k components.
4. Effective-rank measures: participation ratio `(Σλ)² / Σλ²` and cumulative 80%
variance count.
5. Write `eigenanalysis.csv` + a one-page summary.
## Acceptance
- If effective rank ≈ 4 (top-4 explain ~80%+ variance), the concentration claim is
PROVEN and feeds chapter 08 sizing guidance (why topk 10→20 adds no breadth).
- If effective rank is much larger, mark the claim REFUTED.