book: fold Q-campaign (exp 33-43) evidence into ledger, claims, and chapters
- EVIDENCE#022-032: Q01-Q11 runs (2 PASS / 9 FAIL) with run_ids and branches - CLAIMS: promote M2 Sharpe-drift to PROVEN (Q01), refute Kelly (Q06), risk-limit-as-alpha (Q08), standalone reversal (Q11); add label-horizon + weekly-rebalance + long-short-turnover claims - README: TOC + claim inventories for ch 04/05/07/08/09/10/12 updated to the Q-campaign - new chapters 04 (prune), 05 (ensembles), 07 (isolation), 08 (construction), 09 (cost/turnover), 10 (risk limits & gates), 12 (synthesis); ch 00/02/03 updated - Q08 calibration evidence persisted under book/data/evidence/q08-risklimit/
This commit is contained in:
@@ -12,7 +12,7 @@ The metrics vocabulary of ch. 01 is only useful if the numbers can be trusted. T
|
||||
|
||||
## Why traceability is the methodology
|
||||
|
||||
TradeAC ran 31+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:
|
||||
TradeAC ran 40+ experiments. Without the branch+run+notes discipline, the desk could not have told which improvements were real. Two concrete failures make the case:
|
||||
|
||||
- **The dirty lake.** Pre-reset exp 8–18 reported strong results that did not survive a clean rebuild `(PROVEN → exp 21)`. Only because the exact YAML, branch, and run were recorded could the desk reproduce — and falsify — the reference. Traceability is what turned a false belief into evidence.
|
||||
- **Post-hoc cherry-picking.** With 31+ experiments, the best-looking number is expected to be inflated by selection. The counter is pre-registration: hypothesis, change, and acceptance metric are fixed in the run notes *before* the run `(REFERENCED — research practice; see CLAIMS.md multiple-testing note)`. Where the book quotes an experiment whose hypothesis was recorded after the fact, it says so.
|
||||
@@ -21,11 +21,11 @@ The pre-reset experiments (exp 8–18) are therefore treated as **idea material,
|
||||
|
||||
## The isolation discipline
|
||||
|
||||
A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio. Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.
|
||||
A traced experiment proves nothing unless one variable changed. The campaign's clean-lake sequence shows the discipline: exp 26 (n_drop 2→1) established the reference; exp 28 changed only seed count; exp 29 only the momentum bundle; exp 30 only the Sharpe-drift feature; exp 31 only the GARCH trio; and the Q-campaign (exp 33–43) changed exactly one thing per run against that same reference — features (Q01, Q11), seeds (Q02), topk (Q03), label horizon (Q04/Q05), sizing (Q06), rebalance cadence (Q07), risk limits (Q08), construction (Q09), and an entry gate (Q10). Because each changed one thing against the same reference, each verdict is attributable `(PROVEN → exp 28–31; EVIDENCE#022–032 → exp 33–43)`. Where isolation was lost (exp 12 pre-reset re-validations, exp 30's mixed metrics), the book marks the claim HYPOTHESIS.
|
||||
|
||||
## Falsification is the output
|
||||
|
||||
Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.
|
||||
Most additions failed. The loop's value is not that it produced winners — it is that it stopped wrong directions at the cost of a few runs: OU features (exp 25), momentum (exp 29), GARCH (exp 31), stochastic-control construction (exp 13/14), and nine of the eleven Q-campaign runs (longer labels, wider books, Kelly sizing, long-short, regime gates, standalone reversal — exp 34–38, 40–43). The campaign's refuted runs were as valuable as its wins `(PROVEN → refuted runs recorded; REFERENCED for the falsification principle)`. This is the stance carried through the book: a hypothesis that survives the loop becomes proved practice; one that fails becomes a recorded negative that the next hypothesis must beat.
|
||||
|
||||
## From here
|
||||
|
||||
|
||||
Reference in New Issue
Block a user