Brings the durable artifacts of research/portfolio-capacity-rebalancing onto main so the rationale for raising the count cap lives with the code that cites it. The matrix runner, the research simulator hooks and the study's unit tests are deliberately left behind; they remain at tag research/portfolio-capacity-final. Corrects conclusions that were reached on EV per trade and are now superseded: the findings doc's decisions 1 (keep cap 10) and 4 (run the risk-floor A/B) are struck through and answered in a new correction section, and the research README and phase-A matrix entries are updated to match. The frozen specification itself is untouched -- its recorded SHA-256 f1e37783 still verifies. effective-risk-floor-ab.md is retained but marked CLOSED/NEGATIVE: the study it proposes is already answered by cap15 vs cash_unbounded (-0.753pp CAGR while EV/trade rises), and its EV-based pass rule would have shipped it. scripts/research_rankings.py replaces a fourth copy of the historical rank-map helper; run_research_matrix, run_execution_recovery_matrix and run_daily_reentry_matrix now share it. The shared version adds a duplicate observation guard and a deterministic symbol tie-break the copies lacked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
133 lines
6.4 KiB
Markdown
133 lines
6.4 KiB
Markdown
# Phase A research matrix (2026-07-18) — results and decisions
|
||
|
||
Report: `reports/research-matrix-phase-a.json` / `.md`
|
||
Branch: `research/portfolio-vol-and-followups`
|
||
Validation split: entries ≥ **2024-07-01** (called *validation*, not holdout — this window has been opened before).
|
||
Cadence: daily, production gate/rank/trail + gate-reset re-entry.
|
||
Pre-registered N for DSR: **20**.
|
||
|
||
## Pre-registered promotion rule (unchanged after run start)
|
||
|
||
Promote only if **all** of:
|
||
|
||
1. Validation Sharpe ≥ control
|
||
2. Validation max DD not worse by more than **2pp**
|
||
3. Train Sharpe not worse (both-windows consistency)
|
||
|
||
Always report whether validation ΔSharpe exceeds **1 × SE** (expect most will not).
|
||
|
||
Mechanics guards confirmed before reading results: calendar truncation asserted on every arm; next-open re-anchors stop to fill − 1.5×ATR(signal); vol scalars apply at entry only.
|
||
|
||
---
|
||
|
||
## Control baseline
|
||
|
||
| Window | Sharpe | SE | CAGR | MaxDD | Calmar | Trades |
|
||
|---|---:|---:|---:|---:|---:|---:|
|
||
| Train | 1.75 | 0.68 | 49.8% | 17.9% | 2.78 | 240 |
|
||
| **Validation** | **1.68** | **0.72** | **41.6%** | **20.9%** | **1.99** | **239** |
|
||
| Full (close-fill) | 1.77 | 0.50 | 48.3% | 21.6% | 2.23 | 472 |
|
||
|
||
**Capacity correction (2026-08-05):** the full close-fill control also records
|
||
skipped_book_full = 519 versus 472 admitted trades, so the ten-slot book
|
||
refuses 52.4% of admitted+blocked qualified opportunities. The older weekly
|
||
claim that the cap never bound is stale and does not apply to this daily
|
||
gate-reset configuration. Capacity was isolated in the
|
||
[focused bracket study](portfolio-capacity-bracket.md) and **resolved: the count
|
||
cap was raised 10 → 15 so it no longer binds (+1.075pp CAGR paired, 51 paths
|
||
better / 2 worse, drawdown unchanged).** Note that the blocked *count* was a poor
|
||
guide in both directions — one path had 244 blocked entries and relieving all of
|
||
them moved CAGR by −0.1pp. See the
|
||
[findings correction](portfolio-capacity-bracket-findings.md#correction-2026-08-05-ev-per-trade-was-the-wrong-lens).
|
||
|
||
Validation SE ≈ 0.72 — almost no arm clears a 1-SE delta.
|
||
|
||
---
|
||
|
||
## Per-arm decisions
|
||
|
||
### A2 — Max hold {30, 45, 60, 90} — **note and move on**
|
||
|
||
| Hold | Val Sharpe | Val DD | Train Sharpe | Trades train |
|
||
|---:|---:|---:|---:|---:|
|
||
| 30 | 1.68 | 20.9 | 1.75 | 240 |
|
||
| 45 | **2.07** | 19.3 | **1.43** | 218 |
|
||
| 60 | **2.12** | 19.3 | **1.11** | 186 |
|
||
| 90 | 2.04 | 23.1 | 1.30 | 179 |
|
||
|
||
Validation-only would have “found” +0.4 Sharpe. Train collapses: longer holds leave stale names blocking slots (240 → 186 trades at hold-60). This is a **regime interaction** (trend validation vs chop train), not a free knob. A regime-conditional hold is a large research program; prior regime-overlay work already argues against that path.
|
||
|
||
**Decision: keep max hold 30. Do not ship longer static holds.**
|
||
|
||
### A3 — Equity-curve vol targeting — **reject as edge; park as optional insurance**
|
||
|
||
Scalars averaged 0.77–1.07 as designed (grid straddled historical book vol ~22–25%). Lower targets de-levered; **vt25** was nearly neutral (val Sharpe 1.63 vs 1.68). Wide clamp ≈ headline clamp. Lookback sensitivity did not unlock a win.
|
||
|
||
This sample has **no major vol-regime shift**, so the run **rejects vol targeting as an edge on this data** — it does **not** reject crash-insurance value in a future high-vol regime. The ~0.02 Sharpe cost at vt25 is a nearly free insurance policy if drawdown tolerance ever tightens.
|
||
|
||
**Decision: do not ship. Settles “Phase 2 = vol-scaled momentum” as an edge plan on this snapshot. Park vt25 as optional risk preference only.**
|
||
|
||
### A5 — Correlation caps — **reject; sector caps stay Phase B with reduced expectations**
|
||
|
||
Best near-miss: **0.6 skip** — val Sharpe 1.70, val DD **17.0%** (tempting), but train Sharpe 1.61 < 1.75 and full-period Sharpe **1.59 vs 1.77** (the cap deletes real momentum concentration profit). Half-size variants were worse.
|
||
|
||
**Decision: no pure corr cap. Sector caps remain Phase B with reduced expectations.**
|
||
|
||
### A4 — Next-open fill — **not a reject; the discovery**
|
||
|
||
| | Close control | Next-open |
|
||
|---|---:|---:|
|
||
| Full Sharpe | 1.77 | **1.20** |
|
||
| Full CAGR | 48.3% | **30.0%** |
|
||
| Val Sharpe | 1.68 | 1.44 |
|
||
| Val DD | 20.9% | **28.2%** |
|
||
|
||
Overnight gap on validation entries: mean **−0.52%**, median −0.18%, p05 −4.5%, p95 +2.1% (n=243).
|
||
|
||
This is not “slippage noise.” It is largely the **overnight momentum drift** that close-fill earns and a 07:00-Berlin scanner (signal yesterday’s close → fill tomorrow’s open) **structurally cannot**. Honest deployable number under that schedule is ~Sharpe 1.2 / CAGR 30%, not 1.77 / 48%.
|
||
|
||
**Decision baseline going forward (until near-close ships):** grade **promotion** under
|
||
`fill_mode=next_open`; keep close-fill as the historical control for comparability
|
||
with prior reports.
|
||
|
||
**Follow-up (done):** execution recovery matrix — see
|
||
**[execution-recovery.md](execution-recovery.md)**. Short version: monotone fill-timing
|
||
gradient + DD recovery prove this is *when you fill*; live bracket **[1.57, 1.77]**;
|
||
gap-cap dead; no more fill-timing sim on this snapshot; ops move R:R scan to NY
|
||
near-close (one scan/day).
|
||
|
||
### `fip_id` re-derivation — **validated**
|
||
|
||
Weekly IC fingerprint on this snapshot: **mean IC −0.045, t = −2.92**, reliable (35 weeks). Matches the July record. Safe to reuse when the universe broadens.
|
||
|
||
---
|
||
|
||
## Promotion table (rule as written)
|
||
|
||
| Outcome | Arms |
|
||
|---|---|
|
||
| Promote | only `a2_hold_30` (identity with control) |
|
||
| Reject | every other arm |
|
||
|
||
No arm cleared ΔSharpe > 1 SE.
|
||
|
||
---
|
||
|
||
## What not to do next
|
||
|
||
- Re-litigate rejected-table items, min_rr, GTL
|
||
- Regime-conditional max-hold as a “small” experiment
|
||
- Treat validation-only max-hold glitter as a free CAGR lift
|
||
- Ship vol targeting as edge without a vol-regime sample
|
||
- More fill-timing simulation on this snapshot (settled — see execution-recovery.md)
|
||
- Dual daily qualifying scans (would break gate-reset validation)
|
||
- Gap-up entry filters (third tail-trim failure)
|
||
|
||
## What to do next
|
||
|
||
1. **Ship near-close execution** — ops checklist in [execution-recovery.md](execution-recovery.md)
|
||
(one R:R scan/day in `America/New_York`, MOC window, partial-bar honesty).
|
||
2. Until that ships: **decision baseline = next_open**.
|
||
3. Strategy work (nasdaq_all, fip_id, sector) only **after** execution path is decided,
|
||
graded under the fill mode you will trade.
|