feat: log Phase A decisions and add execution-recovery matrix
Document Phase A (max-hold/vol/corr closed; next-open as decision baseline). Add stale_close and next_open gap-cap fill modes plus a small matrix to test whether near-close scheduling recovers overnight momentum drift.
This commit is contained in:
+22
-4
@@ -105,18 +105,36 @@ and it would also sever the last dependency the *gate* has on the weak S/R detec
|
||||
|
||||
---
|
||||
|
||||
## 4. Open leads
|
||||
## 4. Phase A matrix (2026-07-18) — closed
|
||||
|
||||
Full write-up: **[phase-a-matrix.md](phase-a-matrix.md)** ·
|
||||
`reports/research-matrix-phase-a.json`.
|
||||
|
||||
| Arm | Decision |
|
||||
|---|---|
|
||||
| Max-hold {45,60,90} | **Note and move on** — validation glitter, train collapse (regime interaction) |
|
||||
| Equity-curve vol targeting | **Reject as edge** on this sample; park vt25 as optional DD insurance only |
|
||||
| Correlation caps | **Reject**; sector caps stay Phase B with reduced expectations |
|
||||
| Next-open fill | **Discovery, not reject** — honest deployable ~Sharpe 1.2 / CAGR 30%. Decision baseline for future promotion = `next_open` |
|
||||
| `fip_id` re-derive | **Validated** (IC −0.045, t = −2.92) |
|
||||
|
||||
**Highest-leverage open work:** near-close execution recovery (scheduling, not a new signal). Simulator: `scripts/run_execution_recovery_matrix.py` (`stale_close` + gap-cap).
|
||||
|
||||
---
|
||||
|
||||
## 5. Open leads
|
||||
|
||||
| Lead | Why it's interesting | Blocker |
|
||||
|---|---|---|
|
||||
| **`fip_id`** (information discreteness over the 12-1 window) | **Strongest cross-sectional signal measured on this universe** — IC −0.045, t = −2.91, correct sign | Doesn't improve *this* book (the momentum gate already captures it in-sample). Revisit when the universe broadens |
|
||||
| **Near-close / MOC execution** | Recovers the overnight momentum drift a 07:00-Berlin scanner leaves on the table (~0.5 Sharpe / ~18pp CAGR vs close-fill) | Prove with `stale_close` arm; then schedule change |
|
||||
| **`fip_id`** (information discreteness over the 12-1 window) | **Strongest cross-sectional signal measured on this universe** — IC −0.045, t = −2.91, correct sign; re-derived fingerprint matched Phase A | Doesn't improve *this* book. Revisit when the universe broadens |
|
||||
| **Broader universe** (`nasdaq_all`) | Strengthens every week's cross-section and the IC t-stat | Also where `fip_id` could become tradeable |
|
||||
| **Forward paper-trade record** | The only true out-of-sample evidence the snapshot cannot give | Time |
|
||||
| **Better target model for clear-air names** | The return is demonstrably there (#2 wins on raw CAGR in *both* train and test); it's the *flat* 3× ATR target that makes it too expensive in risk | Needs a per-name model, not a constant k×ATR |
|
||||
|
||||
---
|
||||
|
||||
## 5. Method rules learned the hard way
|
||||
## 6. Method rules learned the hard way
|
||||
|
||||
1. **Nested lookback windows are NOT out-of-sample.** The clear-air result (#2) was
|
||||
clean, large, and consistent across five nested windows — and still died on a
|
||||
@@ -133,7 +151,7 @@ and it would also sever the last dependency the *gate* has on the weak S/R detec
|
||||
|
||||
---
|
||||
|
||||
## 6. Why we stay with the current strategy
|
||||
## 7. Why we stay with the current strategy
|
||||
|
||||
Everything we've tried to add has either failed the backtest, failed
|
||||
out-of-sample, or turned out to be measuring something other than what it claimed.
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
# Phase A research matrix (2026-07-18) — results and decisions
|
||||
|
||||
Report: `reports/research-matrix-phase-a.json` / `.md`
|
||||
Branch: `research/portfolio-vol-and-followups`
|
||||
Validation split: entries ≥ **2024-07-01** (called *validation*, not holdout — this window has been opened before).
|
||||
Cadence: daily, production gate/rank/trail + gate-reset re-entry.
|
||||
Pre-registered N for DSR: **20**.
|
||||
|
||||
## Pre-registered promotion rule (unchanged after run start)
|
||||
|
||||
Promote only if **all** of:
|
||||
|
||||
1. Validation Sharpe ≥ control
|
||||
2. Validation max DD not worse by more than **2pp**
|
||||
3. Train Sharpe not worse (both-windows consistency)
|
||||
|
||||
Always report whether validation ΔSharpe exceeds **1 × SE** (expect most will not).
|
||||
|
||||
Mechanics guards confirmed before reading results: calendar truncation asserted on every arm; next-open re-anchors stop to fill − 1.5×ATR(signal); vol scalars apply at entry only.
|
||||
|
||||
---
|
||||
|
||||
## Control baseline
|
||||
|
||||
| Window | Sharpe | SE | CAGR | MaxDD | Calmar | Trades |
|
||||
|---|---:|---:|---:|---:|---:|---:|
|
||||
| Train | 1.75 | 0.68 | 49.8% | 17.9% | 2.78 | 240 |
|
||||
| **Validation** | **1.68** | **0.72** | **41.6%** | **20.9%** | **1.99** | **239** |
|
||||
| Full (close-fill) | 1.77 | 0.50 | 48.3% | 21.6% | 2.23 | 472 |
|
||||
|
||||
Validation SE ≈ 0.72 — almost no arm clears a 1-SE delta.
|
||||
|
||||
---
|
||||
|
||||
## Per-arm decisions
|
||||
|
||||
### A2 — Max hold {30, 45, 60, 90} — **note and move on**
|
||||
|
||||
| Hold | Val Sharpe | Val DD | Train Sharpe | Trades train |
|
||||
|---:|---:|---:|---:|---:|
|
||||
| 30 | 1.68 | 20.9 | 1.75 | 240 |
|
||||
| 45 | **2.07** | 19.3 | **1.43** | 218 |
|
||||
| 60 | **2.12** | 19.3 | **1.11** | 186 |
|
||||
| 90 | 2.04 | 23.1 | 1.30 | 179 |
|
||||
|
||||
Validation-only would have “found” +0.4 Sharpe. Train collapses: longer holds leave stale names blocking slots (240 → 186 trades at hold-60). This is a **regime interaction** (trend validation vs chop train), not a free knob. A regime-conditional hold is a large research program; prior regime-overlay work already argues against that path.
|
||||
|
||||
**Decision: keep max hold 30. Do not ship longer static holds.**
|
||||
|
||||
### A3 — Equity-curve vol targeting — **reject as edge; park as optional insurance**
|
||||
|
||||
Scalars averaged 0.77–1.07 as designed (grid straddled historical book vol ~22–25%). Lower targets de-levered; **vt25** was nearly neutral (val Sharpe 1.63 vs 1.68). Wide clamp ≈ headline clamp. Lookback sensitivity did not unlock a win.
|
||||
|
||||
This sample has **no major vol-regime shift**, so the run **rejects vol targeting as an edge on this data** — it does **not** reject crash-insurance value in a future high-vol regime. The ~0.02 Sharpe cost at vt25 is a nearly free insurance policy if drawdown tolerance ever tightens.
|
||||
|
||||
**Decision: do not ship. Settles “Phase 2 = vol-scaled momentum” as an edge plan on this snapshot. Park vt25 as optional risk preference only.**
|
||||
|
||||
### A5 — Correlation caps — **reject; sector caps stay Phase B with reduced expectations**
|
||||
|
||||
Best near-miss: **0.6 skip** — val Sharpe 1.70, val DD **17.0%** (tempting), but train Sharpe 1.61 < 1.75 and full-period Sharpe **1.59 vs 1.77** (the cap deletes real momentum concentration profit). Half-size variants were worse.
|
||||
|
||||
**Decision: no pure corr cap. Sector caps remain Phase B with reduced expectations.**
|
||||
|
||||
### A4 — Next-open fill — **not a reject; the discovery**
|
||||
|
||||
| | Close control | Next-open |
|
||||
|---|---:|---:|
|
||||
| Full Sharpe | 1.77 | **1.20** |
|
||||
| Full CAGR | 48.3% | **30.0%** |
|
||||
| Val Sharpe | 1.68 | 1.44 |
|
||||
| Val DD | 20.9% | **28.2%** |
|
||||
|
||||
Overnight gap on validation entries: mean **−0.52%**, median −0.18%, p05 −4.5%, p95 +2.1% (n=243).
|
||||
|
||||
This is not “slippage noise.” It is largely the **overnight momentum drift** that close-fill earns and a 07:00-Berlin scanner (signal yesterday’s close → fill tomorrow’s open) **structurally cannot**. Honest deployable number under that schedule is ~Sharpe 1.2 / CAGR 30%, not 1.77 / 48%.
|
||||
|
||||
**Decision baseline going forward:** grade **promotion** under `fill_mode=next_open`; keep close-fill as the historical control for comparability with prior reports.
|
||||
|
||||
**Highest-leverage follow-up (not a strategy change):** near-close / MOC-style execution (~15:45 ET) so live fills sit near the close the signal is built on. Simulator proof arm: `stale_close` (signal t−1 close, fill t close). Secondary: next-open **gap-cap** (skip open > +2% vs signal close) — measure, don’t assume.
|
||||
|
||||
### `fip_id` re-derivation — **validated**
|
||||
|
||||
Weekly IC fingerprint on this snapshot: **mean IC −0.045, t = −2.92**, reliable (35 weeks). Matches the July record. Safe to reuse when the universe broadens.
|
||||
|
||||
---
|
||||
|
||||
## Promotion table (rule as written)
|
||||
|
||||
| Outcome | Arms |
|
||||
|---|---|
|
||||
| Promote | only `a2_hold_30` (identity with control) |
|
||||
| Reject | every other arm |
|
||||
|
||||
No arm cleared ΔSharpe > 1 SE.
|
||||
|
||||
---
|
||||
|
||||
## What not to do next
|
||||
|
||||
- Re-litigate rejected-table items, min_rr, GTL
|
||||
- Regime-conditional max-hold as a “small” experiment
|
||||
- Treat validation-only max-hold glitter as a free CAGR lift
|
||||
- Ship vol targeting as edge without a vol-regime sample
|
||||
|
||||
## What to do next
|
||||
|
||||
1. **Execution recovery matrix** — `stale_close` vs close vs next_open; optional gap-cap under next_open (`scripts/run_execution_recovery_matrix.py`).
|
||||
2. If `stale_close` ≈ close control: schedule scan near the US close (not a signal rewrite).
|
||||
3. Until near-close execution ships live: **decision baseline = next_open**.
|
||||
4. Phase B data work only when wanted: `nasdaq_all` + `fip_id`, earnings calendar, sector residual/caps.
|
||||
Reference in New Issue
Block a user