research: fip breadth diagnostics + compositional read

Add lagged/tier/prod-subset/mom-conditional checks on research.sqlite.
Log: unconditional sign is a winner/bleeder tug-of-war; mom-conditional
fip stays negative and reliable; warn on high-vol tilt if universe broadens.
This commit is contained in:
2026-07-18 21:40:05 +02:00
parent d34c7a21b7
commit ceaaadc49f
5 changed files with 995 additions and 25 deletions
+2 -2
View File
@@ -141,8 +141,8 @@ knobs.
| Lead | Why it's interesting | Blocker |
|---|---|---|
| **Near-close / MOC execution (ops)** | Recovers overnight momentum drift left on the table by a morning EU scan; evidence closed | Implement schedule + partial-bar scan path; one qualifying scan/day only |
| **`fip_id`** (information discreteness over the 12-1 window) | **Strongest cross-sectional signal measured on this universe** IC 0.045, t = 2.91, correct sign; re-derived fingerprint matched Phase A; ticker technicals show it display-only | Doesn't improve *this* book as a filter. **Phase B tooling ready:** liquid-breadth IC on research.sqlite — see [fip-breadth-ic.md](fip-breadth-ic.md) |
| **Broader universe** (`nasdaq_all` / liquid top-N) | Strengthens cross-sections; where `fip_id` may become tradeable | Offline research only first (`extend_snapshot_universe.py`); not prod scan |
| **`fip_id`** | Fingerprint IC 0.045 / t 2.91 on prod book; display-only on ticker technicals | **Phase B:** unconditional liquid-Nasdaq IC fails iron-rule **sign**; **mom-conditional** fip IC 0.088 / t 4.58 (alive as tilt candidate only). See [fip-breadth-ic.md](fip-breadth-ic.md) |
| **Broader universe** | Composition changes factor signs (fip tug-of-war; high-vol junk) | Any prod broaden must **re-validate 80/20 high-vol tilt** first; offline research only for now |
| **Forward paper-trade record** | The only true out-of-sample evidence the snapshot cannot give | Time; mark entries at actual near-close fill once ops ships |
| **Better target model for clear-air names** | The return is demonstrably there (#2 wins on raw CAGR in *both* train and test); it's the *flat* 3× ATR target that makes it too expensive in risk | Needs a per-name model, not a constant k×ATR |