The README opened with "find the path of least resistance, key S/R zones, and asymmetric R:R setups" — a description of a strategy we do not run. What we run is a long-only cross-sectional momentum book with a trailing exit. The S/R engine, the composite score, sentiment and fundamentals are screening and display; none has a measured edge. - Rewrites the intro/philosophy around the real strategy, and says plainly what is NOT the edge. - Adds a mermaid decision graph, universe -> qualified -> ranked -> opened -> closed, with the real exit distribution on the terminal nodes: initial stop 45%, trailing stop 31%, max hold 24%, S/R target 0%. Validated against the mermaid parser, not eyeballed. - Documents that the R:R and touch-probability are GATE INPUTS, not forecasts of the trade — the single easiest way to misread this app. - Adds win rate, best/worst R and the exit-reason split to the production baseline table. - New docs/research/README.md: every strategy tested, the result, the decision, and why we stay with the current one. 12 rejected ideas (take-profit exits, clear-air gate relaxation, EV gate, regime overlay, inverse-vol sizing, shorts, standalone vol, FIP, ...), the confirmed tuning knobs, the open leads, and the method rules we learned the hard way (nested lookbacks are not out-of-sample; a rising win rate is a warning, not a win). - Documents the research flags and the holdout harness, and warns that the portfolio_monitor lookbacks are nested windows, NOT a holdout. - Notes the snapshot must copy paper_% settings or it silently diverges from prod. All baseline numbers re-verified against reports/backtest-20260711-prod-baseline.json (506 tickers, 1,089 qualified, CAGR 50.4%, +413.8% vs SPY +95.7%, DD -21.4%, Sharpe 2.04, 320 trades, 15.3d avg hold, and all five promotion contenders). No corrections were needed — the numbers were right, the framing was not. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Research log — what we tested, what happened, what we kept
Every strategy question we've put to the backtest, in one table. The point is to
stop re-litigating settled questions: if a row says "rejected", the experiment
was run and the data said no. Detail lives in the linked docs and in
reports/*.json (all committed).
The one-line summary of the whole platform: it is a long-only cross-sectional momentum book — buy the top quintile by beta-adjusted 12-1 momentum, tilt toward higher volatility, hold ≤ 10 names, cut at 1.5× ATR, then trail at 3× ATR for up to 30 trading days. Everything else in the app (composite score, S/R levels, sentiment, fundamentals) is display or screening, not edge.
1. What survived — the production strategy
| Component | Status | Why it's there |
|---|---|---|
| Residual 12-1 momentum, top 20%, long-only | The edge. Everything else is scaffolding | Only component with a measured cross-sectional IC. Promoted July 2026 |
| 80/20 residual-momentum / 6m-volatility rank | Ranking tilt | Buys ~2pp CAGR over momentum-only; costs ~6pp drawdown |
| 1.5× ATR initial stop | Real exit | Cuts losers fast |
| 3× ATR trailing stop, 30-day max hold | Real exit | Best Sharpe of every exit tested |
| Max 10 concurrent positions, 1% risk per trade | Sizing | Cap never binds in practice |
2. Rejected — do not resurrect without new data
| # | Experiment | Result | Decision | Evidence |
|---|---|---|---|---|
| 1 | S/R target as a take-profit (exit at the target, with or without the trail) | Sharpe 2.04 → 1.47, CAGR halved (50.4% → 28.9%). Win rate rose (37.5% → 40.0%) — the tell: it truncates the right tail | Rejected. The target must never become an exit | sr-levels-and-exits.md · backtest-20260712-sr-target-exit.json |
| 2 | Clear-air fallback — synthesize a 3× ATR target so 52-week-high breakouts stop being vetoed by "no resistance above" | Looked strictly better in-sample (Sharpe 2.07, CAGR 62.3%, DD 20.1%) but failed a real out-of-sample holdout: Sharpe 2.78 → 2.45, higher drawdown | Rejected. Gate stays as-is | sr-levels-and-exits.md · backtest-20260712-holdout-*.json |
| 3 | Blanket S/R fallback (any missing target, not just clear air) | Sharpe 1.82, per-setup expectancy 0.583 → 0.280 R | Rejected. 65% of what it admitted were ATR/R:R filter misses, which are actively bad | sr-levels-and-exits.md |
| 4 | Expected-value gate (min_expected_value replacing the R:R + probability pair) |
Structurally favoured distant lottery targets; selected worse-than-random setups | Removed June 2026. Settings dropped in migration 020 | migration 009, 020 |
| 5 | Blue-sky projected targets (invent a target above when none exists) | Dilutive under the ATR-trail exit | Reverted July 2026. Same root cause as #2 — better targets can't help when the exit ignores them | — |
| 6 | SPY 200d-MA regime overlay (block entries / go flat) | Halves return (315% → 138%), zero drawdown benefit | Rejected. The ATR trail already manages downside; the filter blocks the recovery entries that make the money | backtest-20260708-regime-overlay.json |
| 7 | Short setups | Fight the trend, drag expectancy | Excluded while the momentum gate is active | — |
| 8 | Standalone volatility ranking (high-vol 80, no momentum) | CAGR 31.6%, DD −34.8%, Sharpe 1.12 | Rejected. Vol is a tilt, not a signal | prod-baseline |
| 9 | Low-volatility ranking | CAGR 2.7%, Sharpe 0.29 | Rejected. No edge | prod-baseline |
| 10 | Inverse-vol position sizing | The apparent "win" was mis-attributed: the 20% notional cap bound on 95% of entries, so it measured concentration, not vol-sizing. Genuine inverse-vol cuts DD to −18.2% but costs ~58pp return at flat Sharpe | Rejected as edge; it's a risk-preference trade | backtest-20260709-position-sizing*.json |
| 11 | FIP path-smoothness as tie-breaker/filter | Non-monotonic within the qualified set; thinning the entry stream costs more compounding than the tilt returns | Rejected as a filter — but see §4, it's the strongest raw signal we've measured | — |
| 12 | Fixed take-profit sweep (R-multiples) | No interior optimum ever found — the best TP is "no TP" | Rejected. Momentum's edge lives in the right tail | backtest_service.py:450 |
3. Tuned and confirmed — don't retest on this snapshot
A systematic single-variable sweep (July 2026) confirmed every production setting. Re-running these against the same ~4-year snapshot is wasted compute and invites overfitting.
| Knob | Verdict |
|---|---|
| ATR trail multiple {1.5–4.0} | Keep 3.0 — ≤2.0 whipsaws out the right tail; ≥2.5 is a plateau |
| Momentum lookback (6-1, 3-1, 12-7 Novy-Marx, composites) | Keep residual 12-1 — the others have IC ≈ 0 or weaker t-stats |
| Selection cutoff {70…90} × book size {10, 15, 20} | Keep 80 × 10 — monotonically worse in both directions |
| Position sizing (equal-weight, inverse-vol, risk-% sweep) | Keep 1% fixed-fractional |
| Primary-target probability floor | Keep 20% — pruned lottery targets, 1,428 → 1,089 qualified, lifted Sharpe |
| Exit policy (hold / SMA50 / 20-day low / technical-40 / ATR trail) | Keep 3× ATR trail — best Sharpe (2.04) |
4. Open leads
| Lead | Why it's interesting | Blocker |
|---|---|---|
fip_id (information discreteness over the 12-1 window) |
Strongest cross-sectional signal measured on this universe — IC −0.045, t = −2.91, correct sign | Doesn't improve this book (the momentum gate already captures it in-sample). Revisit when the universe broadens |
Broader universe (nasdaq_all) |
Strengthens every week's cross-section and the IC t-stat | Also where fip_id could become tradeable |
| Forward paper-trade record | The only true out-of-sample evidence the snapshot cannot give | Time |
| Better target model for clear-air names | The return is demonstrably there (#2 wins on raw CAGR in both train and test); it's the flat 3× ATR target that makes it too expensive in risk | Needs a per-name model, not a constant k×ATR |
| S/R detector quality | POC/VAH/VAL computed then discarded; HVN = "any above-mean bin"; volume double-counted 1.48×; "touch" counts pass-throughs; no round numbers | Worth fixing for the levels users see — but it does not reach P&L, so don't justify it on returns |
5. Method rules learned the hard way
- Nested lookback windows are NOT out-of-sample. The clear-air result (#2) was
clean, large, and consistent across five nested windows — and still died on a
proper train/test split by entry date. Use
BACKTEST_HOLDOUT_SPLIT. - Check what population an ablation actually admits. The blanket fallback (#3) looked like it tested the "resistance famine" hypothesis. It didn't — 65% of the setups it let in were a different population entirely, and they drove the result.
- A rising win rate is a warning, not a win. Both #1 and #12 raised the hit rate while destroying returns. In a right-tailed strategy, "winning more often" usually means you clipped the winners.
- The iron rule: a signal earns its way into selection only through the
factor harness — |mean IC| ≳ 0.03, consistent sign,
reliable: true(≥ 12 non-overlapping windows). Never let an unvalidated score gate setups.
6. Why we stay with the current strategy
Everything we've tried to add has either failed the backtest, failed out-of-sample, or turned out to be measuring something other than what it claimed. What's left is a boring, well-documented result: cross-sectional momentum works; the machinery around it mostly doesn't.
The S/R engine, the composite score, the sentiment and fundamentals dimensions are all still in the product — they make the app legible and are useful context for a human — but none of them has a measured edge, and the platform is honest about that in the UI (see the exit plan and base-rate panels on every setup card). The one component that does have an edge is the momentum gate, and every knob on it has been swept and confirmed.
The next real evidence is forward, not backward: the live paper-trade record.