min_rr = 2.0 was hand-set in Admin (2026-06-24) and never swept — the gate ablation only tested the floor on-vs-off, never its level. It was the last un-swept knob in the live gate. Swept against portfolio Sharpe under the real exit, with a parity self-check (reproduces_production_gate: the row at the live floor must rebuild production's exact 1,089-setup qualified set — it does). min_rr qualified in-sample Sh/CAGR OOS Sh/CAGR (entries >= 2024-07) 0.0 6636 1.98 / 58.5% 2.02 / 66.2% 1.2 3897 1.34 / 33.9% 1.12 / 28.8% 1.5 3127 1.20 / 29.6% 1.12 / 28.8% 1.75 1974 1.64 / 44.5% 1.15 / 27.4% 2.0 (live) 1089 2.04 / 50.4% 2.78 / 73.3% 2.25 577 1.64 / 31.8% 1.71 / 31.9% 2.5 286 1.67 / 29.0% 0.68 / 8.7% KEEP 2.0. It is the optimum in both windows, and a peak that reproduces in data it was never fitted to is real evidence. But treat it as fragile: unlike the ATR trail (a plateau), this is a spike with a trough beside it — +/-0.25 costs ~0.4 Sharpe in-sample and ~1.6 out-of-sample — and the curve is bimodal (floor-off is good, 1.2-1.75 is bad, 2.0 is good). The hand-set value landed on the peak by luck, not by tuning. Do not nudge it. Worth knowing: turning the floor OFF entirely is the second-best row in both windows, with substantially higher CAGR (58.5% / 66.2%) and more trades. If CAGR ever outranks Sharpe here, "no R:R floor" is a live option — and it would sever the gate's last dependency on the weak S/R detector. Also fixes a metric artifact in the holdout harness. The train book's equity curve ran to the end of the data while its entries stopped at the split, so it sat in flat cash for two years and deflated its own CAGR/Sharpe (reported 0.95 / 14.6%; actually 1.31 / 29.6%). _simulate_portfolio now truncates the calendar to hold_days after the last entry when end_date is set — it only triggers on the holdout train window, so no other number moves. The clear-air OOS verdict is unaffected: it rests on the test row, whose entries and curve both start at the split and were always clean. Both holdout reports regenerated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
20 KiB
S/R levels: detection quality, the target exit, and the entry gate
Date: 2026-07-12 Question that started it: are our support/resistance levels built the way best practice says they should be, and do we actually use them that way?
Short answer: the detector is weak against best practice, but its reach into P&L runs entirely through the entry gate — not the exit. Honoring the target as a take-profit was tested and is decisively worse. Whether the S/R-derived gate is net-positive is the open question, tracked below.
1. How the levels are built today
app/services/sr_service.py::detect_sr_levels, over all stored history
(query_ohlcv with no date range — 5 years / ~1260 daily bars per ticker):
- Candidates = volume-profile HVN and LVN bins + pivot swing highs/lows.
- Strength = share of bars that "touched" the level, scaled so ~20% of bars → 100.
- Nearby levels merged within 0.5%; tagged
supportif below spot, elseresistance.
Where that departs from best practice
Measured on backtest_snapshots/prod.sqlite (AAPL, 1261 bars, spot $308.63):
| Gap | Evidence |
|---|---|
POC / VAH / VAL are computed then discarded. sr_service reads only hvn/lvn. The canonical volume-profile levels never become S/R. |
POC $148.32, VAH $230.45 — unused |
| HVN = "any bin above the mean", so nearly every bin is a candidate. A real HVN is a local peak in the histogram. | 8 of 20 bins HVN, other 12 LVN → 73 levels, median spacing $2.44 (0.79% of spot) — a price grid, not detected structure |
| HVN and LVN are scored and used identically, though they encode opposite dynamics (acceptance vs. rejection). | both appended as plain candidates |
| Volume is double-counted: a bar's full volume is added to every bin it spans rather than distributed. | binned total = 1.48× true volume |
| "Touch" = level fell inside the bar's range — a pass-through counts the same as a rejection. Strength therefore measures how central a price is in the 5-year range, not how often price reversed there. | 21 of 73 levels pin at exactly 100 → the sr_strength magnet in the probability model is near-constant |
| No recency decay, unbounded lookback. | 35 AAPL "support" levels sit >35% below spot |
| Round-number levels absent — the mechanism with the best empirical support (Osler 2000). | not implemented |
Not a defect: the pivot window=2 is the standard 5-bar Williams fractal. What it
lacks is a prominence filter — AAPL yields 338 pivots over 1261 bars, one every
~3.7 bars.
The structural problem: resistance famine
Levels are tagged relative to spot, so a stock near its highs has almost nothing above it. Across all 504 tickers in the snapshot, grouped by proximity to the 52-week high:
| Population | Median resistance levels | % with <3 |
|---|---|---|
| ≥98% of 52w high — what the momentum gate buys | 1 | 64% |
| 90–98% | 7 | 12% |
| <90% | 22 | 0% |
AAPL at $308: 71 support levels, 2 resistance levels. 26 tickers have zero.
This matters because _build_setups_at does if not targets: continue — no
resistance above ⇒ no long setup at all, and <3 targets adds a
target-availability conflict. The detector is structurally most blind exactly
where the momentum edge is strongest. recommendation_service.py:521 already says
so out loud: "price is extended near highs (no resistance target above), so no
high-conviction long setup is available."
2. Does the target even act as an exit? No.
Production runs paper_exit_mode = "atr_trailing" (DEFAULT_EXIT_MODE), and
_atr_trailing_close(direction, entry, init_stop, atr_multiplier, hold_days, ...)
does not take target as a parameter. Only the non-default mode == "target"
branch consults it.
So the S/R target's real causal role is:
- the entry gate —
rr ≥ 1.5, primary-targetprob ≥ 20%,≥3 targets, and no-resistance-above ⇒ no setup; - the displayed target table.
It decides whether you enter and plays no part in how you exit.
3. Experiment: should the target be honored as a take-profit? — REJECTED
Dennis's proposal: keep the 3× ATR trail, but also take profit when the S/R target
is hit. This was not representable in the simulator (exit_policy is one string;
the existing "target" policy replaces the trail). Added atr_trail3_target,
which runs both — trade ends at whichever comes first.
Identical 80/20 momentum entry and qualification for every row — the exit is the only
policy that changes. (Trade counts still vary 303–427, because exits free portfolio
slots at different times; the conclusion is robust to that.)
Report: reports/backtest-20260712-sr-target-exit.json
| exit policy | Sharpe | CAGR | MaxDD | trades | win% |
|---|---|---|---|---|---|
low20 |
2.04 | 53.2% | 21.9% | 305 | 37.7 |
atr_trail3 (production) |
2.04 | 50.4% | 21.4% | 320 | 37.5 |
hold |
2.00 | 51.9% | 22.2% | 303 | 38.6 |
technical40 |
2.00 | 51.9% | 22.2% | 303 | 38.6 |
sma50 |
1.76 | 40.8% | 24.4% | 385 | 33.8 |
target (take-profit, no trail) |
1.59 | 32.4% | 25.0% | 415 | 41.9 |
atr_trail3_target (trail + take-profit) |
1.47 | 28.9% | 23.5% | 427 | 40.0 |
Decision: do not ship. The target must not become an exit.
- Adding the take-profit to the trail: Sharpe 2.04 → 1.47, CAGR halved (50.4% → 28.9%), and drawdown got worse (21.4% → 23.5%). No risk compensation.
- Win rate rose (37.5% → 40.0%) — the tell. You win more often and earn far less: the take-profit converts the few 5R/8R/15R runners into 1.5R wins while every loser still costs a full −1R. Momentum's edge is that right tail.
- The combination (1.47) is worse than the take-profit alone (1.59): once upside is capped at the target, the trail's benefit (riding a winner far past any target) is gone but its cost (shakeouts on pullbacks) remains. Worst of both.
This confirms and extends the note at backtest_service.py:450 — swept fixed
take-profits never found an interior optimum; the S/R target is no better.
Caveat: single in-sample run over full history. The effect is large (CAGR halved), not marginal.
4. Open: is the S/R entry gate net-positive?
The target is now proven useless as an exit, so the gate is its only justification — and the gate is what starves the momentum names.
Evidence for keeping it (gate_ablation, prod baseline): dropping the R:R floor
halves per-setup expectancy, 0.583 → 0.301 net avg R (momentum_only = 0.345).
Inside the tradeable pool the S/R-derived R:R floor is doing real selection — it
favors names whose nearest resistance is far away, i.e. clear air above.
But that ablation cannot see the famine: it re-qualifies candidates that already
exist, and starved names never enter the candidate set (if not targets: continue).
So "remove the floors" is the wrong test — it just reproduces the rows above.
The right test changes target generation: when a direction has no S/R target,
synthesize one at k×ATR so the name becomes a candidate. One variable moves; the
previously-vetoed names now trade. Implemented behind
BACKTEST_ATR_TARGET_FALLBACK=<k> (k=3 matches the trail; rr = 3/1.5 = 2.0, which
clears the 1.5 floor, and an aligned momentum name lands ~34% probability, clearing
the 20% floor).
Read the result against the production baseline (atr_trail3, Sharpe 2.04):
- > 2.04 → the veto costs money; let the breakouts in.
- ≈ 2.04 → famine is a wash; a detector rewrite is cosmetic.
- < 2.04 → the veto earns its keep by keeping us out of over-extended names, and S/R gating is vindicated.
Result: the veto EARNS ITS KEEP. Keep S/R in the gate.
Treatment reports/backtest-20260712-sr-gate-ablation-treatment.json vs control
reports/backtest-20260711-prod-baseline.json. Admitting the vetoed names is a big
change: qualified setups go 1089 → 4230 (~4×).
Production exit (atr_trail3), full history:
| Sharpe | CAGR | MaxDD | Calmar | trades | |
|---|---|---|---|---|---|
| control (veto ON) | 2.04 | 50.4% | 21.4% | 2.36 | 320 |
| treatment (veto OFF) | 1.82 | 58.6% | 21.0% | 2.79 | 404 |
Full history alone looks like a genuine trade-off — more return, more volatility, Sharpe down but Calmar up. The lookback split is what settles it:
| window | control (veto ON) | treatment (veto OFF) |
|---|---|---|
| 6m | Sharpe 2.87, CAGR 76.6%, DD 8.1% | Sharpe 1.50, CAGR 53.6%, DD 15.8% |
| 1y | Sharpe 2.47, CAGR 66.8%, DD 8.8% | Sharpe 1.69, CAGR 66.0%, DD 15.8% |
| 3y | Sharpe 2.12, CAGR 52.3%, DD 17.7% | Sharpe 2.11, CAGR 75.9%, DD 21.0% |
| 5y | Sharpe 1.83, CAGR 38.8%, DD 21.4% | Sharpe 1.62, CAGR 44.9%, DD 21.0% |
| all | Sharpe 2.04, CAGR 50.4%, DD 21.4% | Sharpe 1.82, CAGR 58.6%, DD 21.0% |
Per-setup expectancy: all_floors net avg R 0.583 → 0.280.
Decision: do not ship the fallback. Keep the gate as it is. The flat 3×ATR fallback is worse on Sharpe and on per-setup expectancy, and production needs no further defense than that.
But be careful what this run does and does not prove
It does not isolate the famine hypothesis. The fallback fires on any empty
generate_targets result — and that includes the ATR/R:R distance filters in
TargetGenerator (target closer than 1 ATR, or beyond max_atr_multiple), not just
"no resistance above." Measured at the last bar across 502 tickers, of the long
setups the fallback admits:
- 26 (35%) have genuinely no resistance above — the clear-air famine case
- 49 (65%) do have resistance above; the ATR/R:R filters rejected it — a different population entirely
That matches the report's own tell: qualified setups exploded 1089 → 4230 (~4×) while candidates rose only ~16%. So the degradation may be driven mostly by that 65%, and the clear-air breakouts this investigation was about are a minority of what was admitted.
What the run actually supports: "a flat 3×ATR fallback for all S/R-starved setups degrades performance." It does not support the stronger claim that a stock in clear air is a worse risk-adjusted buy, or that the veto is functioning as an over-extension filter. That mechanism is unproven.
The window split is also less clean than it first looks. The verdict rests on the two smallest samples — 6m (n=30) and 1y (n=72) — where Sharpe 2.87 is noise-dominated. The statistically sturdier 3y window (n=230 → 303) shows equal Sharpe (2.12 vs 2.11) with substantially higher treatment CAGR (52.3% → 75.9%). Full-history Calmar also favors the treatment (2.79 vs 2.36). So the result is metric- and window-dependent; only the flat-fallback rejection is solid.
Second contamination (by design): the fallback gives every admitted name the same
rr = 3/1.5 = 2.0, so there is no R:R discrimination within the admitted set.
To actually test the famine, the fallback must fire only when there is no resistance above (not on ATR/R:R filter misses). That is the clear-air run below.
4b. The clean test: fire the fallback ONLY in clear air — the veto DOES cost money
BACKTEST_FALLBACK_CLEAR_AIR_ONLY=1 restricts the fallback to setups with no S/R
level ahead at all, excluding the 65% that were merely ATR/R:R distance-filter
misses. Same whole-portfolio simulation, same 10-slot book, same momentum ranking.
Report: reports/backtest-20260712-sr-gate-ablation-clearair.json
Production exit (atr_trail3), full history:
| arm | Sharpe | CAGR | MaxDD | Calmar | trades |
|---|---|---|---|---|---|
| control — veto ON (production) | 2.04 | 50.4% | 21.4% | 2.36 | 320 |
| blanket fallback (contaminated) | 1.82 | 58.6% | 21.0% | 2.79 | 404 |
| clear-air-only fallback | 2.07 | 62.3% | 20.1% | 3.10 | 363 |
Strictly better than production on all three headline metrics at once — higher Sharpe, ~12 points more CAGR, and lower drawdown. Qualified setups 1089 → 2072 (vs. 4230 for the blanket version).
By lookback (production strategy):
| window | control (veto ON) | clear-air fallback |
|---|---|---|
| 6m (n=30→38) | Sharpe 2.87, CAGR 76.6%, DD 8.1% | Sharpe 2.21, CAGR 78.8%, DD 13.0% |
| 1y (n=72→86) | Sharpe 2.47, CAGR 66.8%, DD 8.8% | Sharpe 2.28, CAGR 87.6%, DD 12.9% |
| 3y (n=230→270) | Sharpe 2.12, CAGR 52.3%, DD 17.7% | Sharpe 2.18, CAGR 68.6%, DD 14.3% |
| 5y (n=320→363) | Sharpe 1.83, CAGR 38.8%, DD 21.4% | Sharpe 1.85, CAGR 47.6%, DD 20.1% |
| all | Sharpe 2.04, CAGR 50.4%, DD 21.4% | Sharpe 2.07, CAGR 62.3%, DD 20.1% |
In every statistically sturdy window (3y, 5y, all) the clear-air fallback wins on Sharpe, CAGR and drawdown. The short windows (6m n=30, 1y n=72 — noise-dominated) favor control on Sharpe/DD while the treatment still earns more (1y CAGR 66.8% → 87.6%).
This reverses §4 and confirms the hypothesis that opened the investigation. The S/R veto on clear-air names was costing money; the blanket run masked it because the 65% loophole population (ATR/R:R filter misses) is genuinely bad and dominated the result. Isolate the two, and they pull in opposite directions:
- clear-air names (no resistance above): portfolio-accretive
- ATR/R:R distance-filter misses: portfolio-destructive
Per-setup expectancy is consistent with this: net avg R all_floors — control 0.583,
clear-air 0.402, blanket 0.280. The admitted clear-air setups are individually a bit
weaker, but they carry the highest momentum ranks, so they win book slots and deliver
outsized portfolio returns.
Caveats before shipping: in-sample, single snapshot; the fallback still assigns a
constant rr = 2.0 to every admitted name (no discrimination within the set). Needs
an out-of-sample run before production — see §4c, which is where it comes undone.
4c. Out-of-sample holdout — the §4b result does NOT survive
Everything in §4b is in-sample: the rule was chosen by looking at the same 5 years it
was then graded on. The portfolio_monitor lookbacks (6m/1y/3y/5y) are not a
holdout — they are nested windows all ending today, so each one overlaps the data the
idea came from.
Real split (BACKTEST_HOLDOUT_SPLIT=2024-07-01, production strategy, disjoint books):
- train = entries before 2024-07-01 (~3y)
- test = entries on/after 2024-07-01 (~2y, never informed the rule)
Reports: reports/backtest-20260712-holdout-control.json,
reports/backtest-20260712-holdout-clearair.json
| window | arm | Sharpe | CAGR | MaxDD | trades |
|---|---|---|---|---|---|
| train (2022-06 → 2024-08) | control | 1.31 | 29.6% | 21.4% | 174 |
| train | clear-air | 1.63 | 42.6% | 20.1% | 191 |
| test (2024-07 → 2026-07) | control | 2.78 | 73.3% | 11.7% | 150 |
| test | clear-air | 2.45 | 83.0% | 14.3% | 176 |
Harness bug, found and fixed 2026-07-12. The train row first reported Sharpe 0.95 / CAGR 14.6% — wrong. Its equity curve ran to the end of the data while its entries stopped at the split, so the book sat in flat cash for two years and deflated its own metrics.
_simulate_portfolionow truncates the calendar tohold_daysafter the last entry wheneverend_dateis set. The verdict is unaffected — it rests on the test row, whose entries and curve both start at the split and were always clean. But the broken numbers looked like a result, and nearly produced a false conclusion ("the first half of the sample was mediocre"). Corrected numbers above.
In train the clear-air rule wins on every metric. Out of sample it does not. On the held-out two years it delivers more raw return (+9.7pp CAGR) but at lower Sharpe (2.78 → 2.45) and higher drawdown (11.7% → 14.3%).
So the §4b headline — "strictly better on all three metrics" — was an in-sample artifact. Out of sample the rule is not a free win; it is a risk/return trade: it buys extra return by taking more risk, and on a risk-adjusted basis it is slightly worse than production.
Decision: do NOT ship the clear-air fallback. This project's decision metric is Sharpe throughout (every ranking in the report, and the 2026-07-10 primary-target A/B was accepted on Sharpe 1.51 → 2.00). By that standard the honest read of the only uncontaminated evidence is no improvement.
Notes for anyone revisiting:
- Both arms show a large regime shift (train Sharpe ~1.3–1.6, test Sharpe ~2.5–2.8) — the test window was simply a much better market. That is why relative comparison within a window is the only valid read.
- n = 150/176 in test is decent but not large; the Sharpe gap (0.33) is not overwhelming. This is "not confirmed," not "definitively refuted."
- The famine hypothesis is therefore real but not exploitable as tried: the clear-air names do add return (consistently, in both train and test), but the flat 3×ATR target admits them at a risk cost that eats the risk-adjusted benefit. A better target model for those names (§ next runs) is the remaining avenue.
A wholesale "ATR target for everyone" variant was deliberately not run as the
headline: with a fixed k×ATR target and a 1.5×ATR stop, rr = k/1.5 is constant
across every name, which erases the very selection the 0.301 credits. A loss there
would be uninterpretable.
5. Reproducing
Both research paths are off by default — the default report is byte-identical to the shipped baseline (5 exit rows, no fallback), and the full unit suite including the backtest↔prod parity guard passes.
# Exit book incl. the rejected take-profit rows
BACKTEST_RESEARCH_EXITS=1 python scripts/run_backtest_snapshot.py \
backtest_snapshots/prod.sqlite --workers 7 --allow-spawn
# S/R gate ablation: synthesize a 3xATR target where S/R offers none
BACKTEST_ATR_TARGET_FALLBACK=3 python scripts/run_backtest_snapshot.py \
backtest_snapshots/prod.sqlite --workers 7 --allow-spawn
Note --allow-spawn is required on Windows: _mp_context() has no fork/
forkserver there and silently falls back to a single thread without it.
6. Standing decisions
Measured:
- The target must not be an exit. Tested, rejected, decisively — Sharpe 2.04 → 1.47, CAGR halved. Momentum's edge is the right tail; a take-profit truncates it. (§3)
- Do NOT ship the clear-air fallback — it failed out-of-sample. In-sample it looked strictly better (Sharpe 2.04 → 2.07, CAGR 50.4% → 62.3%, DD 21.4% → 20.1%), but on a genuine holdout (entries after 2024-07-01, never seen by the rule) it is worse on Sharpe (2.78 → 2.45) and Calmar, better only on raw CAGR (+9.7pp). The in-sample "free win" was an artifact. Production gate stays as-is. (§4b, §4c)
- The famine is real, but not exploitable as tried. Clear-air names do add return consistently (train and test) — the veto genuinely leaves money on the table. But a flat 3×ATR target admits them at a risk cost that cancels the risk-adjusted benefit. (§4c)
- Do NOT relax the veto indiscriminately. The ATR/R:R distance-filter misses (65% of a blanket fallback) are portfolio-destructive and swamp everything — Sharpe 1.82, net avg R 0.280. The two populations pull in opposite directions and must be separated. (§4)
Reasoned, not measured — treat as hypotheses:
- The detector's flaws probably don't reach P&L directly. No run ever varied detection quality — "good S/R vs bad S/R → P&L" has never been measured. Fix the §1 gaps for the displayed levels and the UX; do not promise a return improvement.
Method note (the expensive lesson): the in-sample result in §4b was clean, large, consistent across five nested windows — and still didn't survive a holdout. Nested lookbacks are not out-of-sample. Split by entry date before believing anything.
Next runs, if picked back up:
- A per-name target model for clear-air setups instead of a constant k×ATR. This is the one avenue left: the return is demonstrably there (§4c), it's the flat target that makes it too expensive in risk. Grade on the §4c holdout, not full history.
- Sweep k (fallback distance); only k=3 was tried. Grade on the holdout.
- A volatility-aware admission rule for clear-air names — the OOS failure is a drawdown/vol story (11.7% → 14.3%), so sizing them down may recover the Sharpe.