Investigated whether our support/resistance detection follows best practice
and whether we actually use it that way. Three findings, all backed by runs
against the prod snapshot and written up in docs/research/sr-levels-and-exits.md:
- The S/R target must NOT become an exit. Honoring it as a take-profit on top
of the 3x ATR trail drops Sharpe 2.04 -> 1.47 and halves CAGR. Win rate rises
(37.5% -> 40.0%), which is the tell: it truncates the right tail where
momentum's edge lives.
- The clear-air fallback (synthesize a 3xATR target where no resistance exists,
so 52-week-high breakouts stop being vetoed) looked strictly better in-sample
(Sharpe 2.04 -> 2.07, CAGR 50.4% -> 62.3%, DD 21.4% -> 20.1%) but FAILED a
real out-of-sample holdout: on entries after 2024-07-01 it is worse on Sharpe
(2.78 -> 2.45) and Calmar, better only on raw CAGR. Not shipped.
- The detector itself is weak vs best practice (POC/VAH/VAL computed then
discarded, HVN = any above-mean bin, 1.48x volume double-counting, "touch"
counts pass-throughs, no round numbers), but its only causal path to P&L is
the entry gate. Fix it for the displayed levels, not for returns.
Method note: nested lookback windows are NOT out-of-sample. The in-sample result
was clean, large, and consistent across five windows, and still did not survive
a proper entry-date split.
All research paths are off by default and the default report is unchanged:
BACKTEST_RESEARCH_EXITS=1 take-profit exit rows
BACKTEST_ATR_TARGET_FALLBACK=k synthetic k*ATR target when S/R offers none
BACKTEST_FALLBACK_CLEAR_AIR_ONLY=1 restrict that to genuinely clear air
BACKTEST_HOLDOUT_SPLIT=YYYY-MM-DD train/test split by entry date
Also fixes two reproducibility holes found while reconciling our local baseline
against the live report:
- create_backtest_snapshot.py now copies paper_% settings. The production
monitor row replays the runtime exit policy via get_exit_policy(); without
those keys a snapshot silently falls back to code defaults, so a live-tuned
exit would never be reflected.
- Migration 020 drops activation_min_expected_value and
activation_min_target_probability. Both are orphans of the June EV-gate
redesign, read by no code path, but prod carries min_target_probability = 50.0
which implies a probability floor that is not enforced (the real floor is the
20% constant in qualification.py).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The reports are the evidence behind the production baseline, so they belong
next to the README that quotes them rather than living only on one machine.
Un-ignores reports/*.json (~2.6 MB compressed for all 11); the snapshot DBs
they run against stay ignored.
Renames the reports to a single dated scheme so they sort chronologically and
say what they measured. Each name is derived from the report's own contents
(the atr_trail_sweep / regime_overlay / blue_sky_projected / sizing_test
sections, and the qualified counts that identify the A/B arms), not from the
ad-hoc slugs they carried before. The run the README quotes is now
backtest-20260711-prod-baseline.json.
reports/compare_reports.py loads every report into one sortable table
(portfolio monitor, entry variants, exit policies, portfolio sim), filters by
report and lookback, and highlights the best row for a chosen metric — max
drawdown correctly ranking lowest-as-best. Stdlib tkinter, no dependencies.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>