Sweep the R:R floor; fix a holdout metric artifact
Deploy / lint (push) Successful in 10s
Deploy / test (push) Successful in 1m9s
Deploy / deploy (push) Successful in 37s

min_rr = 2.0 was hand-set in Admin (2026-06-24) and never swept — the gate
ablation only tested the floor on-vs-off, never its level. It was the last
un-swept knob in the live gate.

Swept against portfolio Sharpe under the real exit, with a parity self-check
(reproduces_production_gate: the row at the live floor must rebuild production's
exact 1,089-setup qualified set — it does).

  min_rr   qualified   in-sample Sh/CAGR   OOS Sh/CAGR (entries >= 2024-07)
  0.0        6636      1.98 / 58.5%        2.02 / 66.2%
  1.2        3897      1.34 / 33.9%        1.12 / 28.8%
  1.5        3127      1.20 / 29.6%        1.12 / 28.8%
  1.75       1974      1.64 / 44.5%        1.15 / 27.4%
  2.0 (live) 1089      2.04 / 50.4%        2.78 / 73.3%
  2.25        577      1.64 / 31.8%        1.71 / 31.9%
  2.5         286      1.67 / 29.0%        0.68 /  8.7%

KEEP 2.0. It is the optimum in both windows, and a peak that reproduces in data
it was never fitted to is real evidence. But treat it as fragile: unlike the ATR
trail (a plateau), this is a spike with a trough beside it — +/-0.25 costs ~0.4
Sharpe in-sample and ~1.6 out-of-sample — and the curve is bimodal (floor-off is
good, 1.2-1.75 is bad, 2.0 is good). The hand-set value landed on the peak by
luck, not by tuning. Do not nudge it.

Worth knowing: turning the floor OFF entirely is the second-best row in both
windows, with substantially higher CAGR (58.5% / 66.2%) and more trades. If CAGR
ever outranks Sharpe here, "no R:R floor" is a live option — and it would sever
the gate's last dependency on the weak S/R detector.

Also fixes a metric artifact in the holdout harness. The train book's equity curve
ran to the end of the data while its entries stopped at the split, so it sat in
flat cash for two years and deflated its own CAGR/Sharpe (reported 0.95 / 14.6%;
actually 1.31 / 29.6%). _simulate_portfolio now truncates the calendar to
hold_days after the last entry when end_date is set — it only triggers on the
holdout train window, so no other number moves. The clear-air OOS verdict is
unaffected: it rests on the test row, whose entries and curve both start at the
split and were always clean. Both holdout reports regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-12 16:45:59 +02:00
co-authored by Claude Opus 4.8
parent 906d1db7d1
commit ea11efe3d1
7 changed files with 220674 additions and 4783 deletions
+36
View File
@@ -58,6 +58,42 @@ invites overfitting.
| Position sizing (equal-weight, inverse-vol, risk-% sweep) | **Keep 1% fixed-fractional** |
| Primary-target probability floor | **Keep 20%** — pruned lottery targets, 1,428 → 1,089 qualified, lifted Sharpe |
| Exit policy (hold / SMA50 / 20-day low / technical-40 / ATR trail) | **Keep 3× ATR trail** — best Sharpe (2.04) |
| **Activation R:R floor `min_rr`** (swept 2026-07-12) | **Keep 2.0** — best in-sample *and* out-of-sample. But it is a **spike, not a plateau** — see below |
### The `min_rr` sweep (2026-07-12)
`min_rr = 2.0` had been hand-set in Admin and **never swept** — the gate ablation only
tested the floor *on vs off*, never its level. Swept against portfolio Sharpe under the
real exit, with a parity self-check (`reproduces_production_gate: true` — the row at 2.0
rebuilds production's exact 1,089-setup qualified set).
Reports: `backtest-20260712-min-rr-sweep.json` (in-sample), `-oos.json` (test window only).
| min_rr | qualified | In-sample Sharpe / CAGR | **OOS** Sharpe / CAGR (entries ≥ 2024-07) |
|---|---|---|---|
| 0.0 (floor off) | 6636 | 1.98 / 58.5% | 2.02 / 66.2% |
| 1.2 (code default) | 3897 | 1.34 / 33.9% | 1.12 / 28.8% |
| 1.5 | 3127 | 1.20 / 29.6% | 1.12 / 28.8% |
| 1.75 | 1974 | 1.64 / 44.5% | 1.15 / 27.4% |
| **2.0 (live)** | 1089 | **2.04 / 50.4%** | **2.78 / 73.3%** |
| 2.25 | 577 | 1.64 / 31.8% | 1.71 / 31.9% |
| 2.5 | 286 | 1.67 / 29.0% | 0.68 / 8.7% |
| 3.0 | 89 | 1.09 / 9.1% | 0.87 / 5.0% |
**Verdict: keep 2.0.** It is the optimum in **both** windows, and the peak reproducing in
data it was never fitted to is real evidence — the one thing the clear-air experiment
couldn't show.
**But treat it as fragile, and do not nudge it.** Unlike the ATR trail (a plateau above
2.5), this is a **spike with a trough beside it**: ±0.25 costs ~0.4 Sharpe in-sample and
~1.6 Sharpe out-of-sample. A knob that sharp is not a robustly identified parameter, and
the curve is *bimodal* (floor-off is good, 1.21.75 is bad, 2.0 is good) — which is not
how a well-behaved threshold behaves. We got lucky: the hand-set value landed on the peak.
**Also worth knowing:** turning the floor **off entirely** is the second-best row in both
windows — nearly the same Sharpe with **substantially higher CAGR** (58.5% / 66.2%) and
more trades. If CAGR ever matters more than Sharpe here, "no R:R floor" is a live option,
and it would also sever the last dependency the *gate* has on the weak S/R detector.
---
+18 -10
View File
@@ -283,17 +283,25 @@ Real split (`BACKTEST_HOLDOUT_SPLIT=2024-07-01`, production strategy, disjoint b
Reports: `reports/backtest-20260712-holdout-control.json`,
`reports/backtest-20260712-holdout-clearair.json`
| window | arm | Sharpe | CAGR | MaxDD | Calmar | trades |
|---|---|---|---|---|---|---|
| train | control | 0.95 | 14.6% | 21.4% | 0.68 | 174 |
| train | **clear-air** | **1.18** | **20.5%** | **20.1%** | **1.02** | 191 |
| **test** | **control** | **2.78** | 73.3% | **11.7%** | **6.26** | 150 |
| **test** | clear-air | 2.45 | **83.0%** | 14.3% | 5.80 | 176 |
| window | arm | Sharpe | CAGR | MaxDD | trades |
|---|---|---|---|---|---|
| train (2022-06 → 2024-08) | control | 1.31 | 29.6% | 21.4% | 174 |
| train | **clear-air** | **1.63** | **42.6%** | **20.1%** | 191 |
| **test** (2024-07 → 2026-07) | **control** | **2.78** | 73.3% | **11.7%** | 150 |
| **test** | clear-air | 2.45 | **83.0%** | 14.3% | 176 |
> **Harness bug, found and fixed 2026-07-12.** The train row first reported Sharpe 0.95 /
> CAGR 14.6% — wrong. Its equity curve ran to the *end of the data* while its entries
> stopped at the split, so the book sat in flat cash for two years and deflated its own
> metrics. `_simulate_portfolio` now truncates the calendar to `hold_days` after the last
> entry whenever `end_date` is set. **The verdict is unaffected** — it rests on the test
> row, whose entries and curve both start at the split and were always clean. But the
> broken numbers *looked* like a result, and nearly produced a false conclusion ("the
> first half of the sample was mediocre"). Corrected numbers above.
**In train the clear-air rule wins on every metric. Out of sample it does not.** On
the held-out two years it delivers **more raw return (+9.7pp CAGR)** but at
**lower Sharpe (2.78 → 2.45), higher drawdown (11.7% → 14.3%) and lower Calmar
(6.26 → 5.80)**.
**lower Sharpe (2.78 → 2.45)** and **higher drawdown (11.7% → 14.3%)**.
So the §4b headline — *"strictly better on all three metrics"* — was **an in-sample
artifact.** Out of sample the rule is not a free win; it is a **risk/return trade**:
@@ -306,8 +314,8 @@ was accepted on Sharpe 1.51 → 2.00). By that standard the honest read of the o
uncontaminated evidence is *no improvement*.
Notes for anyone revisiting:
- Both arms show a large regime shift (train Sharpe ~1, test Sharpe ~2.52.8) — the
test window was simply a much better market. That is why *relative* comparison
- Both arms show a large regime shift (train Sharpe ~1.31.6, test Sharpe ~2.52.8) —
the test window was simply a much better market. That is why *relative* comparison
within a window is the only valid read.
- n = 150/176 in test is decent but not large; the Sharpe gap (0.33) is not
overwhelming. This is "not confirmed," not "definitively refuted."