Demote/relabel the setup-outcome check; add exit reason to My Trades

The "tracking/drift" chip compared the live target/stop/expired outcome cohort
against the backtest's target/stop bucket (overall_qualified) — a like-for-like
pipeline check — but sat directly under the portfolio monitor, which shows the
promoted 3x-ATR-trailing book. That juxtaposition (plus "faithfully implementing
it" copy) made a plumbing/QA signal read as validation of the ATR-trail strategy
you actually trade. It validates neither the trailing-stop book nor real trades.

- Move the check out of the monitor block into the "Track-record maintenance"
  disclosure, relabelled "Setup-outcome pipeline check" with copy that says it
  checks the setup-grading pipeline (no look-ahead/config/data drift), NOT the
  ATR-trail production book. The genuine live validation stays My Trades (real
  paper trades, same ATR-trail exits) up top.
- Add a compact "Exit" column to My Trades showing close_reason
  (Stop/Trail/Target/Time/Manual) — the field was already plumbed to the
  frontend PaperTrade type, so this is frontend-only.

tsc -b && vite build pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-04 09:51:13 +02:00
co-authored by Claude Opus 4.8
parent 02b28f5ea6
commit 2a4bdd16a8
3 changed files with 99 additions and 62 deletions
@@ -1,7 +1,6 @@
import { useMemo, useState } from 'react'; import { useMemo, useState } from 'react';
import { useMutation, useQueryClient } from '@tanstack/react-query'; import { useMutation, useQueryClient } from '@tanstack/react-query';
import { useBacktestReport } from '../../hooks/useMarketRegime'; import { useBacktestReport } from '../../hooks/useMarketRegime';
import { usePerformance } from '../../hooks/usePerformance';
import { triggerJob } from '../../api/admin'; import { triggerJob } from '../../api/admin';
import { Button } from '../ui/Button'; import { Button } from '../ui/Button';
import { Callout } from '../ui/Callout'; import { Callout } from '../ui/Callout';
@@ -10,14 +9,6 @@ import { Section } from '../ui/Section';
import { useToast } from '../ui/Toast'; import { useToast } from '../ui/Toast';
import type { BacktestCurvePoint, BacktestPortfolioMonitorRun } from '../../lib/types'; import type { BacktestCurvePoint, BacktestPortfolioMonitorRun } from '../../lib/types';
// Need at least this many matured setups before a live-vs-backtest verdict means
// anything; below it the live sample is too noisy to compare.
const MIN_MATURED = 20;
// Live expectancy this far (in R) below the backtest counts as drift, not noise.
const DRIFT_TOLERANCE_R = 0.2;
type TrackingStatus = 'building' | 'tracking' | 'drift' | 'no-backtest';
function fmtR(v: number | null | undefined): string { function fmtR(v: number | null | undefined): string {
if (v === null || v === undefined) return '—'; if (v === null || v === undefined) return '—';
return `${v > 0 ? '+' : ''}${v.toFixed(2)}R`; return `${v > 0 ? '+' : ''}${v.toFixed(2)}R`;
@@ -67,17 +58,6 @@ function Stat({ label, value, valueClass = 'text-gray-100', sub }: {
); );
} }
function VerdictChip({ status }: { status: TrackingStatus }) {
const styles: Record<TrackingStatus, { cls: string; label: string }> = {
tracking: { cls: 'border-emerald-500/30 bg-emerald-500/15 text-emerald-300', label: '✓ tracking' },
drift: { cls: 'border-amber-500/30 bg-amber-500/15 text-amber-300', label: '⚠ drift' },
building: { cls: 'border-white/10 bg-white/[0.05] text-gray-400', label: 'building' },
'no-backtest': { cls: 'border-white/10 bg-white/[0.05] text-gray-400', label: 'no backtest' },
};
const s = styles[status];
return <span className={`shrink-0 rounded-full border px-2.5 py-1 text-xs font-medium ${s.cls}`}>{s.label}</span>;
}
function curvePath( function curvePath(
points: BacktestCurvePoint[], points: BacktestCurvePoint[],
min: number, min: number,
@@ -158,7 +138,6 @@ function EquityCurveChart({ run }: { run: BacktestPortfolioMonitorRun }) {
export function BacktestPanel() { export function BacktestPanel() {
const { data: report, isLoading } = useBacktestReport(); const { data: report, isLoading } = useBacktestReport();
const { data: perf } = usePerformance({ qualified_only: true });
const queryClient = useQueryClient(); const queryClient = useQueryClient();
const toast = useToast(); const toast = useToast();
const [selectedStrategy, setSelectedStrategy] = useState(''); const [selectedStrategy, setSelectedStrategy] = useState('');
@@ -178,22 +157,6 @@ export function BacktestPanel() {
[monitor, activeStrategy, activeLookback], [monitor, activeStrategy, activeLookback],
); );
// Live matured qualified cohort vs the backtest's qualified expectancy — the
// out-of-sample check that the running system faithfully implements the backtest.
const liveAvgR = perf?.overall.avg_r ?? null;
const liveN = perf?.overall.total ?? 0;
const btAvgR = report?.overall_qualified.avg_r ?? null;
let status: TrackingStatus = 'building';
if (liveAvgR != null && liveN >= MIN_MATURED) {
status = btAvgR == null ? 'no-backtest' : liveAvgR >= btAvgR - DRIFT_TOLERANCE_R ? 'tracking' : 'drift';
}
const verdictNote: Record<TrackingStatus, string> = {
building: `Fewer than ~${MIN_MATURED} matured setups so far — until then the backtest is the edge estimate. This turns into a live check as setups age past their ~30-day window.`,
'no-backtest': 'Run the backtest to get a baseline to compare the live record against.',
tracking: 'Live setups are resolving in line with the backtest — the running system is faithfully implementing it.',
drift: 'Live expectancy is running materially below the backtest — small-sample noise, a regime shift, or a live/backtest gap. Worth a look.',
};
const run = useMutation({ const run = useMutation({
mutationFn: () => triggerJob('backtest'), mutationFn: () => triggerJob('backtest'),
onSuccess: (res) => { onSuccess: (res) => {
@@ -208,7 +171,7 @@ export function BacktestPanel() {
}); });
return ( return (
<Section title="Is the strategy working?" hint="portfolio simulation vs S&P 500, validated against the live record"> <Section title="Is the strategy working?" hint="portfolio simulation of the promoted strategy vs S&P 500">
<div className="space-y-4"> <div className="space-y-4">
<div className="flex flex-wrap items-start justify-between gap-3"> <div className="flex flex-wrap items-start justify-between gap-3">
<Disclosure summary="How this is measured"> <Disclosure summary="How this is measured">
@@ -320,25 +283,6 @@ export function BacktestPanel() {
</div> </div>
)} )}
{/* Live-vs-backtest validation: does the running system realize what the backtest promised? */}
<div className="glass-sm space-y-2 p-4">
<div className="flex flex-wrap items-center justify-between gap-x-6 gap-y-2">
<div className="flex flex-wrap items-baseline gap-x-5 gap-y-1">
<span className="text-sm text-gray-400">
Live <span className={`num font-semibold ${rColor(liveAvgR)}`}>{fmtR(liveAvgR)}</span>
</span>
<span className="text-sm text-gray-400">
Backtest <span className={`num font-semibold ${rColor(btAvgR)}`}>{fmtR(btAvgR)}</span>
</span>
<span className="text-xs text-gray-500">
{liveN} matured{perf ? ` · ${perf.maturing} maturing` : ''} · qualified expectancy
</span>
</div>
<VerdictChip status={status} />
</div>
<p className="text-[11px] leading-relaxed text-gray-500">{verdictNote[status]}</p>
</div>
{monitor.note && <p className="text-[11px] text-gray-600">{monitor.note}</p>} {monitor.note && <p className="text-[11px] text-gray-600">{monitor.note}</p>}
</div> </div>
) : ( ) : (
@@ -374,7 +318,7 @@ export function BacktestPanel() {
<p className="text-[11px] text-gray-600"> <p className="text-[11px] text-gray-600">
Strategy research — gate tuning, exit sweeps, factor rank-IC — now runs locally against a Strategy research — gate tuning, exit sweeps, factor rank-IC — now runs locally against a
database snapshot (see README). This page keeps only what says whether the promoted strategy database snapshot (see README). This page keeps only what says whether the promoted strategy
is worth trading and being delivered live. is worth trading; your realized results up top show what it is actually delivering.
</p> </p>
</> </>
)} )}
@@ -19,6 +19,18 @@ function color(v: number | null): string {
return 'text-gray-300'; return 'text-gray-300';
} }
// How the trade was closed — useful context on real trades at almost no cost.
function reasonMeta(reason: string | null): { label: string; cls: string } {
switch (reason) {
case 'stop': return { label: 'Stop', cls: 'text-red-400' };
case 'trailing': return { label: 'Trail', cls: 'text-amber-400' };
case 'target': return { label: 'Target', cls: 'text-emerald-400' };
case 'time': return { label: 'Time', cls: 'text-gray-400' };
case 'manual': return { label: 'Manual', cls: 'text-blue-300' };
default: return { label: '—', cls: 'text-gray-500' };
}
}
function Stat({ label, value, valueClass = 'text-gray-100', sub }: { function Stat({ label, value, valueClass = 'text-gray-100', sub }: {
label: string; value: string; valueClass?: string; sub?: string; label: string; value: string; valueClass?: string; sub?: string;
}) { }) {
@@ -85,6 +97,7 @@ export function MyTradesPanel() {
<th className="px-4 py-2.5 text-right">P&L</th> <th className="px-4 py-2.5 text-right">P&L</th>
<th className="px-4 py-2.5 text-right">R</th> <th className="px-4 py-2.5 text-right">R</th>
<th className="px-4 py-2.5 text-right">Alpha</th> <th className="px-4 py-2.5 text-right">Alpha</th>
<th className="px-4 py-2.5">Exit</th>
<th className="px-4 py-2.5 text-right">Closed</th> <th className="px-4 py-2.5 text-right">Closed</th>
</tr> </tr>
</thead> </thead>
@@ -102,6 +115,11 @@ export function MyTradesPanel() {
<td className={`num px-4 py-2.5 text-right font-semibold ${p ? color(p.pnl) : 'text-gray-500'}`}>{p ? money(p.pnl) : '—'}</td> <td className={`num px-4 py-2.5 text-right font-semibold ${p ? color(p.pnl) : 'text-gray-500'}`}>{p ? money(p.pnl) : '—'}</td>
<td className={`num px-4 py-2.5 text-right ${p?.r != null ? color(p.r) : 'text-gray-500'}`}>{p?.r != null ? fmtR(p.r) : '—'}</td> <td className={`num px-4 py-2.5 text-right ${p?.r != null ? color(p.r) : 'text-gray-500'}`}>{p?.r != null ? fmtR(p.r) : '—'}</td>
<td className={`num px-4 py-2.5 text-right ${t.alpha_pct != null ? color(t.alpha_pct) : 'text-gray-500'}`} title="Return vs. S&P 500 over the holding period">{t.alpha_pct != null ? `${t.alpha_pct >= 0 ? '+' : ''}${t.alpha_pct.toFixed(1)}%` : '—'}</td> <td className={`num px-4 py-2.5 text-right ${t.alpha_pct != null ? color(t.alpha_pct) : 'text-gray-500'}`} title="Return vs. S&P 500 over the holding period">{t.alpha_pct != null ? `${t.alpha_pct >= 0 ? '+' : ''}${t.alpha_pct.toFixed(1)}%` : '—'}</td>
<td className="px-4 py-2.5">
<span className={`num text-[10px] font-semibold uppercase tracking-wider ${reasonMeta(t.close_reason).cls}`} title="How the trade was closed">
{reasonMeta(t.close_reason).label}
</span>
</td>
<td className="num px-4 py-2.5 text-right text-gray-500">{t.closed_at ? new Date(t.closed_at).toLocaleDateString() : '—'}</td> <td className="num px-4 py-2.5 text-right text-gray-500">{t.closed_at ? new Date(t.closed_at).toLocaleDateString() : '—'}</td>
</tr> </tr>
))} ))}
@@ -1,4 +1,6 @@
import { useMutation, useQueryClient } from '@tanstack/react-query'; import { useMutation, useQueryClient } from '@tanstack/react-query';
import { usePerformance } from '../../hooks/usePerformance';
import { useBacktestReport } from '../../hooks/useMarketRegime';
import { triggerJob, resetTrackRecord } from '../../api/admin'; import { triggerJob, resetTrackRecord } from '../../api/admin';
import { Button } from '../ui/Button'; import { Button } from '../ui/Button';
import { Disclosure } from '../ui/Disclosure'; import { Disclosure } from '../ui/Disclosure';
@@ -6,10 +8,63 @@ import { useToast } from '../ui/Toast';
import { BacktestPanel } from './BacktestPanel'; import { BacktestPanel } from './BacktestPanel';
import { MyTradesPanel } from './MyTradesPanel'; import { MyTradesPanel } from './MyTradesPanel';
// Need at least this many matured setups before the pipeline check means anything;
// below it the live sample is too noisy to compare.
const MIN_MATURED = 20;
// Live expectancy this far (in R) below the backtest counts as drift, not noise.
const DRIFT_TOLERANCE_R = 0.2;
type PipelineStatus = 'building' | 'tracking' | 'drift' | 'no-backtest';
function fmtR(value: number | null): string {
if (value === null) return '—';
return `${value > 0 ? '+' : ''}${value.toFixed(2)}R`;
}
function rColor(value: number | null): string {
if (value === null) return 'text-gray-400';
if (value > 0) return 'text-emerald-400';
if (value < 0) return 'text-red-400';
return 'text-gray-300';
}
function StatusChip({ status }: { status: PipelineStatus }) {
const styles: Record<PipelineStatus, { cls: string; label: string }> = {
tracking: { cls: 'border-emerald-500/30 bg-emerald-500/15 text-emerald-300', label: '✓ in sync' },
drift: { cls: 'border-amber-500/30 bg-amber-500/15 text-amber-300', label: '⚠ drift' },
building: { cls: 'border-white/10 bg-white/[0.05] text-gray-400', label: 'building' },
'no-backtest': { cls: 'border-white/10 bg-white/[0.05] text-gray-400', label: 'no backtest' },
};
const s = styles[status];
return <span className={`shrink-0 rounded-full border px-2.5 py-1 text-xs font-medium ${s.cls}`}>{s.label}</span>;
}
export function TrackRecordPanel() { export function TrackRecordPanel() {
const queryClient = useQueryClient(); const queryClient = useQueryClient();
const toast = useToast(); const toast = useToast();
// Setup-outcome pipeline check: does the live outcome evaluator reproduce the
// backtest's target/stop grading? Both sides use the SAME target/stop/expired
// model — this is a plumbing/QA signal (no look-ahead, config or data drift),
// NOT validation of the ATR-trail production strategy shown in the monitor.
const { data: perf } = usePerformance({ qualified_only: true });
const { data: report } = useBacktestReport();
const liveAvgR = perf?.overall.avg_r ?? null;
const liveN = perf?.overall.total ?? 0;
const btAvgR = report?.overall_qualified.avg_r ?? null;
let status: PipelineStatus = 'building';
if (liveAvgR != null && liveN >= MIN_MATURED) {
status = btAvgR == null ? 'no-backtest' : liveAvgR >= btAvgR - DRIFT_TOLERANCE_R ? 'tracking' : 'drift';
}
const statusNote: Record<PipelineStatus, string> = {
building: `Fewer than ~${MIN_MATURED} matured setups so far — too few to compare.`,
'no-backtest': 'Run the backtest to get a target/stop baseline to check against.',
tracking:
"Live setup outcomes are resolving in line with the backtest's target/stop model — the outcome-evaluation pipeline shows no look-ahead, config or data drift. (Checks the setup-grading pipeline, not the ATR-trail production book above.)",
drift:
"Live setup outcomes are running materially below the backtest's target/stop model — small-sample noise, a regime shift, or a live/backtest pipeline gap. Worth a look.",
};
const evaluateMutation = useMutation({ const evaluateMutation = useMutation({
mutationFn: () => triggerJob('outcome_evaluator'), mutationFn: () => triggerJob('outcome_evaluator'),
onSuccess: () => { onSuccess: () => {
@@ -46,22 +101,42 @@ export function TrackRecordPanel() {
return ( return (
<div className="space-y-6"> <div className="space-y-6">
{/* Your real, realized results come first; the strategy validation follows. */} {/* Your real, realized results come first; the strategy simulation follows. */}
<MyTradesPanel /> <MyTradesPanel />
<div className="border-t border-white/[0.06]" /> <div className="border-t border-white/[0.06]" />
<BacktestPanel /> <BacktestPanel />
<Disclosure summary="Track-record maintenance"> <Disclosure summary="Track-record maintenance">
<div className="space-y-3 pt-1"> <div className="space-y-4 pt-1">
<p className="max-w-2xl text-xs text-gray-500"> <p className="max-w-2xl text-xs text-gray-500">
The live check replays every setup against the daily bars after detection: target before stop = The live check replays every setup against the daily bars after detection: target before stop =
win, stop first = loss (both in one bar counts conservatively as a loss), neither within 30 win, stop first = loss (both in one bar counts conservatively as a loss), neither within 30
trading days = expired at 0R. Only setups whose full window has elapsed count; younger ones are trading days = expired at 0R. Only setups whose full window has elapsed count; younger ones are
still maturing (near stops resolve fast, far targets need time, so early numbers skew negative). still maturing (near stops resolve fast, far targets need time, so early numbers skew negative).
The evaluator scores <span className="text-gray-300">all</span> setups qualified or not, so The evaluator scores <span className="text-gray-300">all</span> setups qualified or not, so
unqualified ones stay a control group and runs nightly. Reset permanently clears all setups and unqualified ones stay a control group and runs nightly.
their outcomes; live setups regenerate on the next scan.
</p> </p>
{/* Diagnostic, not strategy validation: live target/stop outcomes vs the backtest's target/stop model. */}
<div className="glass-sm space-y-2 p-4">
<div className="flex flex-wrap items-center justify-between gap-x-6 gap-y-2">
<div className="flex flex-wrap items-baseline gap-x-5 gap-y-1">
<span className="text-sm text-gray-300">Setup-outcome pipeline check</span>
<span className="text-sm text-gray-400">
Live <span className={`num font-semibold ${rColor(liveAvgR)}`}>{fmtR(liveAvgR)}</span>
</span>
<span className="text-sm text-gray-400">
Backtest <span className={`num font-semibold ${rColor(btAvgR)}`}>{fmtR(btAvgR)}</span>
</span>
<span className="text-xs text-gray-500">
{liveN} matured{perf ? ` · ${perf.maturing} maturing` : ''} · qualified target/stop
</span>
</div>
<StatusChip status={status} />
</div>
<p className="text-[11px] leading-relaxed text-gray-500">{statusNote[status]}</p>
</div>
<div className="flex flex-wrap items-center gap-2"> <div className="flex flex-wrap items-center gap-2">
<Button onClick={() => evaluateMutation.mutate()} loading={evaluateMutation.isPending}> <Button onClick={() => evaluateMutation.mutate()} loading={evaluateMutation.isPending}>
{evaluateMutation.isPending ? 'Evaluating…' : 'Evaluate Now'} {evaluateMutation.isPending ? 'Evaluating…' : 'Evaluate Now'}