feat(jobs): persist each job's last run so it survives a restart

Job outcomes lived only in scheduler._job_runtime, an in-memory dict. Every
deploy wiped it, so Admin -> Jobs could report "Active" with no indication a job
had ever run or how it ended -- which is the main thing that page is for.

New job_run_state table (migration 031): one row per job, upserted on job_name.
Deliberately not history -- system_events already grows unbounded with no
retention job, and a second append-only operational table would repeat that
debt. Adding history later is purely additive.

Written from two hooks, NOT from _runtime_finish. That looked cheapest (one
function, ~40 call sites) but unit tests invoke job coroutines directly, so it
would fire detached DB writes at the real session factory throughout the suite,
and there is no testing flag to guard on.

  - An APScheduler EVENT_JOB_EXECUTED/ERROR listener covers everything the
    scheduler fires, including manual triggers. Its detached task is held in a
    module-level set (a bare create_task result can be collected mid-flight) and
    drained in the app lifespan before engine.dispose().
  - _run_pipeline persists directly, and must: pipeline steps are plain
    coroutine calls that emit no scheduler events, so the listener cannot see
    them. The step persist sits AFTER the except that swallows step errors --
    inside it, exactly the failed runs worth seeing would be skipped. The
    orchestrator persists in the finally, and the disabled early-return persists
    too, or "skipped" is silently dropped.

_persist_job_run never raises: a persistence failure must not break an otherwise
successful pipeline.

The API reports this as last_run_* and leaves runtime_* meaning strictly live
in-memory state. Reusing runtime_status would have been a regression, not a
no-op: JobControls drives the status chip from it (a job that errored eight days
ago would read "Last run error" forever instead of "Active") and picks the
rate-limit banner from it (a week-old rate limit would pin the banner
permanently). Tests pin the split.

The table starts empty; each job fills its row the next time it finishes. No
backfill from system_events, which records only warning/error outcomes under a
different status vocabulary and would invent successes that never happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-08 12:19:29 +02:00
co-authored by Claude Opus 5
parent 083c9dbf7c
commit 22bee28ac7
10 changed files with 462 additions and 19 deletions
+23 -2
View File
@@ -207,14 +207,28 @@ export function backfillTickerNames() {
}
// Jobs
export type JobCategory = 'pipeline' | 'pipeline_step' | 'scheduled' | 'manual';
export type NextRunSource = 'own_schedule' | 'via_pipeline' | 'manual_only';
export interface JobStatus {
name: string;
label: string;
enabled: boolean;
next_run_at: string | null;
via_pipeline?: boolean;
registered: boolean;
category?: JobCategory;
/** Server-assigned ordering; the payload already arrives grouped by it. */
sort_order?: [number, number];
/** Parent pipelines for a step. Many-to-many: data_collector runs in all four. */
pipelines?: string[];
/** Step names, for a pipeline row. */
steps?: string[];
next_run_at: string | null;
next_run_source?: NextRunSource;
/** For a step: the soonest enabled parent's next run, and which parent. */
via_next_run_at?: string | null;
via_next_run_job?: string | null;
running?: boolean;
/** runtime_* is live, in-memory state only — it resets when the app restarts. */
runtime_status?: string | null;
runtime_processed?: number | null;
runtime_total?: number | null;
@@ -223,6 +237,13 @@ export interface JobStatus {
runtime_started_at?: string | null;
runtime_finished_at?: string | null;
runtime_message?: string | null;
/** last_run_* is persisted and survives restarts. Kept separate from
* runtime_* so a stale error cannot pin the status chip or the banner. */
last_run_at?: string | null;
last_run_status?: string | null;
last_run_message?: string | null;
last_run_processed?: number | null;
last_run_total?: number | null;
}
export interface TriggerJobResponse {
+33 -14
View File
@@ -1,4 +1,5 @@
import { useJobs, useToggleJob, useTriggerJob } from '../../hooks/useAdmin';
import type { JobStatus } from '../../api/admin';
import { SkeletonTable } from '../ui/Skeleton';
function formatNextRun(iso: string | null): string {
@@ -29,10 +30,34 @@ function lastRunColor(status: string | null | undefined): string {
return 'text-gray-500';
}
/** One consistent answer per job: its own timer, its parent's, or "manual only".
* A step has no schedule of its own, so reporting one was the original bug. */
function NextRun({ job, labels }: { job: JobStatus; labels: Record<string, string> }) {
const muted = 'text-[11px] text-gray-500';
if (job.next_run_source === 'manual_only') {
return <span className={muted}>manual only</span>;
}
if (job.next_run_source === 'via_pipeline') {
if (!job.via_next_run_at || !job.via_next_run_job) {
return <span className={muted}>runs via pipeline</span>;
}
return (
<span className={muted}>
Next via {labels[job.via_next_run_job] ?? job.via_next_run_job}{' '}
{formatNextRun(job.via_next_run_at)}
</span>
);
}
if (!job.next_run_at) return null;
return <span className={muted}>Next run {formatNextRun(job.next_run_at)}</span>;
}
export function JobControls() {
const { data: jobs, isLoading } = useJobs();
const toggleJob = useToggleJob();
const triggerJob = useTriggerJob();
// Job id -> display label, so a step can name its parent pipeline.
const labels = Object.fromEntries((jobs ?? []).map((job) => [job.name, job.label]));
const anyJobRunning = (jobs ?? []).some((job) => job.running);
const runningJob = jobs?.find((job) => job.running);
const pausedJob = jobs?.find((job) => !job.running && job.runtime_status === 'rate_limited');
@@ -148,24 +173,18 @@ export function JobControls() {
? 'Active'
: 'Inactive'}
</span>
{job.via_pipeline ? (
<span className="text-[11px] text-gray-500">runs via pipeline</span>
) : (
job.enabled && job.next_run_at && (
<span className="text-[11px] text-gray-500">
Next run {formatNextRun(job.next_run_at)}
</span>
)
)}
{job.enabled && <NextRun job={job} labels={labels} />}
{!job.registered && (
<span className="text-[11px] text-red-400">Not registered</span>
)}
</div>
{!job.running && job.runtime_finished_at && (
<div className={`mt-1 text-[11px] ${lastRunColor(job.runtime_status)}`}>
Last run {formatAgo(job.runtime_finished_at)}
{job.runtime_status ? ` · ${job.runtime_status}` : ''}
{job.runtime_message ? `${job.runtime_message}` : ''}
{/* Persisted, so this survives a deploy — unlike runtime_*,
which the status chip above still reads for live state. */}
{!job.running && job.last_run_at && (
<div className={`mt-1 text-[11px] ${lastRunColor(job.last_run_status)}`}>
Last run {formatAgo(job.last_run_at)}
{job.last_run_status ? ` · ${job.last_run_status}` : ''}
{job.last_run_message ? `${job.last_run_message}` : ''}
</div>
)}
{job.running && (