Files
signal-platform/docs/dolt-integration-plan.md
T
dennisthiessenandClaude Opus 5 e1607ddbff
Deploy / lint (push) Successful in 11s
Deploy / test (push) Successful in 1m24s
Deploy / deploy (push) Successful in 41s
fix: don't double-report a failed SEC run, and correct the rollback doc
Three follow-ups from review of the A6 commits.

The cache-summary re-finalize raised a second durable system event on a failed
run: _runtime_finish emits for `error`/`rate_limited`, and the dedup key
includes the message, so "SEC unavailable" and "SEC unavailable · cache 511 · 2
score inputs changed" landed as two unacknowledged Admin events. Adds
`emit_event` so a re-finalize that only rewords an outcome stays silent, with a
regression test asserting exactly one event.

The rollback section still claimed disabling the SEC job freezes the cache — the
opposite of what the same page says two lines earlier, and of what the code now
does. Rewritten: there is no Admin cache-off switch, restoring `fundamental_data`
alone is temporary because the next run rebuilds it from the same snapshots and
code, and a real freeze means stopping the service.

Remaining "shadow" wording: the two import jobs have never been shadow since
activation, so `_run_shadow_import` -> `_run_source_import`, its section heading,
the deployment doc's job label, and the plan doc's "production switch remains"
handoff paragraph are all brought up to date.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:46:34 +02:00

32 KiB
Raw Blame History

Dolt bulk-data integration — implementation plan

Status: approved 2026-07-21, revised through four review rounds; direction: KISS backend, UI value first. Hand-off document for the implementing agent; self-contained.

Objective

Replace the free-tier fundamentals APIs (FMP, Finnhub, Alpha Vantage) with bulk data: SEC Company Facts for fundamentals, the DoltHub earnings repo for the earnings calendar/history, and — later, independently — the DoltHub stocks repo for historical OHLCV. PostgreSQL stays the production system of record.

Delivery order: two independent workstreams.

  • Workstream A (build first): SEC fundamentals + Dolt earnings + API v1 + FundamentalsPanel + decommission FMP/Finnhub/Alpha Vantage. Valuation uses the existing Alpaca closes already in ohlcv_records. This alone achieves the goal (killing the quota-limited APIs) and delivers all the UI value.
  • Workstream B (later, optional until needed): replace historical OHLCV with the Dolt stocks repo. The most complex machinery (4.7 GB clone, split adjustment, source-bar table, reconciliation) lives here and blocks nothing in A.

Guiding principle: KISS. Plain daily importers with staging and atomic promotion — no forensic replay, no permanent archive store, no conflict tables, no aggregate tables. Engineering budget goes into the UI (quarter trends, peer comparison). Deferred until a concrete need: exact source replay, point-in-time backtest enforcement, fundamental metrics in scoring.

Non-negotiables

  • The application never queries Dolt/DoltHub or SEC at request time. All access is batch import → PostgreSQL. If a sync fails or the source is unchanged, production continues on the last successfully imported data.
  • Do not replace PostgreSQL with Dolt/Doltgres. Never commit to the upstream clones.
  • No owned SEC Dolt repo: SEC JSON is normalized straight into PostgreSQL.
  • Scoring code is unchanged, but swapping the data source changes production behavior: app/services/scoring_service.py (~line 450) scores pe_ratio / revenue_growth / earnings_surprise from fundamental_data, so new definitions change rankings even with identical code. Cutover of fundamental_data population requires the score-parity gate (phase A5) — never silently. All new metrics are display-only.
  • Intraday (10:0015:00), near-close (15:30) and after-close (16:45) pipelines stay on Alpaca unchanged.

Data sources

  1. SEC Company Facts + submissions bulk files (free, no key; costs bandwidth, CPU and disk — optimize accordingly) — XBRL facts per issuer (CIK), not per ticker. Tickers resolve to a CIK via SEC company_tickers.json (multi-class issuers like GOOGL/GOOG share one CIK and one set of fundamentals). Submissions also supply the SIC code (peer grouping) and acceptanceDateTime. Handle unit variants and fiscal-period alignment (derive Q4 = FY Q1..Q3 where needed). Amendments: retain every accession immutably; readers select the newest valid accepted_at snapshot per reporting period at read time. No flags, no mutation.
  2. post-no-preference/earnings (DoltHub) — announcement date, BMO/AMC session (partial), period end, EPS estimate/actual, surprise history. Small clone. scripts/import_dolthub_earnings.py is a research/SQLite importer — reuse its normalization/alignment logic (calendar↔EPS-history monotonic alignment, SUE scaling) but write a production PostgreSQL importer; do not extend the script.
  3. post-no-preference/stocks (DoltHub, workstream B) — daily raw OHLCV (unadjusted), symbol metadata, splits, dividends. Publishes ~01:30 ET the following calendar day. Clone is ~4.7 GB.

Licensing (phase A0) — DECIDED 2026-07-22: post-no-preference/earnings is approved for private/internal ingestion under CC BY-SA 4.0. Conditions the A2 importer must honor: preserve the upstream license, attribution, and transformation notes (retain a CC BY-SA 4.0 reference + attribution to post-no-preference/earnings and a note of the transformations applied — e.g. in a repo NOTICE/attribution file and the importer module); no public API, bulk export, or redistribution of the data; re-review licensing before any public or commercial access. The post-no-preference/stocks repo (workstream B) is not covered here and will be reviewed separately if B begins.

Schema

Migration 026 (workstream A) — current head: 025_trade_setup_scan_run_id:

  • data_import_runs — lean: id, source (sec_facts | dolt_earnings | dolt_stocks), revision (Dolt commit hash, or SEC archive SHA-256), status (running/validated/promoted/no_op/failed), source_max_date, row_counts JSON, validation JSON (includes reconciliation/discrepancy summaries — no separate conflicts table; details go to structured logs), started_at, completed_at, error_details. One run per source at a time (Postgres advisory lock keyed by source).
  • earnings_events — ticker_id, announce_date, session (bmo/amc/unknown), period_end, eps_estimate, eps_actual, source, import_run_id. Unique (ticker_id, announce_date). Rescheduling: within each promotion transaction, delete this source's future-dated rows (announce_date > today) and re-insert from the new snapshot, so moved or cancelled dates never linger. Past rows (results) are never deleted.
  • tickers — add nullable cik, sic, sic_description (from SEC submissions / company_tickers.json; refreshed by the SEC import; multi-class tickers share values). The only ticker↔issuer join point.
  • fundamental_snapshotsCIK-keyed, one immutable row per accession: cik, accession (unique), form, filed_date, accepted_at (kept although PIT enforcement is deferred — one timestamp now vs painful retrofit later), period_start, period_end, fiscal_year, fiscal_period (the filing's own dei/us-gaap period identity — required to align non-calendar fiscal years and to derive discrete quarters from cumulative facts), and the price-independent raw facts so metrics are recomputable. Store facts as the filing reports them, not as derived quarters: duration facts (revenue, net income, diluted EPS, CFO, capex, EBITDA inputs) retain the filing's normalized cumulative YTD/FY value for the (period_start → period_end) span; balance-sheet facts (cash+ST investments, total debt, shares outstanding) are period-end values. shares_outstanding is a point-in-time count (dei:EntityCommonStockSharesOutstanding), not the weighted-average diluted share count — both consumers (est. market cap, YoY dilution) want a point-in-time value. Nothing derived is frozen into a row: discrete quarters (10-Q YTD deltas, Q4 = FY Q1..Q3), TTM, YoY and the four-period metric histories are all computed at read time by picking the newest valid accepted_at snapshot for each required period — so non-calendar fiscal years resolve correctly and a later amendment to a prior quarter is reflected automatically without ever storing a stale derived quarter. Readers pick the newest valid accepted_at per period; history powers the UI reference comparisons and deterministic reads.
  • Keep fundamental_data (app/models/fundamental.py) as the latest-value compat cache, repopulated by the daily SEC job — but only after the phase-A5 parity gate.

Migration 027 (workstream B, written when B starts):

  • ohlcv_source_bars — source-truth bar table, required because ohlcv_records allows one row per (ticker_id, date) (app/models/ohlcv.py:12) and Alpaca ingestion upserts it in place (app/services/price_service.py:82) — Dolt and Alpaca bars cannot coexist there. Holds Dolt raw (unadjusted) bars only — Alpaca bars are already split-adjusted at the provider (app/providers/alpaca.py:77 requests Adjustment.SPLIT) and live exclusively in ohlcv_records. Columns: source (dolt), adjustment (raw — explicit), ticker_id, date, OHLCV, import_run_id; unique (source, ticker_id, date). Changed bars are counted in the run's validation JSON and logged before overwrite.
  • corporate_actions — ticker_id, type (split/dividend), ex_date, ratio/amount, source, import_run_id. Unique (ticker_id, type, ex_date).
  • ohlcv_records — add nullable import_run_id FK and source text (default 'alpaca').

Import framework

Every importer: idempotent per revision (same Dolt commit / archive checksum → no_op, zero row changes); stage into a representation outside the live tables first (in-memory for the small workstream-A sources; a file/table handle is fine if workstream B ever needs it); promotion in one transaction; safe to retry; a failed or unchanged run leaves the current dataset untouched. Record every attempt in data_import_runs.

SEC access requirements (operational safeguards, per SEC fair-access policy): send an identifying User-Agent with a contact email on every request; stay far below the 10 req/s limit (the bulk endpoints need only a handful of requests per run); exponential backoff on 429; a 403 means the User-Agent or request pattern is wrong — alert and stop, never retry-loop. See SEC developer resources (https://www.sec.gov/about/developer-resources).

Reproducibility scope (deliberately limited): the normalized snapshots in PostgreSQL are the durable record. Keep only the last ~2 SEC archives on disk for debugging. Byte-level replay of old runs is out of scope until a concrete need.

Dolt access: dolt pull on the persistent clone, record the resulting commit hash, read via dolt sql -r csv (no long-running sql-server). The scheduler shares one event loop with the API (app/scheduler.py:73) — run dolt/unzip/download subprocesses via asyncio.create_subprocess_exec (or an executor), never blocking calls. Check free disk space before pulling; alert and skip if below threshold.

Deployment constraints: deploy is rsync --delete of the repo tree (.gitea/workflows/deploy.yml:127), so clones and archives must live outside the deployment path — an env-configured persistent directory (e.g. DOLT_DATA_DIR=/var/lib/signal-platform/dolt). The dolt binary is a new prod runtime dependency: install it once with the version-pinned provisioner in deploy/provision_fundamentals.sh; operational steps are in docs/fundamentals-deployment.md. The clone is reproducible from DoltHub; the normalized PostgreSQL rows remain part of the normal database backup.

Validation gates (block promotion, raise an alert via the existing system-events path): source freshness as expected; tracked-universe coverage; no duplicate business keys; fundamental units/periods consistent; row-count deltas within reason; upstream schema change stops promotion. Workstream B adds: OHLC sanity (high ≥ open/close/low, low ≤ open/close/high, volume ≥ 0); no unexplained split discontinuities.

Split adjustment (workstream B): ohlcv_source_bars + corporate_actions are the source of truth; canonical ohlcv_records is generated from them to match Alpaca Adjustment.SPLIT, selecting only adjustment = raw rows as input so adjustment is applied exactly once. A newly published split rewrites the symbol's entire adjusted history — treat whole-symbol rewrites as a normal import event (exempt that symbol from the row-count-delta gate for that run) and stamp rows with the import_run_id (a backtest↔prod parity guard exists; changed history changes backtests).

Scheduling (app/scheduler.pySCHEDULE_DEFAULTS / _CRON_JOBS, ~line 1451)

Follow the existing pattern: cron strings in SystemSettings via app/services/settings_store.py, day-of-week as names never numbers, logging via _log_event.

Workstream A:

  • Dolt earnings import: daily ~02:30 ET (with the future-row replacement above).
  • SEC fundamentals job: daily ~04:00 ET. One job, three steps: (a) detect a composite revision from the latest EDGAR daily-index date, the exact tracked index rows, and the tracked-universe fingerprint — an unchanged revision is a no_op before Company Facts are fetched; (b) when changed, fetch and parse Company Facts only for tracked-universe CIKs that filed, plus full available history for the first run or a newly added issuer, through validation→atomic promotion. Universe resolution and the exact index inputs are cached during revision detection and reused during staging; CIKs are resolved from company_tickers.json without writes until promotion; (c) always, locally, and only after production activation (the phase-A5 parity approval): refresh the legacy fundamental_data fields and mark affected cached fundamental scores stale. Sources differ per field — do not assume all five come from SEC: pe_ratio and market_cap from the newest valid snapshots × latest PostgreSQL close, each with its own formula — pe_ratio = latest close / TTM diluted EPS; market_cap = issuer-wide shares outstanding × latest close; revenue_growth from the snapshots alone; earnings_surprise and next_earnings_date from earnings_events (the Dolt earnings feed — these two do not exist in SEC facts). Before activation the job imports snapshots only (shadow). Step (c) must run identically when SEC is unreachable — prices move daily even when filings don't, and the earnings-derived fields already live in PostgreSQL. The new API valuation object is not stored anywhere — it is computed at request time (below). No valuation cache or table exists.

Workstream B:

  • Dolt OHLCV+splits pull/import: 0 2 * * tue-sat ET. If source_max_date is not fresh, retry hourly until ~06:00, then give up quietly. After a successful import, reconcile the previous session's Dolt-derived bars against the Alpaca bars; summary into the run's validation JSON, details to logs.
  • Move schedule_daily_pipeline_cron (morning refresh) from 0 2 * * * to 0 3 * * * (only needed once the 02:00 slot is taken by the OHLCV pull). Late Dolt publication is a non-event: the canonical scan runs at 15:30 on Alpaca, so the morning pipeline runs normally even when the import hasn't landed — no gating, no defensive coupling.

Metrics catalog (curated — TTM basis)

Snapshots store price-independent per-period facts (the "snapshot" column below means derived from stored snapshots, assembled across periods at read time — see Schema — not frozen at import); price-dependent ratios are never frozen into snapshots and have no storage location at all: the API computes them at request time from the stored snapshots + the latest ohlcv_records close (both already in PostgreSQL, so this works identically when SEC is unreachable). The only stored price-dependent values are the legacy fundamental_data fields that scoring already reads, refreshed daily by step (c) after activation.

Metric Definition Where computed
Revenue growth YoY TTM revenue vs prior TTM snapshot
EPS growth YoY TTM diluted EPS vs prior TTM snapshot
Operating margin + 4q trend TTM operating income / revenue snapshot
FCF margin (TTM CFO capex) / revenue snapshot
Net debt total debt (cash + ST investments); positive = net debt snapshot
Net debt / EBITDA net debt / TTM EBITDA snapshot
Share count Δ YoY shares outstanding vs year ago snapshot
Trailing P/E price / TTM diluted EPS request time
FCF yield TTM FCF / est. market cap request time
Est. market cap issuer-wide shares outstanding × ticker price request time
Earnings surprise history last 4+ from earnings_events query

Market cap is an estimate (issuer-wide shares outstanding × one ticker's price — approximate for multi-class issuers). Share count comes from a single consolidated value, not a class sum: prefer the one dei:EntityCommonStockSharesOutstanding cover-page fact; if absent (e.g. Alphabet) fall back to us-gaap:CommonStockSharesOutstanding at period end. companyfacts is non-dimensional, so class-specific facts can't be summed reliably — never do that, and never substitute weighted-average/diluted shares; if conflicting values remain, store null. Label it "est." in the UI and round aggressively rather than withholding it; false precision is the failure mode, not the approximation.

Units follow existing app conventions: percentages are percentage points (21.0 = 21%), P/E and net-debt/EBITDA are multiples, market cap and net debt are dollars.

Deliberately excluded: ROIC (invested-capital/NOPAT normalization too noisy), gross margin (COGS tagging too inconsistent), any new composite score.

Peer comparison (read-time only)

  • Peer group = tracked-universe issuers sharing the first two SIC digits, deduplicated by CIK — GOOG and GOOGL are one issuer, one observation, in medians, percentiles and peer_count.
  • Computed at read time from current snapshots — no aggregate tables until performance demonstrates a need.
  • Medians exclude null/invalid values. Fewer than 5 valid peer issuers → omit the peer result entirely rather than showing a misleading universe comparison.
  • Percentile direction respects metric polarity (higher-is-better for FCF yield, lower-is-better for P/E and leverage).

API contract (additive v1)

Every existing top-level field is preserved unchanged (name, type, position) — backend and frontend ship independently, no breaking interval. New objects, exact names and types:

{
  // ...all existing legacy fields, unchanged...
  "earnings": {
    "next": {"date": "YYYY-MM-DD", "session": "bmo|amc|unknown", "days_until": 12} | null,
    "recent": [   // newest first, max 4, may be empty
      {"announce_date": "YYYY-MM-DD", "period_end": "YYYY-MM-DD|null",
       "eps_estimate": 1.02|null, "eps_actual": 1.10|null, "surprise_pct": 7.8|null}
    ]
  },
  "metrics": [    // fixed row set — every key always present, value null when unavailable
    {
      "key": "revenue_growth_yoy",   // revenue_growth_yoy | eps_growth_yoy | operating_margin |
                                     // fcf_margin | net_debt | net_debt_to_ebitda | share_count_change_yoy
      "value": 18.0,                 // number | null — pp / multiples / dollars per units above
      "history": [                   // oldest→newest, max 4 points, [] when unavailable
        {"period_end": "YYYY-MM-DD", "value": 8.0}
      ],
      "industry": {                  // object | null — null when < 5 valid peer issuers (CIK-deduped)
        "label": "SIC 73 peers",     // truthful 2-digit group label — grouping IS 2-digit,
                                     // so no 4-digit description like "Prepackaged Software"
        "median": 11.0,
        "favorable_percentile": 82,  // 0-100, polarity-aware (higher = more favorable)
        "peer_count": 12             // issuers, not tickers
      },
      "period_end": "YYYY-MM-DD|null",
      "filed_date": "YYYY-MM-DD|null",
      "source": "sec|dolt|legacy_api"
    }
  ],
  "valuation": {  // object | null (null until SEC snapshots exist, phase A3); same industry sub-object rules
    // computed at REQUEST TIME from stored snapshots + latest PostgreSQL close —
    // no valuation cache or table; unaffected by SEC availability
    "pe": 29.2|null, "fcf_yield": 3.8|null,
    "market_cap_est": 1.2e9|null,    // estimated — UI labels "est."
    "pe_industry": {...}|null, "fcf_yield_industry": {...}|null,
    "price_date": "YYYY-MM-DD"       // close used for the ratios
  }
}

Null/freshness semantics: absent data is null with the row still present (the UI shows "n/a", never hides rows); every metric carries its own source, period and filing date — no panel-wide source label. The objects may serve partial data during rollout (e.g. earnings live, metrics still legacy_api); the shape never changes.

UI — frontend/src/components/ticker/FundamentalsPanel.tsx

One distinctive visual device — the Reference Rails — in an otherwise restrained panel. Preserve the app's dark glass styling and numeric typography.

Fundamentals
Growth accelerating · margins improving · valuation priced above peers
Next earnings  Aug 3 · AMC       EPS surprises  ▂ ▅ ▃ ▆

Operating trend              less favorable ← ref → more favorable
Revenue growth                                             18%
───────────────│━━━━●                   +3pp vs prior · accelerating
Share count YoY                                           −1.7%
───────────────│━━━━●                                    buying back

Valuation & balance           less favorable ← median → more favorable
P/E                                                       29.2×
────●━━━━━━━━━━│────                  priced above peers · median 23.5× · 12 peers
  • Growth and margins: horizontal rails compare the latest value with the prior quarter or prior-period average; share-count YoY compares with zero. The rail is normalized so right is always more favorable, including buybacks.
  • P/E, FCF yield and leverage: horizontal favorable-percentile rails with a peer-median marker. No decorative rail when industry is null (< 5 peers).
  • Every row keeps the exact value and one deterministic comparison caption; missing values render n/a, and insufficient peers render peers n/a.
  • Earnings: four bars around a shared zero baseline — cyan beats, coral misses, gray unavailable — plus next date and BMO/AMC session countdown.
  • Accessibility: color is always paired with text; neutral/ambiguous stays gray; rails and earnings bars expose complete ARIA descriptions.
  • Remove the hard-coded "FMP" source label; surface filing and price-date provenance in the footer.

Deterministic reads — one shared rule set. Implement as a single function with named constants; the metric reads and the header sentence use identical outputs. No LLM, no new composite score. Defaults (tunable constants, not scattered literals):

  • A series read requires ≥ 3 periods; otherwise show "—" and no read.
  • Growth metrics (pp): latest prior ≥ +2.0 → "accelerating"; ≤ 2.0 → "decelerating"; else "steady".
  • Margins (latest vs mean of prior periods, pp): ≥ +1.0 → "improving"; ≤ 1.0 → "deteriorating"; else "stable" (phrased "above/below own average" where the layout calls for it).
  • Share count YoY: > +1.0% → "N% dilution"; < 1.0% → "buying back"; else "flat".
  • Peer-relative: favorable_percentile ≥ 60 → favorable ("above peers"); ≤ 40 → adverse ("priced above peers" for P/E, "elevated leverage" for net-debt/EBITDA); else "in line".
  • Header sentence: join the growth read, margin read and peer-relative valuation read with " · ", omitting segments that have no read (e.g. "Growth accelerating · margins stable · valuation above industry median"). Segment sources are fixed: growth = revenue growth read; margins = operating margin read; valuation = P/E peer-relative read, falling back to FCF yield when P/E is null. This keeps the header unambiguous when sibling metrics (EPS vs revenue growth, P/E vs FCF yield) point in different directions.

Decommissioning (end of workstream A)

Remove completely: FMP (app/providers/fmp.py), Finnhub + Alpha Vantage (app/providers/fundamentals_chain.py), their config keys (app/config.py), and their wiring in app/scheduler.py, app/routers/ingestion.py, app/services/ticker_universe_service.py. Retain: Alpaca (prices), FRED, sentiment provider, Telegram. Note: decommissioning does not depend on workstream B — Alpaca remains the price source throughout.

Rollout

Workstream A:

  • A0. License review DONE (earnings approved for private/internal use under CC BY-SA 4.0, no redistribution — see Licensing above). The Dolt version, persistent DOLT_DATA_DIR, clone, and production checks are captured in deploy/provision_fundamentals.sh and docs/fundamentals-deployment.md.
  • A1. Migration 026, import-run framework.
  • A2. Earnings ingestion in shadow (writes earnings_events, prod untouched); verify forward-calendar coverage and rescheduling behavior.
  • A3. SEC daily job in shadow (writes fundamental_snapshots). Primary technical risk here: Q4 derivation and fiscal-period alignment — non-calendar fiscal years, restatements/amendments, and XBRL unit/dimension variants; budget accordingly.
  • A4. API v1 + FundamentalsPanel + peer comparison — served from snapshots and earnings_events, independent of the scoring cutover (the additive API supports partial data). UI value ships before anything touches scoring inputs.
  • A5. Score-parity gatefundamental_data cutover: compute candidate pe_ratio/revenue_growth/earnings_surprise from SEC/Dolt side by side with the API values across the tracked universe, report per-field deltas and resulting fundamental-score/ranking changes, require explicit approval. Definition changes (e.g. TTM vs provider convention) called out, not averaged away. Status 2026-07-24: the gate has been exercised and the evidence supports approval — see the handoff section below. Step (c) is implemented behind the default-off fundamental_data_sec_dolt_cutover_enabled SystemSetting; the remaining production action is flipping that switch on and observing it.
  • A6. DONE 2026-08-07. FMP/Finnhub/Alpha Vantage removed, along with the weekly fundamental_collector job, the A5 cutover toggle (SEC+Dolt is now the unconditional path) and the parity report. Migration 029 tombstones the two behavior-bearing settings rows for the rollback window; the archived parity bundles stay as the A5 evidence trail.

Workstream B (independent, start when wanted):

  • B0. Stocks clone (~4.7 GB) provisioned; migration 027.
  • B1. OHLCV + split adjustment in shadow (writes ohlcv_source_bars only; Alpaca keeps owning ohlcv_records); historical backfill.
  • B2. Reconciliation window (≥ 2 weeks) vs Alpaca; review validation summaries.
  • B3. Promote Dolt as historical OHLCV source (canonical rebuilt from raw source bars + splits); morning pipeline → 03:00.

Test plan

  • Daily SEC job: changed revision imports; unchanged conditional-HTTP check is a no_op with zero downloads and zero row changes; validation failure leaves production untouched; source unavailable still runs the local fundamental_data refresh (step c, post-activation); before activation the job never writes fundamental_data.
  • Valuation endpoint returns identical values with SEC reachable and unreachable (pure PostgreSQL computation); no valuation rows exist in any table.
  • Multi-class tickers resolve to the same CIK snapshots; peer medians and peer_count are CIK-deduplicated (GOOG+GOOGL = one observation).
  • Earnings rescheduling: a moved future date replaces the old row atomically; a cancelled date disappears; historical results are never touched.
  • Amendment selection: for a period with multiple accessions, the newest valid accepted_at wins at read time; older rows remain unchanged.
  • History arrays are chronological, ≤ 4 points.
  • Percentage-point units stay compatible with existing formatters and scoring inputs.
  • Deterministic reads: threshold boundary cases (exactly +2.0pp, exactly 60th percentile) resolve per the stated rules; header uses identical outputs and falls back from P/E to FCF yield for the valuation segment when P/E is null.
  • Peer comparison disappears below 5 peer issuers; favorable-percentile direction correct for both polarities.
  • Workstream B: split-adjusted OHLCV matches Alpaca on representative normal / split / reverse-split symbols.
  • UI states: positive, adverse, neutral, insufficient history, insufficient peers; mobile layout; non-color accessibility.
  • Unit, integration, scheduler and frontend suites pass.

Acceptance criteria

  • App works normally with Dolt/DoltHub/SEC unreachable.
  • Re-running the same revision: zero duplicate or changed rows.
  • Upcoming earnings dates present and timely for the tracked universe — the forward calendar is the hardest thing to replace and gates decommissioning.
  • Coverage meets the tracked-universe target; scheduler runs cleanly with FMP/Finnhub/AV keys removed from the environment.
  • Score-parity diff reviewed and approved before fundamental_data cutover.
  • Scheduled imports never block the API event loop.

Handoff — remaining work after the A5 parity investigation (2026-07-24)

The 2026-07-23 parity report surfaced coverage gaps and wrong values; a nine-pass investigation traced every one to parser/identity bugs (not source data), fixed them, and reparsed production twice. Full evidence trail: reports/fundamentals-parity-20260723-findings.md (root causes, decisions, validation) plus the before/after reports (fundamentals-parity-20260723T… / …20260724T….json). Post-fix: candidate scores 504 of 511 vs legacy's 507 (gap = PSKY/Q new registrants + FITB, all explained); revenue-growth agreement 0.0038 median abs delta where both exist. Dennis reviewed the evidence 2026-07-24 and directed proceeding to cutover.

Task 1 — A5 activation: DONE. Implemented 2026-07-24, switched on and observed in production, and made unconditional by A6 (2026-08-07) — there is no longer a switch, an Admin card, or a weekly legacy collector to skip. The local refresh of fundamental_data derives pe_ratio and market_cap from newest valid snapshots × latest PostgreSQL close, revenue_growth from snapshots, and earnings_surprise/next_earnings_date from earnings_events; it marks affected cached fundamental scores stale and runs identically when SEC is unreachable. It consumes fundamentals_derivation.derive() outputs, NOT raw snapshot fields — that path carries the split guard (ttm_diluted_eps nulls when contaminated, with ttm_diluted_eps_caveat) and the multi-class share fallback (shares_outstanding + shares_outstanding_estimated). See docs/fundamentals-deployment.md for current operations and rollback.

Task 2 — A6 decommissioning: DONE 2026-08-07. The cutover ran on and was observed in production, so the legacy providers, their config/env keys, the weekly collector job and the parity report were all removed. Two consequences to carry: (1) fundamental_data now has no provider fallback — recovery is restore-from-backup; (2) disabling SEC Fundamentals Import stops the SEC fetch only, because the local cache refresh was deliberately moved outside the job-enable check. Remaining follow-up: delete the migration-029 tombstone rows once the rollback window closes.

Known caveats to carry (documented in the findings report, not bugs to fix):

  • KLAC-class post-filing splits: P/E wrong until the next 10-Q; undetectable from snapshots. Workstream B's corporate_actions table is the natural future fix.
  • BRK-B: no share count exists anywhere in companyfacts → no market cap, correctly.
  • FITB: unscored (split guard + no taggable revenue) — the one name that lost its score relative to legacy; composite renormalises.
  • Share-change guard at 25% nulls P/E for stock-funded M&A too (COF, WAT…); revisit only if the ~3% universe hit-rate proves painful.
  • sec_cik_overrides SystemSetting pins XOM → 34088 (applied in prod); the no_xbrl_filings SystemEvent says when a new pin is needed.
  • After any future parser change, stored rows need scripts/reparse_fundamentals.py (dry-run default; --apply rewrites) — snapshots are otherwise immutable.

Deferred (explicitly, until a concrete need appears)

  • Workstream B itself is deferred relative to A and blocks nothing in A.
  • Exact byte-level source replay of historical imports; permanent archive store.
  • Point-in-time backtest enforcement (accepted_at is stored now; derivation and backtest visibility rules are built only when fundamentals enter scoring/backtesting).
  • Fundamental metrics in the score; sector-relative scoring.
  • Aggregate/rollup tables for peer statistics.
  • Any valuation cache or table (request-time computation from snapshots + latest close suffices).
  • A dedicated conflicts table (validation JSON + logs suffice).