31 KiB
Dolt bulk-data integration — implementation plan
Status: approved 2026-07-21, revised through four review rounds; direction: KISS backend, UI value first. Hand-off document for the implementing agent; self-contained.
Objective
Replace the free-tier fundamentals APIs (FMP, Finnhub, Alpha Vantage) with bulk data: SEC Company Facts for fundamentals, the DoltHub earnings repo for the earnings calendar/history, and — later, independently — the DoltHub stocks repo for historical OHLCV. PostgreSQL stays the production system of record.
Delivery order: two independent workstreams.
- Workstream A (build first): SEC fundamentals + Dolt earnings + API v1 +
FundamentalsPanel + decommission FMP/Finnhub/Alpha Vantage. Valuation uses the
existing Alpaca closes already in
ohlcv_records. This alone achieves the goal (killing the quota-limited APIs) and delivers all the UI value. - Workstream B (later, optional until needed): replace historical OHLCV with the Dolt stocks repo. The most complex machinery (4.7 GB clone, split adjustment, source-bar table, reconciliation) lives here and blocks nothing in A.
Guiding principle: KISS. Plain daily importers with staging and atomic promotion — no forensic replay, no permanent archive store, no conflict tables, no aggregate tables. Engineering budget goes into the UI (quarter trends, peer comparison). Deferred until a concrete need: exact source replay, point-in-time backtest enforcement, fundamental metrics in scoring.
Non-negotiables
- The application never queries Dolt/DoltHub or SEC at request time. All access is batch import → PostgreSQL. If a sync fails or the source is unchanged, production continues on the last successfully imported data.
- Do not replace PostgreSQL with Dolt/Doltgres. Never commit to the upstream clones.
- No owned SEC Dolt repo: SEC JSON is normalized straight into PostgreSQL.
- Scoring code is unchanged, but swapping the data source changes production
behavior:
app/services/scoring_service.py(~line 450) scores pe_ratio / revenue_growth / earnings_surprise fromfundamental_data, so new definitions change rankings even with identical code. Cutover offundamental_datapopulation requires the score-parity gate (phase A5) — never silently. All new metrics are display-only. - Intraday (10:00–15:00), near-close (15:30) and after-close (16:45) pipelines stay on Alpaca unchanged.
Data sources
- SEC Company Facts + submissions bulk files (free, no key; costs bandwidth,
CPU and disk — optimize accordingly) — XBRL facts per issuer (CIK), not per
ticker. Tickers resolve to a CIK via SEC
company_tickers.json(multi-class issuers like GOOGL/GOOG share one CIK and one set of fundamentals). Submissions also supply the SIC code (peer grouping) andacceptanceDateTime. Handle unit variants and fiscal-period alignment (derive Q4 = FY − Q1..Q3 where needed). Amendments: retain every accession immutably; readers select the newest validaccepted_atsnapshot per reporting period at read time. No flags, no mutation. post-no-preference/earnings(DoltHub) — announcement date, BMO/AMC session (partial), period end, EPS estimate/actual, surprise history. Small clone.scripts/import_dolthub_earnings.pyis a research/SQLite importer — reuse its normalization/alignment logic (calendar↔EPS-history monotonic alignment, SUE scaling) but write a production PostgreSQL importer; do not extend the script.post-no-preference/stocks(DoltHub, workstream B) — daily raw OHLCV (unadjusted), symbol metadata, splits, dividends. Publishes ~01:30 ET the following calendar day. Clone is ~4.7 GB.
Licensing (phase A0) — DECIDED 2026-07-22: post-no-preference/earnings is
approved for private/internal ingestion under CC BY-SA 4.0. Conditions the A2
importer must honor: preserve the upstream license, attribution, and transformation
notes (retain a CC BY-SA 4.0 reference + attribution to post-no-preference/earnings
and a note of the transformations applied — e.g. in a repo NOTICE/attribution file
and the importer module); no public API, bulk export, or redistribution of the
data; re-review licensing before any public or commercial access. The
post-no-preference/stocks repo (workstream B) is not covered here and will be
reviewed separately if B begins.
Schema
Migration 026 (workstream A) — current head: 025_trade_setup_scan_run_id:
data_import_runs— lean: id, source (sec_facts|dolt_earnings|dolt_stocks), revision (Dolt commit hash, or SEC archive SHA-256), status (running/validated/promoted/no_op/failed), source_max_date, row_counts JSON, validation JSON (includes reconciliation/discrepancy summaries — no separate conflicts table; details go to structured logs), started_at, completed_at, error_details. One run per source at a time (Postgres advisory lock keyed by source).earnings_events— ticker_id, announce_date, session (bmo/amc/unknown), period_end, eps_estimate, eps_actual, source, import_run_id. Unique (ticker_id, announce_date). Rescheduling: within each promotion transaction, delete this source's future-dated rows (announce_date > today) and re-insert from the new snapshot, so moved or cancelled dates never linger. Past rows (results) are never deleted.tickers— add nullablecik,sic,sic_description(from SEC submissions /company_tickers.json; refreshed by the SEC import; multi-class tickers share values). The only ticker↔issuer join point.fundamental_snapshots— CIK-keyed, one immutable row per accession: cik, accession (unique), form, filed_date, accepted_at (kept although PIT enforcement is deferred — one timestamp now vs painful retrofit later), period_start, period_end, fiscal_year, fiscal_period (the filing's owndei/us-gaapperiod identity — required to align non-calendar fiscal years and to derive discrete quarters from cumulative facts), and the price-independent raw facts so metrics are recomputable. Store facts as the filing reports them, not as derived quarters: duration facts (revenue, net income, diluted EPS, CFO, capex, EBITDA inputs) retain the filing's normalized cumulative YTD/FY value for the (period_start → period_end) span; balance-sheet facts (cash+ST investments, total debt, shares outstanding) are period-end values.shares_outstandingis a point-in-time count (dei:EntityCommonStockSharesOutstanding), not the weighted-average diluted share count — both consumers (est. market cap, YoY dilution) want a point-in-time value. Nothing derived is frozen into a row: discrete quarters (10-Q YTD deltas, Q4 = FY − Q1..Q3), TTM, YoY and the four-period metric histories are all computed at read time by picking the newest valid accepted_at snapshot for each required period — so non-calendar fiscal years resolve correctly and a later amendment to a prior quarter is reflected automatically without ever storing a stale derived quarter. Readers pick the newest valid accepted_at per period; history powers the UI reference comparisons and deterministic reads.- Keep
fundamental_data(app/models/fundamental.py) as the latest-value compat cache, repopulated by the daily SEC job — but only after the phase-A5 parity gate.
Migration 027 (workstream B, written when B starts):
ohlcv_source_bars— source-truth bar table, required becauseohlcv_recordsallows one row per (ticker_id, date) (app/models/ohlcv.py:12) and Alpaca ingestion upserts it in place (app/services/price_service.py:82) — Dolt and Alpaca bars cannot coexist there. Holds Dolt raw (unadjusted) bars only — Alpaca bars are already split-adjusted at the provider (app/providers/alpaca.py:77requestsAdjustment.SPLIT) and live exclusively inohlcv_records. Columns: source (dolt), adjustment (raw— explicit), ticker_id, date, OHLCV, import_run_id; unique (source, ticker_id, date). Changed bars are counted in the run's validation JSON and logged before overwrite.corporate_actions— ticker_id, type (split/dividend), ex_date, ratio/amount, source, import_run_id. Unique (ticker_id, type, ex_date).ohlcv_records— add nullableimport_run_idFK andsourcetext (default'alpaca').
Import framework
Every importer: idempotent per revision (same Dolt commit / archive checksum →
no_op, zero row changes); stage into a representation outside the live tables
first (in-memory for the small workstream-A sources; a file/table handle is fine
if workstream B ever needs it); promotion in one transaction;
safe to retry; a failed or unchanged run leaves the current dataset untouched.
Record every attempt in data_import_runs.
SEC access requirements (operational safeguards, per SEC fair-access policy):
send an identifying User-Agent with a contact email on every request; stay far
below the 10 req/s limit (the bulk endpoints need only a handful of requests per
run); exponential backoff on 429; a 403 means the User-Agent or request pattern is
wrong — alert and stop, never retry-loop. See SEC developer resources
(https://www.sec.gov/about/developer-resources).
Reproducibility scope (deliberately limited): the normalized snapshots in PostgreSQL are the durable record. Keep only the last ~2 SEC archives on disk for debugging. Byte-level replay of old runs is out of scope until a concrete need.
Dolt access: dolt pull on the persistent clone, record the resulting commit hash,
read via dolt sql -r csv (no long-running sql-server). The scheduler shares one
event loop with the API (app/scheduler.py:73) — run dolt/unzip/download
subprocesses via asyncio.create_subprocess_exec (or an executor), never blocking
calls. Check free disk space before pulling; alert and skip if below threshold.
Deployment constraints: deploy is rsync --delete of the repo tree
(.gitea/workflows/deploy.yml:127), so clones and archives must live outside the
deployment path — an env-configured persistent directory (e.g.
DOLT_DATA_DIR=/var/lib/signal-platform/dolt). The dolt binary is a new prod
runtime dependency: install it once with the version-pinned provisioner in
deploy/provision_fundamentals.sh; operational steps are in
docs/fundamentals-deployment.md. The clone is reproducible from DoltHub; the
normalized PostgreSQL rows remain part of the normal database backup.
Validation gates (block promotion, raise an alert via the existing system-events path): source freshness as expected; tracked-universe coverage; no duplicate business keys; fundamental units/periods consistent; row-count deltas within reason; upstream schema change stops promotion. Workstream B adds: OHLC sanity (high ≥ open/close/low, low ≤ open/close/high, volume ≥ 0); no unexplained split discontinuities.
Split adjustment (workstream B): ohlcv_source_bars + corporate_actions are
the source of truth; canonical ohlcv_records is generated from them to match
Alpaca Adjustment.SPLIT, selecting only adjustment = raw rows as input so
adjustment is applied exactly once. A newly published split rewrites the symbol's
entire adjusted history — treat whole-symbol rewrites as a normal import event
(exempt that symbol from the row-count-delta gate for that run) and stamp rows with
the import_run_id (a backtest↔prod parity guard exists; changed history changes
backtests).
Scheduling (app/scheduler.py — SCHEDULE_DEFAULTS / _CRON_JOBS, ~line 1451)
Follow the existing pattern: cron strings in SystemSettings via
app/services/settings_store.py, day-of-week as names never numbers, logging via
_log_event.
Workstream A:
- Dolt earnings import: daily ~02:30 ET (with the future-row replacement above).
- SEC fundamentals job: daily ~04:00 ET. One job, three steps:
(a) detect a composite revision from the latest EDGAR daily-index date, the
exact tracked index rows, and the tracked-universe fingerprint — an unchanged
revision is a
no_opbefore Company Facts are fetched; (b) when changed, fetch and parse Company Facts only for tracked-universe CIKs that filed, plus full available history for the first run or a newly added issuer, through validation→atomic promotion. Universe resolution and the exact index inputs are cached during revision detection and reused during staging; CIKs are resolved fromcompany_tickers.jsonwithout writes until promotion; (c) always, locally, and only after production activation (the phase-A5 parity approval): refresh the legacyfundamental_datafields and mark affected cached fundamental scores stale. Sources differ per field — do not assume all five come from SEC:pe_ratioandmarket_capfrom the newest valid snapshots × latest PostgreSQL close, each with its own formula —pe_ratio= latest close / TTM diluted EPS;market_cap= issuer-wide shares outstanding × latest close;revenue_growthfrom the snapshots alone;earnings_surpriseandnext_earnings_datefromearnings_events(the Dolt earnings feed — these two do not exist in SEC facts). Before activation the job imports snapshots only (shadow). Step (c) must run identically when SEC is unreachable — prices move daily even when filings don't, and the earnings-derived fields already live in PostgreSQL. The new API valuation object is not stored anywhere — it is computed at request time (below). No valuation cache or table exists.
Workstream B:
- Dolt OHLCV+splits pull/import:
0 2 * * tue-satET. If source_max_date is not fresh, retry hourly until ~06:00, then give up quietly. After a successful import, reconcile the previous session's Dolt-derived bars against the Alpaca bars; summary into the run's validation JSON, details to logs. - Move
schedule_daily_pipeline_cron(morning refresh) from0 2 * * *to0 3 * * *(only needed once the 02:00 slot is taken by the OHLCV pull). Late Dolt publication is a non-event: the canonical scan runs at 15:30 on Alpaca, so the morning pipeline runs normally even when the import hasn't landed — no gating, no defensive coupling.
Metrics catalog (curated — TTM basis)
Snapshots store price-independent per-period facts (the "snapshot" column below
means derived from stored snapshots, assembled across periods at read time — see
Schema — not frozen at import); price-dependent ratios are never frozen into
snapshots and have no storage location at all: the API computes
them at request time from the stored snapshots + the latest ohlcv_records close
(both already in PostgreSQL, so this works identically when SEC is unreachable).
The only stored price-dependent values are the legacy fundamental_data fields
that scoring already reads, refreshed daily by step (c) after activation.
| Metric | Definition | Where computed |
|---|---|---|
| Revenue growth YoY | TTM revenue vs prior TTM | snapshot |
| EPS growth YoY | TTM diluted EPS vs prior TTM | snapshot |
| Operating margin + 4q trend | TTM operating income / revenue | snapshot |
| FCF margin | (TTM CFO − capex) / revenue | snapshot |
| Net debt | total debt − (cash + ST investments); positive = net debt | snapshot |
| Net debt / EBITDA | net debt / TTM EBITDA | snapshot |
| Share count Δ YoY | shares outstanding vs year ago | snapshot |
| Trailing P/E | price / TTM diluted EPS | request time |
| FCF yield | TTM FCF / est. market cap | request time |
| Est. market cap | issuer-wide shares outstanding × ticker price | request time |
| Earnings surprise history | last 4+ from earnings_events |
query |
Market cap is an estimate (issuer-wide shares outstanding × one ticker's price —
approximate for multi-class issuers). Share count comes from a single consolidated
value, not a class sum: prefer the one dei:EntityCommonStockSharesOutstanding
cover-page fact; if absent (e.g. Alphabet) fall back to
us-gaap:CommonStockSharesOutstanding at period end. companyfacts is
non-dimensional, so class-specific facts can't be summed reliably — never do that,
and never substitute weighted-average/diluted shares; if conflicting values remain,
store null. Label it "est." in the UI and round aggressively rather than withholding
it; false precision is the failure mode, not the approximation.
Units follow existing app conventions: percentages are percentage points (21.0 = 21%), P/E and net-debt/EBITDA are multiples, market cap and net debt are dollars.
Deliberately excluded: ROIC (invested-capital/NOPAT normalization too noisy), gross margin (COGS tagging too inconsistent), any new composite score.
Peer comparison (read-time only)
- Peer group = tracked-universe issuers sharing the first two SIC digits, deduplicated by CIK — GOOG and GOOGL are one issuer, one observation, in medians, percentiles and peer_count.
- Computed at read time from current snapshots — no aggregate tables until performance demonstrates a need.
- Medians exclude null/invalid values. Fewer than 5 valid peer issuers → omit the peer result entirely rather than showing a misleading universe comparison.
- Percentile direction respects metric polarity (higher-is-better for FCF yield, lower-is-better for P/E and leverage).
API contract (additive v1)
Every existing top-level field is preserved unchanged (name, type, position) — backend and frontend ship independently, no breaking interval. New objects, exact names and types:
{
// ...all existing legacy fields, unchanged...
"earnings": {
"next": {"date": "YYYY-MM-DD", "session": "bmo|amc|unknown", "days_until": 12} | null,
"recent": [ // newest first, max 4, may be empty
{"announce_date": "YYYY-MM-DD", "period_end": "YYYY-MM-DD|null",
"eps_estimate": 1.02|null, "eps_actual": 1.10|null, "surprise_pct": 7.8|null}
]
},
"metrics": [ // fixed row set — every key always present, value null when unavailable
{
"key": "revenue_growth_yoy", // revenue_growth_yoy | eps_growth_yoy | operating_margin |
// fcf_margin | net_debt | net_debt_to_ebitda | share_count_change_yoy
"value": 18.0, // number | null — pp / multiples / dollars per units above
"history": [ // oldest→newest, max 4 points, [] when unavailable
{"period_end": "YYYY-MM-DD", "value": 8.0}
],
"industry": { // object | null — null when < 5 valid peer issuers (CIK-deduped)
"label": "SIC 73 peers", // truthful 2-digit group label — grouping IS 2-digit,
// so no 4-digit description like "Prepackaged Software"
"median": 11.0,
"favorable_percentile": 82, // 0-100, polarity-aware (higher = more favorable)
"peer_count": 12 // issuers, not tickers
},
"period_end": "YYYY-MM-DD|null",
"filed_date": "YYYY-MM-DD|null",
"source": "sec|dolt|legacy_api"
}
],
"valuation": { // object | null (null until SEC snapshots exist, phase A3); same industry sub-object rules
// computed at REQUEST TIME from stored snapshots + latest PostgreSQL close —
// no valuation cache or table; unaffected by SEC availability
"pe": 29.2|null, "fcf_yield": 3.8|null,
"market_cap_est": 1.2e9|null, // estimated — UI labels "est."
"pe_industry": {...}|null, "fcf_yield_industry": {...}|null,
"price_date": "YYYY-MM-DD" // close used for the ratios
}
}
Null/freshness semantics: absent data is null with the row still present (the UI
shows "n/a", never hides rows); every metric carries its own source, period and
filing date — no panel-wide source label. The objects may serve partial data during
rollout (e.g. earnings live, metrics still legacy_api); the shape never
changes.
UI — frontend/src/components/ticker/FundamentalsPanel.tsx
One distinctive visual device — the Reference Rails — in an otherwise restrained panel. Preserve the app's dark glass styling and numeric typography.
Fundamentals
Growth accelerating · margins improving · valuation priced above peers
Next earnings Aug 3 · AMC EPS surprises ▂ ▅ ▃ ▆
Operating trend less favorable ← ref → more favorable
Revenue growth 18%
───────────────│━━━━● +3pp vs prior · accelerating
Share count YoY −1.7%
───────────────│━━━━● buying back
Valuation & balance less favorable ← median → more favorable
P/E 29.2×
────●━━━━━━━━━━│──── priced above peers · median 23.5× · 12 peers
- Growth and margins: horizontal rails compare the latest value with the prior quarter or prior-period average; share-count YoY compares with zero. The rail is normalized so right is always more favorable, including buybacks.
- P/E, FCF yield and leverage: horizontal favorable-percentile rails with a
peer-median marker. No decorative rail when
industryis null (< 5 peers). - Every row keeps the exact value and one deterministic comparison caption;
missing values render
n/a, and insufficient peers renderpeers n/a. - Earnings: four bars around a shared zero baseline — cyan beats, coral misses, gray unavailable — plus next date and BMO/AMC session countdown.
- Accessibility: color is always paired with text; neutral/ambiguous stays gray; rails and earnings bars expose complete ARIA descriptions.
- Remove the hard-coded "FMP" source label; surface filing and price-date provenance in the footer.
Deterministic reads — one shared rule set. Implement as a single function with named constants; the metric reads and the header sentence use identical outputs. No LLM, no new composite score. Defaults (tunable constants, not scattered literals):
- A series read requires ≥ 3 periods; otherwise show "—" and no read.
- Growth metrics (pp): latest − prior ≥ +2.0 → "accelerating"; ≤ −2.0 → "decelerating"; else "steady".
- Margins (latest vs mean of prior periods, pp): ≥ +1.0 → "improving"; ≤ −1.0 → "deteriorating"; else "stable" (phrased "above/below own average" where the layout calls for it).
- Share count YoY: > +1.0% → "N% dilution"; < −1.0% → "buying back"; else "flat".
- Peer-relative: favorable_percentile ≥ 60 → favorable ("above peers"); ≤ 40 → adverse ("priced above peers" for P/E, "elevated leverage" for net-debt/EBITDA); else "in line".
- Header sentence: join the growth read, margin read and peer-relative valuation read with " · ", omitting segments that have no read (e.g. "Growth accelerating · margins stable · valuation above industry median"). Segment sources are fixed: growth = revenue growth read; margins = operating margin read; valuation = P/E peer-relative read, falling back to FCF yield when P/E is null. This keeps the header unambiguous when sibling metrics (EPS vs revenue growth, P/E vs FCF yield) point in different directions.
Decommissioning (end of workstream A)
Remove completely: FMP (app/providers/fmp.py), Finnhub + Alpha Vantage
(app/providers/fundamentals_chain.py), their config keys (app/config.py), and
their wiring in app/scheduler.py, app/routers/ingestion.py,
app/services/ticker_universe_service.py. Retain: Alpaca (prices), FRED,
sentiment provider, Telegram. Note: decommissioning does not depend on
workstream B — Alpaca remains the price source throughout.
Rollout
Workstream A:
- A0. License review DONE (earnings approved for private/internal use under
CC BY-SA 4.0, no redistribution — see Licensing above). The Dolt version,
persistent
DOLT_DATA_DIR, clone, and production checks are captured indeploy/provision_fundamentals.shanddocs/fundamentals-deployment.md. - A1. Migration 026, import-run framework.
- A2. Earnings ingestion in shadow (writes
earnings_events, prod untouched); verify forward-calendar coverage and rescheduling behavior. - A3. SEC daily job in shadow (writes
fundamental_snapshots). Primary technical risk here: Q4 derivation and fiscal-period alignment — non-calendar fiscal years, restatements/amendments, and XBRL unit/dimension variants; budget accordingly. - A4. API v1 + FundamentalsPanel + peer comparison — served from snapshots and earnings_events, independent of the scoring cutover (the additive API supports partial data). UI value ships before anything touches scoring inputs.
- A5. Score-parity gate →
fundamental_datacutover: compute candidate pe_ratio/revenue_growth/earnings_surprise from SEC/Dolt side by side with the API values across the tracked universe, report per-field deltas and resulting fundamental-score/ranking changes, require explicit approval. Definition changes (e.g. TTM vs provider convention) called out, not averaged away. Status 2026-07-24: the gate has been exercised and the evidence supports approval — see the handoff section below. Step (c) is implemented behind the default-offfundamental_data_sec_dolt_cutover_enabledSystemSetting; the remaining production action is flipping that switch on and observing it. - A6. Remove FMP/Finnhub/Alpha Vantage; keep monitoring + manual fallback.
Workstream B (independent, start when wanted):
- B0. Stocks clone (~4.7 GB) provisioned; migration 027.
- B1. OHLCV + split adjustment in shadow (writes
ohlcv_source_barsonly; Alpaca keeps owningohlcv_records); historical backfill. - B2. Reconciliation window (≥ 2 weeks) vs Alpaca; review validation summaries.
- B3. Promote Dolt as historical OHLCV source (canonical rebuilt from raw source bars + splits); morning pipeline → 03:00.
Test plan
- Daily SEC job: changed revision imports; unchanged conditional-HTTP check is a
no_opwith zero downloads and zero row changes; validation failure leaves production untouched; source unavailable still runs the localfundamental_datarefresh (step c, post-activation); before activation the job never writesfundamental_data. - Valuation endpoint returns identical values with SEC reachable and unreachable (pure PostgreSQL computation); no valuation rows exist in any table.
- Multi-class tickers resolve to the same CIK snapshots; peer medians and peer_count are CIK-deduplicated (GOOG+GOOGL = one observation).
- Earnings rescheduling: a moved future date replaces the old row atomically; a cancelled date disappears; historical results are never touched.
- Amendment selection: for a period with multiple accessions, the newest valid accepted_at wins at read time; older rows remain unchanged.
- History arrays are chronological, ≤ 4 points.
- Percentage-point units stay compatible with existing formatters and scoring inputs.
- Deterministic reads: threshold boundary cases (exactly +2.0pp, exactly 60th percentile) resolve per the stated rules; header uses identical outputs and falls back from P/E to FCF yield for the valuation segment when P/E is null.
- Peer comparison disappears below 5 peer issuers; favorable-percentile direction correct for both polarities.
- Workstream B: split-adjusted OHLCV matches Alpaca on representative normal / split / reverse-split symbols.
- UI states: positive, adverse, neutral, insufficient history, insufficient peers; mobile layout; non-color accessibility.
- Unit, integration, scheduler and frontend suites pass.
Acceptance criteria
- App works normally with Dolt/DoltHub/SEC unreachable.
- Re-running the same revision: zero duplicate or changed rows.
- Upcoming earnings dates present and timely for the tracked universe — the forward calendar is the hardest thing to replace and gates decommissioning.
- Coverage meets the tracked-universe target; scheduler runs cleanly with FMP/Finnhub/AV keys removed from the environment.
- Score-parity diff reviewed and approved before
fundamental_datacutover. - Scheduled imports never block the API event loop.
Handoff — remaining work after the A5 parity investigation (2026-07-24)
The 2026-07-23 parity report surfaced coverage gaps and wrong values; a nine-pass
investigation traced every one to parser/identity bugs (not source data), fixed them,
and reparsed production twice. Full evidence trail:
reports/fundamentals-parity-20260723-findings.md (root causes, decisions, validation)
plus the before/after reports (fundamentals-parity-20260723T… / …20260724T….json).
Post-fix: candidate scores 504 of 511 vs legacy's 507 (gap = PSKY/Q new registrants +
FITB, all explained); revenue-growth agreement 0.0038 median abs delta where both exist.
Dennis reviewed the evidence 2026-07-24 and directed proceeding to cutover.
Task 1 — A5 activation (IMPLEMENTED 2026-07-24; production switch remains). The
post-activation local refresh of fundamental_data derives pe_ratio and
market_cap from newest valid snapshots × latest PostgreSQL close, revenue_growth
from snapshots, earnings_surprise/next_earnings_date from earnings_events; mark
affected cached fundamental scores stale; must run identically when SEC is unreachable.
It consumes fundamentals_derivation.derive() outputs, NOT raw snapshot fields —
that path carries the split guard (ttm_diluted_eps
nulls when contaminated, with ttm_diluted_eps_caveat) and the multi-class share
fallback (shares_outstanding + shares_outstanding_estimated). Parity and activation
share the same candidate builder. Activation is the explicit
fundamental_data_sec_dolt_cutover_enabled SystemSetting and defaults off; see
docs/fundamentals-deployment.md for the production flip and rollback procedure.
Task 2 — A6 decommissioning. After a short observation window: remove
FMP/Finnhub/Alpha Vantage providers, config and env keys; keep monitoring + manual
fallback. Gated by the acceptance criteria above — especially forward-calendar
timeliness from dolt_earnings (its source_max_date ran ~5 weeks ahead as of
2026-07-23, which passes).
Known caveats to carry (documented in the findings report, not bugs to fix):
- KLAC-class post-filing splits: P/E wrong until the next 10-Q; undetectable from
snapshots. Workstream B's
corporate_actionstable is the natural future fix. - BRK-B: no share count exists anywhere in companyfacts → no market cap, correctly.
- FITB: unscored (split guard + no taggable revenue) — the one name that lost its score relative to legacy; composite renormalises.
- Share-change guard at 25% nulls P/E for stock-funded M&A too (COF, WAT…); revisit only if the ~3% universe hit-rate proves painful.
sec_cik_overridesSystemSetting pins XOM → 34088 (applied in prod); theno_xbrl_filingsSystemEvent says when a new pin is needed.- After any future parser change, stored rows need
scripts/reparse_fundamentals.py(dry-run default;--applyrewrites) — snapshots are otherwise immutable.
Deferred (explicitly, until a concrete need appears)
- Workstream B itself is deferred relative to A and blocks nothing in A.
- Exact byte-level source replay of historical imports; permanent archive store.
- Point-in-time backtest enforcement (
accepted_atis stored now; derivation and backtest visibility rules are built only when fundamentals enter scoring/backtesting). - Fundamental metrics in the score; sector-relative scoring.
- Aggregate/rollup tables for peer statistics.
- Any valuation cache or table (request-time computation from snapshots + latest close suffices).
- A dedicated conflicts table (validation JSON + logs suffice).