A conditional clause in the summary sentence carried an f prefix with no
placeholders, failing `ruff check app/` and blocking the deploy. Literal only;
the rendered text is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The v3 cutover run scored 2/4 corrections warned against v2's 3/4, which reads
like a regression and is not one. Only 4 of the 11 detected corrections fall in
the holdout, so recall is one event from a different headline -- and the event
that flips is decided by threshold placement, not by what the score saw. "v3
without the credit sensor" catches 2025-02-21 at a *higher* threshold (35.5)
than shipped v3 misses it at (32.3), because the alarm rule needs a rising edge
and a lower threshold can fire outside the horizon then never reset below.
Two caveats are now computed and surfaced rather than left for the reader to
infer:
- Holdout event count against MIN_EVENTS_FOR_CONFIDENCE. The summary sentence
states how many of the detected corrections actually fall in the test period.
- Warning-sensor coverage across the split. The score renormalises over what is
available, so a training window predating a sensor's history freezes the
threshold on a different construct than the holdout is measured against. At
the cutover that is 39% of training sessions with all three sensors versus
100% of the test period, credit history beginning 2023-07-25.
Restricting the threshold to sensor-matched training sessions was tested and
rejected: those sessions are a calm recent stretch, so the threshold falls from
32.3 to 22.5 and false alarms rise from 3.3 to 8.6/yr. It swaps a coverage bias
for a regime-selection bias. The report states its limits instead.
_warning_series now returns per-session sensor counts alongside the scores.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The LLM-sourced capex/earnings observations carried 12+8 of 100 Warning points,
so both pegged at 100 produced a Warning of 20.0 -- below the event study's 25.3
alarm threshold and still inside the "stable" band. The reading was
arithmetically incapable of changing anything on screen, which is why refreshing
it appeared to do nothing. They are now a qualitative overlay reported beside
the scores rather than diluted into them.
Calibrated against the 408 v2 sessions to 2026-07-24, reproduced offline from
Alpaca + FRED; the harness matched the stored prod distribution exactly before
any parameter was changed.
State:
- P3 used dd_pct * 5, reaching 100 at a 20% drawdown -- the 90th percentile of
the observed distribution -- so 39/408 sessions sat at exactly 100 with no
resolution left during the part of a selloff that matters most. Replaced with
anchored breakpoints keeping headroom past the observed 36% maximum, blended
2:1 like P1/P2 instead of max(). P3's realized share of State falls from 65%
to 40%, matching its nominal weight.
- Credit level is now anchors-only. ICE capped FRED's BAMLH0A0HYM2 at a rolling
3-year window in April 2026, silently turning the 10-year percentile leg into
a 3-year one that scored 20 points of stress at an OAS of 3.5 -- the level its
own anchors call "mild". The anchors already encode the long-run distribution.
Warning:
- Added HY OAS 20-session widening (25%). The level is pinned at zero below the
3.5 anchor; its rate of change is not.
- Divergence tapers to a 0.35 floor instead of a hard price_ret >= 0 gate, which
zeroed the sensor through every decline: on 2026-07-24 the basket shed 10
points of participation in 20 sessions and Warning printed exactly 0.
- The event study and the live monitor now share one sensor definition, so they
cannot silently drift apart.
Bands are per axis (State 20/50/80, Warning 20/40/60) with quadrant dividers at
50/40; v2 Warning never exceeded 64.9 against a shared 60, leaving that half of
the quadrant unreachable. Realized shares: State 73/15/8/3%, Warning 69/20/8/3%.
Snapshots now record credit_history_days and vix_history_days -- the percentile
defect went unnoticed for months because nothing asserted the window the code
claimed.
Cutover: the first run rebuilds 400 sessions automatically; the Event Study job
must be re-run, as its cached report self-invalidates on the methodology check.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Regression found by the post-reparse collision check. _period_identity trusted
submissions.fiscalYearEnd, which is not reliable: Franklin Resources (BEN)
declares 1231 while every one of its 10-Ks ends 09-30.
The effect was data loss, not just a bad label. BEN's real fiscal Q1 (Dec 31)
sat 0 days from the claimed year end, matching no quarter band, so it fell back
to SEC's fy/fp; its fiscal Q2 (Mar 31) computed 275 days out and was labelled
Q1. Both landed on the same key, the collision discarded one, and BEN lost TTM
EPS and revenue growth entirely — values it had before this branch.
A 10-K's reportDate IS the fiscal year end by definition, so resolve_fiscal_
year_end() now prefers the issuer's most recent annual filing and treats the
declared value as a fallback for issuers with no 10-K in the set.
Scanned the full tracked universe: 2 of 506 issuers declare a year end more
than 21 days from their own 10-K — BEN (91d, broken) and DELL (29d, mislabelled
but functionally correct). Both now derive correctly and match the legacy
provider: BEN revenue growth 3.8243 vs 3.82, DELL 38.5735 vs 38.57. Controls
(AAPL, COST, PEP, DPZ, IRM, JPM, CRM, STX, AVY) byte-identical.
826 unit tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Found in review. _merge_amendments rebuilds a period from _MERGED_FIELDS +
_CARRIED_FIELDS alone, so a column in neither list is absent from the merged
row, not just stale — and callers read it with getattr(..., None), which
silently yields None. weighted_avg_diluted_shares was never added when the
market-cap fallback landed (_SNAPSHOT_COLS in the importer was updated, its
counterpart in the derivation was not).
The failure needed both of this branch's fixes at once: a multi-class issuer
with a partial amendment on its latest period (META with a Part-III-only
10-K/A) would silently lose market cap and FCF yield again.
Adds the field, a regression test for that case, and a guard test asserting
the merge/carry lists cover every SnapshotRow field, so the next column added
fails loudly rather than losing data quietly. Confirmed the guard catches the
original bug.
Also from review:
- Expose pe_caveat in the valuation payload, so a P/E suppressed by split
contamination says why instead of looking like missing data (the caveat was
set but never read).
- no_xbrl_filings now names both causes; the old text advised pinning a CIK
override, which is wrong for a genuine new registrant that simply has not
filed yet and clears itself.
- Document that fiscalYearEnd is the issuer's current calendar, so a fiscal-
year-end change degrades old periods (fallback/newest-wins), not current ones.
- Parser-level tests for _select_weighted_avg_shares (shortest-span-wins and
concept priority), which only had derivation-level coverage.
823 unit tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Operational plumbing to land the parser fixes and to make silent resolution
failures visible.
- Reparse: SecFundamentalsImporter(reparse=True) restages every accession with
the current parser and rewrites the ones that now reconstruct differently,
writing the full column set so a row is never half old-parse. Snapshots stay
immutable with respect to SEC; the stored row is our reconstruction, and after
a parser fix keeping it is a stale cache, not history. run_import(force=True)
bypasses the unchanged-revision no-op, since the staleness is on our side, not
the source's. Exposed as scripts/reparse_fundamentals.py, dry-run by default.
- CIK overrides: sec_universe reads a {symbol: cik} pin from
SystemSetting['sec_cik_overrides'], applied ahead of company_tickers.json, for
when SEC maps a ticker to a successor shell with no filings (XOM -> a zero-
filing "ExxonMobil Holdings Corp" while every 10-K/Q is under CIK 34088).
- Resolution validation: a tracked issuer resolving to a registrant with no XBRL
filings now records no_xbrl_filings and raises a warning naming the CIKs and
the override setting, instead of silently yielding nothing on every run.
- _diff_fields compares datetime instants, not representations: accepted_at
round-trips naive from SQLite but tz-aware from Postgres, which otherwise made
a reparse of identical data report every row as changed (and false-positived
the pre-existing discrepancy warning).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The A5 parity report surfaced coverage gaps and wrong values that all traced
to the SEC facts parser and read-time derivation rather than to bad source
data. Fixes, each validated by replaying the production parser + derivation
against live company facts:
- Period identity is derived from period_end against the issuer's fiscal
calendar, not SEC's fy/fp fields, which collide (two period ends on one key,
one silently discarded) and invert (a period sorting before one that precedes
it) often enough to break the quarter chain. Recovers BXP, CRM, CRWD, FRT,
MTD, NTAP, PPL, STX, WDAY. Fixed labels are internal ordering keys only (not
in any API schema), so a filer whose year ends in early January shifting by
one is harmless.
- Revenue concept list gains RevenuesNetOfInterestExpense (banks) and the
IncludingAssessedTax variant (REITs/consumer); EPS gains the continuing-ops
variant (REG/FCX) and, last, basic EPS for a period tagging no diluted
variant at all (PPL). All appended, so any issuer that already resolved keeps
its concept.
- YTD span tolerance 20 -> 25 days, covering 4-4-5 retail calendars whose
36-week YTD-Q3 (251-252d) previously missed by ~2 (COST, PEP, DPZ).
- Amendment resolution is per field: a partial 10-K/A (Part III only, no
financial facts) no longer blanks the period (DVN).
- TTM diluted EPS is suppressed when a split contaminates the trailing window
(BKNG's mixed-unit sum produced a P/E of 1.10 that clamped to a perfect
fundamental sub-score). A post-filing split with no share-count evidence
(KLAC) remains undetectable from this data.
- Multi-class share fallback: weighted_avg_diluted_shares is captured and used
for market cap when the cover-page count is absent (dimensional, so missing
from company facts for META/CMCSA/CHTR/FOXA/NWSA/LEN). Within ~0.6% of the
true count on controls; flagged shares_estimated in the API. BRK-B has no
weighted-average fact either and stays unavailable.
820 unit tests pass; new tests confirmed to fail against the pre-fix code.
Effect is inert until existing rows are reparsed (see reparse path).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- The eps_growth_yoy read was never computed, leaving that fixed by_key entry
null even with sufficient EPS history; now growth_read() is applied to EPS
history just like revenue.
- Tests: same-day earnings returns as next with days_until 0 (and not in
recent); zero close guards valuation to null; eps read populated. Fixture
seeds three fiscal years so YoY growth reads have a >=3 run. 9 API tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1. Multi-class subject is priced by the REQUESTED ticker: the peer group's
representative for the subject CIK is overridden to the requested ticker_id
(other issuers pick a deterministic-by-symbol rep), so GOOGL's P/E uses
GOOGL's price, not GOOG's. Differing-price GOOG/GOOGL test added.
2. reads matches the selected contract: header is null when there is no read;
by_key is a fixed map over every metric key plus pe and fcf_yield, null when
unavailable (was a sparse dict).
3. Earnings use the New York calendar date; same-day is UPCOMING (days_until 0),
recent is strictly earlier.
4. Valuation is null when there is no usable price (> 0 required for P/E and
market cap); when present, price_date is non-null.
Added a real router/API-envelope test with a seeded legacy record (the endpoint,
not just the schema merge). 6 tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GET /fundamentals/{symbol} now returns the additive v1 objects alongside the
unchanged legacy fields (no legacy growth mapped onto the SEC TTM metric).
- earnings: next (date/session/days_until) + recent (<=4, with surprise_pct)
from earnings_events.
- metrics: fixed key set (value + dated history + per-metric SIC-peer industry
object + source=sec); net_debt has no industry (size-dependent).
- valuation: P/E, FCF yield, market_cap_est computed at REQUEST TIME from the
derived TTM inputs x the latest ohlcv close (no stored valuation); guarded to
null on missing/invalid inputs; pe_industry / fcf_yield_industry peer stats.
- reads: deterministic outputs in a SEPARATE object (header + per-metric reads).
Peer queries are batched and CIK-deduplicated by 2-digit SIC; industry omitted
below 5 valid peers. Schema extended with optional typed sub-models; the router
merges legacy + v1 so every existing field is preserved.
Tests: 4 (full assembly incl. peer industry + valuation + additive-merge, no-cik
null metrics, <5-peers omitted, price-guarded valuation).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1. Peer percentile is now a tie-aware rank against the OTHER issuers
((worse + 0.5*tied)/(peers-1)): an all-equal group maps to 50 (not 100), the
median maps to 50, a unique best to 100, a unique worst to 0.
2. Deterministic reads use the consecutive non-null suffix ending at the latest
point (>=3 values): a null latest or an internal gap yields no read, so a read
never reflects a period displayed as n/a.
3. Peer filtering excludes non-finite (NaN/±inf) as well as null, including an
invalid subject.
Tests updated + added (all-equal, median rank, non-finite, latest-null history,
internal gap). 15 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Completes the pure read-time core.
fundamentals_peers.py: median + polarity-aware favorable percentile + peer_count
for a subject within its SIC group (CIK-deduped by the caller); returns None
below MIN_PEERS=5 so the caller omits the industry object. Absolute net_debt is
intentionally NOT peer-eligible (size-dependent) — leverage compares via
net_debt_to_ebitda. HIGHER_IS_BETTER polarity map + two_digit_sic() grouping key.
fundamentals_reads.py: one shared deterministic rule set (no LLM): growth_read
(+-2pp), margin_read (latest vs mean-of-prior, +-1pp), share_count_read (+-1%),
peer_read (60/40 bands, polarity-aware phrasing per metric), header_sentence
(growth · margins · valuation, omitting empty). Tunable named constants; >=3
periods required for a series read.
Tests: 8 peer + 5 reads, anchored on the boundary cases (exactly +2.0pp, exactly
60th percentile, exactly +1.0pp margin). 13 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1. net_debt requires BOTH cash and total_debt; a missing side is null, not
treated as zero (which would be a partial, misleading value).
2. net_debt_to_ebitda is null when TTM EBITDA <= 0 — a negative denominator
would otherwise rank a distressed issuer as favorably low-leverage.
3. The quarter tape is the CONSECUTIVE run ending at the latest period (stops at
a gap), so trend text never compares non-adjacent quarters as if consecutive.
4. YoY growth is null when the prior-year TTM is <= 0 (e.g. loss->profit), which
is not a meaningful percentage.
Also corrected the plan's net-debt formula to total debt − (cash + ST) matching
the positive-means-net-debt implementation. +4 tests. 10 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Derives the display metrics from the stored YTD snapshots at read time (no I/O,
no DB), per the A3 schema decision. Given an issuer's snapshot rows it produces:
- amendment selection (newest accepted_at per fiscal period);
- discrete quarters = YTD(Qn) - YTD(Qn-1), Q4 = YTD(FY) - YTD(Q3);
- TTM = trailing four discrete quarters; missing period -> null, never partial;
- metric series (value + 4-quarter tape, each point dated): revenue_growth_yoy,
eps_growth_yoy, operating_margin, fcf_margin, net_debt, net_debt_to_ebitda,
share_count_change_yoy;
- request-time valuation inputs (ttm_diluted_eps, ttm_fcf, shares_outstanding)
for the API to combine with price.
Units per app convention (percentages = pp, leverage = multiple, dollars).
Tests: 6 (growth+Q4, margins, net-debt/EBITDA+dilution, valuation inputs,
missing-period-null, amendment selection). Verified on real Apple snapshots:
op margin 32.6%, net-debt/EBITDA 0.10, buyback -1.7%/yr, TTM EPS $8.26.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Extend the companyfacts structural check to reject a concept with a
missing/non-dict `units` mapping (not just the top-level `facts`), so a
partially-malformed payload fails promotion instead of silently dropping that
concept's facts. New fixture proves it fails.
- Strengthen the newly-added-issuer test: keep latest_index equal to the prior
run so ONLY the universe fingerprint changes the revision — proving the
fingerprint alone prevents a new ticker from being starved/no_op'd.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1. Removed the 45-day index-walk cap: it discarded the older part of a long
outage while still advancing source_max_date, permanently losing filings.
The walk now covers every unprocessed date (a large gap is one-time cost).
2. Discrepancy detection meets the immutability contract: it compares ALL source
snapshot fields (not five), read-only during stage/validate, reports the
differing accessions + fields in validation_json, and promote emits a warning
system event (in-transaction) — never mutating the stored row.
3. Malformed companyfacts (missing facts/units structure) are recorded separately
and FAIL validation, instead of silently degrading to skipped rows that the
50% backfill coverage floor could still pass.
Also corrected the stale "sum share classes" / DEI-only wording in the snapshot
model docstring and the A3 design doc to describe the us-gaap fallback.
Tests: +4 regressions (>45-day gap loses nothing, newly-added issuer backfills
without filing, malformed payload fails, shares discrepancy detected + evented).
23 passed, 1 skipped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
SecFundamentalsImporter (SourceImporter, source=sec_facts): populates immutable
fundamental_snapshots from Company Facts and back-fills tickers.cik/sic, driven
by the EDGAR daily index. Shadow only. Guardrails per review:
- detect_revision caches the resolved universe + exact tracked index rows and
composes the revision from them; stage consumes those same cached inputs
(no index/universe refetch) so promoted data matches the computed revision.
- Resolution is read-only in stage (proposals only); ticker writes happen in
promote via apply_ticker_updates.
- validate runs the index<->Company-Facts consistency gate before any write:
a tracked XBRL index accession missing from Company Facts fails the run
(they lag independently) so we retry, not record null. Non-XBRL amendments
are skipped with a recorded reason. Backfill has a coverage floor.
- promote inserts ON CONFLICT (accession) DO NOTHING (immutable), reports
differing existing accessions without mutating, and applies ticker updates in
the same transaction.
- Full-history backfill on first run / for newly-added issuers (include_history);
incremental fetch only for issuers that filed.
Parser: split parse result into skipped_filings vs field_issues (coverage must
not count field warnings); header notes the us-gaap shares fallback; added
companyfacts_accessions() for the gate.
Verified live end-to-end (AAPL + GOOGL backfill): 112 snapshots, cik/sic set,
GOOGL shares via us-gaap fallback, AAPL via dei. Tests: 6 importer (backfill,
incremental, consistency-gate fail, non-XBRL skip, read-only-on-failure,
conflict-discrepancy) + parser ParseResult updates.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1. Multi-class shares: prefer the single dei:EntityCommonStockSharesOutstanding
cover-page fact; else fall back to us-gaap:CommonStockSharesOutstanding at
period end (Alphabet has no dei fact). Never sum class facts (companyfacts is
non-dimensional) and never use weighted-average/diluted; conflicting values ->
null, counted as an "ambiguous shares outstanding" note in validation. Plan's
"sum class-specific" wording corrected. Verified live: Alphabet shares now
populate (12.1B), Apple still uses its dei cover date.
2. Fiscal context is the majority (fy, fp) among facts ending at reportDate, with
ties rejected — no longer the arbitrary first fact.
3. Hardening: catalog selectors require taxonomy == "us-gaap"; indexing drops
malformed facts (missing accession/end, non-finite value) so a custom concept
or bad date can't be selected.
Tests: +8 (dei precedence, us-gaap fallback, conflict->null, no weighted-average,
tie-context skip, foreign-taxonomy/malformed ignored, ambiguous-shares note).
14 passed, 1 skipped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pure parser (no I/O/DB) turning one issuer's companyfacts + submissions filing
metadata into per-accession snapshot rows for the filing's primary period.
- Period identity from end == reportDate, never fy/fp (fy/fp is the filing's
context; comparatives repeat it).
- Duration facts stored as cumulative YTD: pick the fact whose span matches the
fiscal-period-to-date length (Q1~3mo..FY~12mo) within tolerance; no YTD-length
fact -> null (never a discrete masquerading as YTD).
- Balance-sheet instants at end == reportDate; shares_outstanding is the dei
cover-page fact whose own end (cover date) is stored in shares_outstanding_date.
- Cash and debt composites are aggregate-first and mutually exclusive (each
source tag counted at most once).
- Carries filing_date through submissions rows (snapshot.filed_date).
Verified on REAL Apple companyfacts: 44 snapshots, 0 skipped, YTD revenue
124.3B->219.7B->313.7B->416.2B across FY2025 (Q4 derives at read time), every
shares_date is the cover date != period_end. Tests: 6 fixture + 1 skip-guarded
live-invariants (monotonic YTD, cover-date shares).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses the slice-1 review:
1. Resolution is now read-only (A1 transaction contract). resolve_ciks /
fetch_sic_updates compute proposals and mutate nothing; a new
apply_ticker_updates issues the writes, called only in promote — so a failed
validation can't leak ticker changes on the framework's failure commit.
2. Only 404 means "missing". Added SecNotFoundError; daily_index /
latest_index_date catch only that. 403, exhausted 429, 5xx, timeouts, and
transport/parse errors now propagate instead of looking like "no index".
3. Fair-access enforced when opening a REAL client (transport=None): reject
blank/placeholder/non-email User-Agent and sub-0.11s spacing. Mock transports
skip it (tests use 0 spacing).
4. submissions(include_history=False) by default — only the one-time full
backfill fetches the history shards; SIC/incremental work makes no extra
requests.
Plus: retry transient 5xx/network errors and honor Retry-After during the 1 GB
backfill; compose_revision rejects a missing index date (no "None:..." revision).
Re-verified live vs real SEC (fair-access validation passes, shard merge intact).
Tests: 18 (added error propagation, read-only resolution, fair-access, recent-only
submissions, reject-None revision).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
First A3 implementation checkpoint (design: docs/dolt-sec-a3-design.md).
- sec_client.py: async SEC EDGAR client honoring fair-access — identifying
User-Agent (config), request spacing < 10 req/s, exponential backoff on 429,
and 403 -> SecForbiddenError (alert and stop, never retry-loop). Fetchers:
company_tickers (normalised, multi-class share CIK), submissions (merges the
paginated filings.files shards so full history is visible), companyfacts,
daily_index (fixed-width form.idx parse), latest_index_date.
- sec_universe.py: resolve_ciks (tickers.cik backfill), refresh_sic
(sic/sic_description), and the composite-revision pieces — universe_fingerprint
(a new ticker changes the revision, so it's never no_op'd/starved),
index_content_hash, compose_revision.
- config + .env.example: SEC_USER_AGENT (must be a real contact email) + spacing
/ retries / timeout.
Verified live against real SEC: AAPL->320193, GOOG==GOOGL, BRK-B resolved;
submissions shard-merge proven (131 filings back to 1993); daily index parsed.
Tests: 11 (mocked-transport parsing + 403/429 handling + resolution/fingerprint).
Full suite 713 passed. Next slice: companyfacts -> snapshot parser + importer.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses the A2 review:
1. Every dolt subprocess is now bounded by a hard timeout
(dolt_command_timeout_seconds, default 600s); on expiry the process is killed
and DoltError raised — a hung pull/sql can no longer pin the import
connection and advisory lock indefinitely. Tested (timeout + non-zero exit).
2. Initial-load validate is stronger: besides zero-future, an initial load now
requires a real forward horizon (>= 21d, under the ~35d observed on the
clone) AND universe coverage >= 50% (a broken symbol join can't seed a hollow
calendar). Subsequent runs keep the 50% collapse gate.
3. Revision uses DOLT_HASHOF('HEAD') — formally HEAD, not dolt_log-by-timestamp.
4. Free-disk floor raised 2 GB -> 5 GB (safe headroom over the ~1.7 GB clone).
Full suite 702 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A SourceImporter that ingests post-no-preference/earnings into earnings_events
for the tracked universe. Shadow by construction (nothing reads earnings_events
until A4).
- earnings_alignment.py: pure calendar<->EPS-history min-cost monotonic DP,
reused from scripts/import_dolthub_earnings.py with identical constants (not
extending that one-off script); symbol/session normalization; unit-tested
against the pinned constants.
- dolt_client.py: async dolt CLI wrapper (pull / current_commit / query_csv via
asyncio.create_subprocess_exec — never blocks the shared event loop) + disk
guard before pull.
- dolt_earnings_importer.py: detect_revision = pull + HEAD hash; stage = query
earnings_calendar + eps_history, dedup, align, map act_symbol->ticker_id
(normalize both sides so dotted BRK.B joins); promote is destructive
(delete future dolt_earnings rows + upsert; past never deleted) so validate is
FAIL-CLOSED — blocks when the staged forward calendar is empty or has collapsed
below 50% of what's loaded (the forward calendar is the acceptance gate).
- NOTICE: CC BY-SA 4.0 attribution; config: DOLT_BINARY / DOLT_DATA_DIR / etc.
Verified end-to-end against the real 1.68 GB clone (5 tickers: 133 events, 128
paired, forward calendar to 2026-08-26, BRK.B joined). Tests: 9 alignment + 7
importer + 1 skip-guarded real-clone smoke. Full suite 699 passed.
Remaining for A2: wire the daily ~02:30 ET shadow cron — deferred to pair with
the deploy-time dolt install + DOLT_DATA_DIR provisioning.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Both real target tables (fundamental_snapshots, earnings_events) carry an
import_run_id; stamping requires the current run's id. A2 (the earnings
importer) is the first real consumer, so promote gains a run_id argument rather
than having importers hack the running row out of the framework. Protocol +
call site updated; the A1 fake importer now stamps and asserts import_run_id.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review fixes to the import-run framework (A1):
1. detect_revision ran outside the failure handler, so a failed revision probe
(the most likely external failure) escaped unrecorded — violating "every
attempt is recorded". Now the running row is created FIRST, then
detect_revision + last-revision lookup + stage + validate + promote all run
inside the same handler; the row converts to no_op when the revision is
unchanged. New test covers a detection exception → recorded failed + alert.
2. asyncio.CancelledError (BaseException, not caught by except Exception) left a
permanent running row on deploy/scheduler shutdown. Now caught explicitly:
best-effort mark failed, then re-raise the cancellation (never swallowed).
New test asserts the run is failed and the error re-propagates.
Full suite 682 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
run_import + SourceImporter Protocol (detect_revision/stage/validate/promote)
giving every bulk importer the plan's non-negotiables, KISS:
- one run per source at a time — Postgres session-level advisory lock held on a
single pinned engine.connect() so it survives the running-row and promotion
commits; no-op on SQLite.
- idempotent per revision — cheap detect_revision compared to the last promoted
run; unchanged revision records a no_op with zero writes (no fetch).
- staging (in-memory, no physical staging tables) → validate (read-only) →
atomic promote + run-row flip in one transaction.
- failed validation or mid-run exception marks the run failed, alerts via
system_event_service, and leaves live tables untouched.
Every attempt recorded in data_import_runs; conflicts summary in validation_json
(no conflicts table). Concrete SEC/earnings importers land in later phases.
Tests: 6 orchestration tests (no_op / promote / new-revision / failed-untouched
/ promote-exception-rollback) + deterministic advisory-key derivation. Full
suite 680 passed. Advisory-lock mutual exclusion is PG-verify-pending (SQLite
no-ops it — flagged, not covered).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The run-id marker proved which scan wrote last, but the shadow book still
selected setups by detected_at >= scan_start. An overlapping manual scan
could insert rows in that same window; if the pipeline's scan wrote the
marker last its id matched and the shadow book proceeded, then swept in --
or ranked highest -- a manual-scan row. The identity check gated entry but
selection did not.
Carry the run id onto the rows. Migration 025 adds an indexed
trade_setups.scan_run_id. scan_all_tickers computes one id per run
(pipeline's when a step, else fresh), passes it to scan_ticker which stamps
every row after enhancement, and writes the same id to the completion
marker. The shadow book selects WHERE scan_run_id == the matched id, so a
concurrent scan's rows are excluded by identity regardless of their
detected_at. The now-unused STARTED marker is dropped; COMPLETED
(freshness) and RUN_ID (identity) remain.
Decisive test: the pipeline's id matches, but a same-window manual row with
a higher rank is present and is excluded -- only the pipeline's own row is
traded. A time-window select would have swept it in and ranked it first.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A manually triggered rr_scanner and the scheduled near-close pipeline are
separate APScheduler jobs; max_instances=1 serialises a job only against
itself, so they can overlap. A manual scan starting just before the
pipeline can finish just after it began and overwrite the scan markers.
Its completion timestamp is then later than the pipeline start, so the
previous 'completed >= pipeline_start' check accepted its batch as though
it were the pipeline's own -- exactly when the pipeline's scan may have
failed.
Replace the timestamp comparison with an exact run-id match. A new
pipeline_run module holds a per-task run-id contextvar (separate module so
the scanner and scheduler import it without a cycle). _run_pipeline binds a
fresh id per invocation; scan_all_tickers stamps that id -- or a fresh one
when run standalone -- into the scan markers, written with started/completed
in a single commit. The shadow step requires the stored run id to equal its
pipeline's id exactly, so a concurrent manual scan (its own id) or a failed
pipeline scan (a prior run's id) can never be mistaken for it. Direct Admin
triggers have no pipeline context and keep the freshness fallback.
Known residual: the id match governs whether shadow proceeds; setup
selection remains detected_at >= scan start, so a fully per-run setup
isolation would need a run_id column on trade_setups (not required here).
Tests cover the reported race (manual scan finishing last is refused), a
failed pipeline scan, the id-match accept path, and contextvar propagation
and non-leakage across tasks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A 6-hour freshness window proves only that some scan ran recently, which a
manual mid-day scan satisfies. Scenario: a manual scan succeeds at 13:00;
the 15:30 near-close pipeline's scan step is disabled or fails; at 15:30
the 13:00 completion is still 'fresh', so the shadow step trades that
earlier batch despite no successful scan in the current pipeline.
_run_pipeline now records its start in a per-task contextvar, visible to
the steps it awaits. run_shadow_book reads it and requires the scan
completion marker to be at/after the pipeline start, so a scan that failed
or was disabled in this pass (marker left at a prior run, before the
pipeline began) cannot be substituted by an earlier manual scan. A direct
Admin trigger has no pipeline context and falls back to the freshness
window -- an explicit operator action, not an automated one.
Tests pin the reported case: a fresh manual scan predating the pipeline
start is refused; the pipeline's own post-start scan is accepted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Second review round on the shadow book; all three findings were real.
- Scan freshness is now proven, not assumed. Pipeline steps run and fail
independently, so a disabled or failed scan step still let the shadow
step run on the newest *stored* setups -- a prior session's picks at
stale prices. scan_all_tickers now records a run boundary
(last_scan_run_started_at / _completed_at) only on successful
completion; the shadow book refuses to trade unless COMPLETED is fresh
and selects only setups with detected_at >= the run start. Deduplication
to the latest row per ticker now happens BEFORE qualification, so a newer
unqualified row suppresses an older qualified one rather than the reverse.
- Shadow selection is hard long-only. setup_qualifies only enforces
long-only when min_momentum_percentile > 0, but 0 is a legal admin
setting, and the cash accounting assumes long positions -- so the
constraint is enforced in shadow selection regardless of gate config.
- The personal setup list excludes only the caller's own open positions.
get_trade_setups gained exclude_open_trade_user_id; the trades route
passes the authenticated user, while the Telegram broadcast stays global
since it has no single owner.
New tests cover stale/absent scan markers, prior-run exclusion, newer
unqualified suppressing older qualified, long-only under a disabled gate,
and both sides of the user-scoped exclusion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review of the shadow book found seven ways the two books could leak into
each other; all are fixed here. The most serious silently invalidated the
comparison the shadow book exists to make.
- Shadow holdings no longer suppress the manual candidate list. The
open-trade exclusion filtered on any book, so shadow taking the
top-ranked names removed exactly those from the user's list and alerts,
confining the discretionary book to leftovers. Scoped to the manual
book. Closed-trade alerts and paper-book equity were leaking the same
way and are likewise scoped.
- Shadow sizing now matches _simulate_portfolio: min(1% risk, 20% notional
cap, available cash) from marked equity, plus the sub- dust guard.
Previously risk-only from realized equity, so a tight stop produced a
multiples-of-equity leveraged position the strategy would never take.
- Shadow only trades setups from the scan that just ran (<6h old) with one
setup per ticker. A failed or disabled scan step could otherwise open
positions from a prior session at stale prices.
- Gate-reset transitions are observed for both books, so a shadow stop-out
completes fail -> requalify instead of staying locked forever.
- Manual list/close endpoints default to the manual book and reject
hand-closing shadow trades; the performance endpoint is scoped to the
caller so 'your picks' is not every user's book.
- run_shadow_book is registered as a paused job so Admin can trigger it.
Also anchors three pre-existing paper-trade tests (and the new alpaca
window test) on the UTC date. They build fixtures from the local date but
the service stamps opened_at in UTC, so they failed only between 00:00 and
02:00 in a UTC+hh timezone -- latent on ba2df8b, exposed by the clock.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The manual paper book only contains trades taken by hand, inside a 20
minute window, on days someone was available. The backtest that validated
this strategy auto-takes the top-ranked qualified setups up to capacity
every session. The forward record was therefore measuring strategy plus
discretion plus availability -- and degrading silently on busy days.
The shadow book closes that gap: it mirrors the backtest's selection rule
(top strategy_rank qualified, up to capacity, 1% fixed-fractional risk)
and shares the manual book's exit policy, so the only difference between
the two books is which setups get taken. Selection ordering reuses the
strategy_rank the scanner already stores rather than recomputing it, so
the two cannot drift apart. It runs as a near-close pipeline step right
after the scan, marking entries at the same prices a human would see.
Gate-reset re-entry state is now scoped per book -- the books diverge as
soon as their entries differ, and each must see only its own stops.
Performance view rewritten around the comparison:
- three series (shadow, manual, SPY) from a new endpoint
- SPY changes from a per-trade cost-basis counterfactual to plain
buy-and-hold %, since one line has to serve two books
- headline stats are R-multiples, not currency: the books size
differently, so only R compares across them
- configurable start date, because the strategy has been revised
repeatedly and pre-cutover trades ran under rules that no longer
exist
Migration 024 also repairs the numeric weekday crons written by 023,
rewriting only rows still holding the broken form so hand-corrected
settings survive. Its literals are inlined because bound parameters
render as NULL under 'alembic upgrade --sql'.
The shadow book is opt-in and writes nothing until enabled. Verify its
first selections match a backtest of that day's cross-section before
trusting any point on the curve.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Drop intermediate history-depth reports, sector-residual runners/map/code hooks
(evidence stays in final reports + docs), and slim MacBook helper to ssl/earnings/
prod-book-matrix only. SSL bootstrap and archived research conclusions retained.
Tier-1 alpha research (local only, no production deploy):
Sector residual momentum: two-factor SPY+sector residual and sector demean signals, IC harness + A/B. Sector resid clears pre-registered bars narrowly (PROMOTE for human wire design only). Sector demean fails t vs market resid.
Earnings: earnings_events backfill (FMP bulk paid; FMP/AV per-symbol), 2a gap diagnostic report-only, 2b SUE IC (PARK; incomplete 48/506 coverage).
History-depth: pre-registered doc + runner for MacBook deep rebuild/harness.
Do not ship production residual or filters from this branch.
Add research-only snapshot extender, PIT dollar-volume mask for signal IC,
rank-only harness path, fingerprint+breadth runner, and docs. Fingerprint
reproduced IC -0.045 / t -2.91 on prod.sqlite. No production gate/schedule changes.
Display-only Da/Gurun/Warachka information discreteness on the ticker
indicator panel. Shared compute with the backtest harness; not wired into
gate or rank.
Move the only qualifying R:R scan to 15:30 ET with chained Telegram alerts,
put outcome eval after a final-bar OHLCV fetch, enforce NY trading-day
requalify semantics, stamp paper trades fill_mode=near_close, and migrate
stored schedule_* keys to America/New_York.
Document Phase A (max-hold/vol/corr closed; next-open as decision baseline).
Add stale_close and next_open gap-cap fill modes plus a small matrix to test
whether near-close scheduling recovers overnight momentum drift.
Ship shared Sharpe SE/PSR diagnostics, next-open fill and equity-curve vol targeting in the portfolio simulator, re-derived fip_id, and a checkpointed offline matrix runner for Mac-side validation sweeps.
Honor custom S/R tolerance as a transient detect, refresh levels after OHLCV
mutations without failing committed price writes, report per-ticker S/R
rebuild failures from admin cleanup, and warn in the admin UI when refresh is partial.