35 Commits
Author SHA1 Message Date
dennisthiessenandClaude Opus 5 8453b87290 fix(sec): compose total debt across the styles filers actually tag
Deploy / lint (push) Successful in 10s
Deploy / test (push) Successful in 1m20s
Deploy / deploy (push) Successful in 37s
total_debt read LongTermDebt, else LongTermDebtNoncurrent/Current, plus one of
ShortTermBorrowings/CommercialPaper. That misses two whole tagging styles, and
it feeds net_debt -> net_debt_to_ebitda -> the categorical leverage read, so the
misses were not absences but confident wrong answers: Coca-Cola scored on 0.25bn
of commercial paper against ~39bn of debt, Verizon on 21.78bn of current
maturities against ~165bn, AT&T and Exxon produced no value at all against 134bn
and 33bn tagged. Measured over 19 large caps and 14 REITs, 11 were wrong or
absent and the rest are unchanged.

Each concept's span is now respected. LongTermDebt already includes current
maturities (Apple tags all three: 71.34 + 11.01 = 82.30), so only true
short-term borrowing is added. LongTermDebtAndCapitalLeaseObligations — what KO,
HD, T, XOM and CVX tag, and nothing read before — is noncurrent and takes a
current complement, and DebtCurrent *is* that whole complement rather than an
addition to it.

The REIT branch needed disambiguating: NotesPayable is not the same line across
issuers. MAA tags NotesPayable 5.66bn = UnsecuredDebt 5.30bn + SecuredDebt
0.36bn exactly, so there it is the total and adding the secured side
double-counts; EQR tags it alongside a larger SecuredDebt, where it is only the
unsecured component. UnsecuredDebt's presence separates them.

A component alone is no longer reported as a total. Chevron tags full debt only
in its 10-K, so its 10-Q carried 0.40bn of short-term borrowing; Boston
Properties tags SecuredDebt 4.28bn against ~15bn real. net_debt needs both sides
and yields nothing when either is missing, so None costs a leverage read where
the fragment produced a confidently wrong one.

Snapshots are immutable, so this corrects new filings only; stored history needs
scripts/reparse_fundamentals.py, which cannot complete until the EQR/931182
collision is retired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Lo99z3jWqu9X9ueBU3Z3D
2026-08-21 17:44:18 +02:00
dennisthiessenandClaude Opus 5 83fe76c506 fix(sec): alert when a filing gap's reprieve lapses instead of re-pausing quietly
An escalated gap stops pausing setups while the issuer's own fundamentals are
still recent. That reprieve ends on its own — the stored filings age past
GAP_GATE_RECENT_FILING_DAYS, or a newer gap arrives and the all-escalated
condition fails — and nothing reported either, because filing_gap_aged only
escalates gaps whose escalated_at is NULL and so never fires twice for the same
gap. For the 43 issuers behind the previous commit that lands around
2026-10-26, when their late-April filings age out together.

sec_filing_gaps.exempted_at (migration 034) makes the transition observable:
stamped quietly while the issuer is exempt, cleared when the exemption lapses,
and the clear is what raises filing_gap_repaused — once per lapse, re-arming if
the issuer's data recovers and ages out again. A gap that was never exempt has
no transition and stays silent; it is simply still paused, which
filing_gap_aged already said.

gap_exempt_ciks is public so the importer alerts on membership changes in
exactly the set the gate reads, rather than restating the rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Lo99z3jWqu9X9ueBU3Z3D
2026-08-21 17:44:04 +02:00
dennisthiessenandClaude Opus 5 c15b51439e fix(sec): tell an attribution collision apart from a changed reconstruction
Deploy / lint (push) Successful in 12s
Deploy / test (push) Successful in 1m43s
Deploy / deploy (push) Successful in 39s
snapshot_discrepancy named the accessions but not the columns, so it could not
distinguish "our numbers moved" from "the same filing is attributed twice".
The fields were already computed for validation_json and simply dropped from
the message; they are now in it.

A difference in cik ALONE is no longer reported as a reconstruction change at
all. Every fact matched, so two tracked CIKs are claiming one filing and the
fix is the universe, not the parser: it raises accession_cik_collision naming
both CIKs and sec_cik_overrides. It also never self-heals — the losing CIK
stores no row, so _ciks_with_snapshots never sees it and it is full-history
backfilled and re-reported every run until its ticker is re-pointed or retired.

Observed 2026-08-19 for EQR: after Equity Residential renamed to Vivmark
Residential (VMRK, CIK 906107), SEC's own company_tickers.json left the old
symbol on ERP Operating LP (CIK 931182), the non-traded co-registrant of their
combined 10-Qs. Both were tracked, both reconstructed the same two filings.

The reparse path now excludes cik-only differences from its rewrite set:
rewriting one would re-stamp the filing onto the co-registrant, taking it from
the issuer that actually filed it, which no parser fix asks for.

No stored value was wrong in that incident — reports/ carries the full
reproduction for this and for the companyfacts staleness behind the gate fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Lo99z3jWqu9X9ueBU3Z3D
2026-08-21 16:46:40 +02:00
dennisthiessenandClaude Opus 5 a13dbc9710 fix(sec): stop an unrecoverable filing gap pausing setups forever
A filing gap pauses its issuer until the filing is ingested or a later one
supersedes it, which assumes the gap is temporary. It is not always: SEC's
per-company Company-Facts files can go stale indefinitely — 43 large caps
whose Q2 10-Qs the frames API carries but whose companyfacts files never
received (Abbott's newest fact was 2026-04-29 in late August) — and because
the supersede rule needs a *successfully ingested* later filing, a stale file
swallows the next quarter too. The pause was open-ended, not seasonal.

So the pause hands off to the alert: once filing_gap_aged has escalated a gap,
it stops gating if the issuer's newest stored 10-K/10-Q is under 180 days old.
An issuer with nothing that recent has no usable fundamentals at all and stays
paused, which is the case the gate was built for.

Applied in the gate service only. active_gaps is deliberately untouched so
_retry_backlog keeps retrying and a recovered filing still resolves normally,
and the bound covers both gate paths — the queue and the validation_json
summary that mirrors the same filings — since bounding one leaves production
behaviour unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Lo99z3jWqu9X9ueBU3Z3D
2026-08-21 16:46:28 +02:00
dennisthiessenandClaude Opus 5 c97a067e0e fix(risk-monitor): drop an unused import and align migration 033 with its model
Deploy / lint (push) Successful in 8s
Deploy / test (push) Successful in 1m27s
Deploy / deploy (push) Successful in 37s
ruff F401 failed the deploy: `import pytest` in test_event_study.py outlived the
pytest.approx assertion it was added for.

Compiling the migration for Postgres while checking that turned up a second
defect worth fixing while the table is still empty. It created a unique
constraint *and* a plain index on effective_date, while the model declares
`unique=True, index=True` -- one unique index. Both enforce uniqueness, but the
pairing left a redundant second index on the column and a permanent diff for
autogenerate to keep trying to reconcile. Now renders byte-for-byte what the
model declares, matching RegimeSnapshot.date.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 11:31:05 +02:00
dennisthiessenandClaude Opus 5 333989eeab feat(risk-monitor): measure the rule that fires, and give fundamentals their own channel
Deploy / lint (push) Failing after 11s
Deploy / test (push) Skipped
Deploy / deploy (push) Skipped
The Warning study measured a fitted percentile crossing that nothing consumes.
What reaches Telegram is a quadrant change: fixed 50/40 dividers, hysteresis,
two-session confirmation, 3-day cooldown. Those thresholds are constants, not
fits, so there is no training set to protect and all 11 detected corrections are
evaluable instead of the 4 that fell in a holdout.

Replaying it: 1/10 corrections, 0.9 false alarms/year. Random alarms at the same
firing rate match or beat that in 65% of draws. The panel now carries ablations
(does the quadrant machinery earn its place?), external baselines (does the score
earn its complexity?), and that null, because a bare "2 of 4" was unreadable in
either direction. Nothing in the alert path was retuned on the strength of it.

Fundamentals become a third channel rather than a term in either score. v3 cut
them arguing 12+8 of 100 points "could not change any published conclusion" --
true only when every technical sensor reads zero; weighted they moved the bar for
the 40 divider from 40 to 25. But no fusion weight is measurable either: with ~10
events and no fundamental history, any weight is a policy preference presented as
a measurement. So the read is a categorical state (supportive/neutral/adverse/
unknown) with an evidence grade, derived by fixed rules from stored facts, read
by confluence. The LLM extracts and explains; it does not score.

Absence stays absence throughout. `unknown` is unreachable by averaging, a stale
or empty observation may display but never confirm, extraction failures map to
`unknown` rather than `mixed`, and the study rows are coverage-matched and marked
not-measurable until enough corrections are covered -- otherwise a fortnight of
observations renders as 0/10 and reads as a failed test.

Observations become a real time series (migration 033); they lived in a single
overwritten settings slot, so no history existed to replay. Pre-rename snapshots
are adapted rather than discarded. METHODOLOGY stays v4 -- no score changed --
so no reseed; STUDY_SCHEMA moves to 3 and discards the cached report.

Post-deploy: re-run Event Study from Admin -> Jobs. The panel reads "not run yet"
until then.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 11:15:09 +02:00
dennisthiessenandClaude Opus 5 3033ad83fd docs(readme): catch the README up with the last 20 commits
The README still described capacity 10, a two-tab Signals page, cascade-delete
ticker retirement, and a discretionary paper book as the out-of-sample proof.
All four are wrong against the current tree.

Capacity: SIM_MAX_POSITIONS is 15 since the 2026-08-05 bracket. The summary and
flowchart now say 15; the re-entry study and tuning rows keep 10 and are marked
as measured at the then-production capacity, because rewriting numbers a study
did not produce is worse than a stale one. The tuning table claimed the 10-slot
cap never binds -- that read came from EV per trade and is what the bracket
reversed. Gate reset was promoted at capacity 10 and capacity 15 is the one arm
where immediate re-entry edged ahead, so that gap is written up as an open
question rather than resolved by edit.

The shadow book was missing entirely, and it contradicts what the README claimed
as the OOS record: the discretionary book measures the strategy plus discretion
and availability, which is the gap the shadow book exists to close. Documented
as opt-in, with its near-close pipeline step, its parity invariant, and the
Dashboard chart that actually renders it (not the Paper Trades tab).

Also: Signals is Setups / Paper Trades / Backtest; the iron rule pointed at a
Signal edge table the UI no longer renders, now redirected to the local report;
delisting replaces cascade delete; SEC promotion ceiling; SEC_USER_AGENT and the
DeepSeek/xAI/Dolt/backtest env vars; ~15 missing endpoints; the systemd unit
filename; npm test no longer exists as a script.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:31:03 +02:00
dennisthiessenandClaude Opus 5 044a3447f6 fix(backtest): rebuild the recommendation on read, not only on run
Deploy / lint (push) Successful in 8s
Deploy / test (push) Successful in 1m13s
Deploy / deploy (push) Successful in 36s
Every fix so far only applied to reports generated after deploy. The cached
report is served verbatim, so it keeps the recommendation the OLD build stored —
quoting the legacy policy book, naming a rejected exit as "recommended", and
carrying no basis_lookback, which let the lookback selector default to 3y and
put 3-year tiles beside an all-history recommendation with no divergence notice.
Exactly the contradiction the last three commits set out to remove, silently
present on the first page load after deploy and until the next scheduled run
overwrote it.

The recommendation is a pure function of the numbers already in the report — its
own note says it is derived from them on every run — so it is now re-derived on
read. A corrected recommendation appears immediately instead of after the next
backtest. On failure it is dropped rather than falling back to the stored one,
which is the stale derivation this replaces.

The test drives the real shape: an old-build report with a legacy recommendation
written straight to the settings row, read back through get_backtest_report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 08:33:32 +02:00
dennisthiessenandClaude Opus 5 21a5fc8a52 fix(backtest): flag a lookback the recommendation was not computed on
Selecting a different window or a comparison strategy silently made the tiles
stop matching the recommendation below, which is baked into the report and
cannot follow a dropdown. On load they now agree by construction; moving off
that basis says so.

Also: an absent production row produced no headline and no benchmark, but any
passing gate finding still rendered a green "no warnings" chip — a success badge
for missing data, directly beside "this report predates the portfolio monitor".
Missing baseline now reads "baseline unavailable".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 08:03:12 +02:00
dennisthiessenandClaude Opus 5 11dffcd695 fix(backtest): one window, and stop calling a rejected exit "recommended"
Two ways the recommendation still disagreed with the page it sits on.

It preferred the "all" monitor row while the UI defaulted its selector to "3y",
so a default page load showed one set of returns in the tiles and a different
set in the recommendation. The row it used is now published as basis_lookback
and the page defaults to it, so the two cannot open on different windows. The
test fixture gains a second monitor row with different numbers — with only an
"all" row present, a lookback mix-up could not fail.

Robustness picked its basis between "the recommended Nd hold" and "the S/R
target exit", naming an exit the production book replaced as recommended. There
is no ATR-trail ex-top-5% figure in the report, so it now always reports the
gate-level grading and says that is what it is, rather than dressing a legacy
number as a verdict on the production book. time_exit_sweep is no longer read
here at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 08:03:12 +02:00
dennisthiessenandClaude Opus 5 3a2d548610 feat(backtest): make the tiles answer "is that good?", and stop the layout jumping
Five UI problems, all reported from using the page.

Expanding "How this is measured" shoved every control down, because the
disclosure and the run controls shared one flex row. They no longer do: run
status and the controls that start a new run sit together on one line, and the
explainer is below them where growing it moves nothing.

A long strategy name wrapped the dropdown trigger onto three lines and dragged
the row out of alignment. The trigger now truncates with the full text on hover
— a wrapping dropdown is broken anywhere, so the fix is in the primitive — and
the twelve-character "Production: " prefix is a bullet.

"Sortino 2.72" answered nothing. Each risk-adjusted metric now carries a meter:
a track showing where the value sits, ticks at the band edges, and the band word.
Colour never travels alone. Bands are deliberately stricter than textbook ranges
because this universe is today's survivors replayed backward, which flatters
every ratio — that caveat is stated next to them rather than left implied.

The two tile rows were different sizes, which read as inconsistent rather than
as hierarchy. Every tile is the same size now and grouping carries the ranking:
top row is raw outcome and takes no meters, second row is risk-adjusted ratios
and all take meters. Sharpe moved down to join them — it is one of those ratios,
and leaving it above made it the only metered tile in a row of bare ones.

The recommendation led with a long bold sentence that describes the
configuration, not a verdict, while the actual findings were small grey text.
Findings now come first, each split into label and detail on the colon the
backend strings already carry, and the configuration is a footer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 23:25:11 +02:00
dennisthiessenandClaude Opus 5 28b3273150 fix(backtest): quote one book, not two
The recommendation's "Book vs SPY" line read from portfolio_sim — the hold/target
policy book — while the tiles directly above read from portfolio_monitor, the
production ATR-trail book. Same SPY figure, different portfolio return, on one
screen. It now reads the same production row the tiles do.

Also removed, for the same reason: the "legacy exit diagnostic" comparing hold
against the S/R target. Both are exits the production book replaced, so a
recommendation between them could not lead to an action. And the fallback
headline, which advised the fixed-hold exit whenever a report had no production
row — a report that cannot describe the production baseline now states none.

portfolio_sim stays in the report payload: scripts/run_backtest_snapshot.py and
reports/compare_reports.py read it, and it is no longer surfaced in the UI. The
test fixture now carries a production monitor whose numbers differ from its
policy sim, so re-sourcing that line from the old place fails rather than passes
unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 23:24:58 +02:00
dennisthiessenandClaude Opus 5 14cfa44fc5 refactor(signals): drop the now-unused fmtMoney helper
Deploy / lint (push) Successful in 9s
Deploy / test (push) Successful in 1m16s
Deploy / deploy (push) Successful in 38s
It was lifted from BacktestPanel during the extraction, then lost its last
caller in the same branch when avg_trade_pnl became an EV / trade tile rendered
with fmtSignedMoney and disappeared from the monitor footnote. formatPrice
already covers a bare unsigned amount if one is ever needed again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:40:37 +02:00
dennisthiessenandClaude Opus 5 13a984a84d refactor(signals): split the Track Record tab and cut the backtest page down
One tab stacked three things that all called themselves a track record:
realized paper P&L, setup-outcome grading under the rejected take-profit model,
and the backtest portfolio simulation. Split into Setups | Paper Trades |
Backtest, one subject each. `track` stays the Paper Trades slug so the legacy
/performance redirect keeps working. The grading diagnostic and its Evaluate /
Reset controls go with Backtest, not Paper Trades — reset_track_record deletes
trade_setups, not paper trades.

BacktestPanel 439 -> 175 lines. Its run settings alone were 106 lines of
hand-rolled sr-only radio cards for two binary choices; they are now two
Dropdowns and a button on one wrapping row, with the per-option prose moved into
the existing explainer. The amber warnings survive as a conditional slot, so a
non-default choice still announces itself but the common path is silent.

The recommendation printed eight findings at equal weight, burying the verdict
in tuning detail. `topic` now splits them: production, benchmark and robustness
stay inline, gate/exit/cutoff collapse behind a disclosure, and any WARNING or
LAGS item is promoted out of the collapsed group regardless of topic. No topic
chips — every backend string already self-prefixes, so a chip would render
"GATE | Gate: ...".

Portfolio metrics are now two tiers: five headline tiles for what the book
returned, then a smaller labelled row for how good that return was (Sortino,
Calmar (MAR), Gain/Pain, Profit Factor $, EV/trade). Reports cached before those
metrics existed hide the second row rather than showing a half-populated line of
dashes.

Extracted EquityCurveChart, PortfolioMonitorPanel and BacktestRecommendationCard,
plus a StatTile primitive and shared formatters for the duplication in the files
this touched. DashboardPage and OpenTradesPanel deliberately keep their own
copies — migrating them is separate scope.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:09:23 +02:00
dennisthiessenandClaude Opus 5 442dc3f04b feat(backtest): add Sortino, Gain-to-Pain and dollar profit factor
Three portfolio metrics computed where their inputs already live in
_simulate_portfolio: Sortino off the existing daily return series, Gain-to-Pain
off a monthly aggregation of the equity curve, profit factor off closed-trade
dollar P&L.

Gain-to-Pain follows Schwager — sum of ALL monthly returns over the absolute
sum of the negative ones. The profit-factor-shaped variant,
sum(positive)/|sum(negative)|, sits exactly 1.0 higher for every input since
sum(all) = sum(pos) - |sum(neg)|; the test asserts against both so the wrong one
cannot pass. Sortino divides by len(rets), the full-sample lower partial moment,
not by the count of down days, which would shrink the denominator and inflate
the ratio.

No MAR field: calmar is already CAGR / max drawdown, the same number under the
other name (docs/research/effective-risk-floor-ab.md).

All three keys are emitted unconditionally even when None — the UI reads an
absent key as "report predates these metrics", so presence is a contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:09:10 +02:00
dennisthiessenandClaude Opus 5 247a7889b9 fix(tickers): honor the effective date instead of retiring on the mark
Deploy / lint (push) Successful in 11s
Deploy / test (push) Successful in 1m22s
Deploy / deploy (push) Successful in 38s
active_only tested delisted_on IS NULL, so a symbol dropped out of signals the
moment a Form 25 was detected — ten days before Rule 12d2-2 makes the removal
effective, while it was demonstrably still trading. A manual future-dated mark
behaved the same way. It now compares against the database's own date, so a
pending delisting stays live until the day it takes effect.

That exposes a second problem the fix would otherwise create. Trading typically
stops before the ten-day delay expires, so across that window the symbol is
correctly active yet produces no bars — and confirm_delisting returned None for
an already-marked row, which would have fired the staleness warning daily for
ten days, the exact noise this flow exists to remove. It now reports the known
effective date on every path where the delisting is established, so the caller
warns only about gaps that are still unexplained.

CURRENT_DATE renders identically on postgres and sqlite, and the OR is
parenthesized when callers chain further where clauses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:58:51 +02:00
dennisthiessenandClaude Opus 5 1d4ed39fd2 fix(tickers): close the delisting review findings
Detection could retire an actively traded symbol — silently, since it then
vanishes from every signal. Three causes:

- Form 25 is filed per security class. An issuer removing its notes, preferred
  or warrants files one while the common keeps trading. The filing's own
  descriptionClassSecurity distinguishes them, so the primary document is now
  fetched and read; anything not recognisably common equity is rejected, as is
  anything unreadable (pre-2009 filings have no primary_doc.xml). Fail closed.
- Form 15 ends a reporting obligation and is no evidence trading stopped. The
  whole family is dropped.
- A historical filing for a long-gone class could retire a symbol whose bars ran
  years later, stamping the old date. Filings before the last bar (less a 30-day
  lead for the exchange) are now ignored.

Rule 12d2-2 makes removal effective ten days after filing, so delisted_on is the
effective date rather than the filing date.

bootstrap_universe(prune_missing=True) still ran a cascading delete over
delisted rows, undoing the retention this branch exists for; it now skips them
and reports kept_delisted so the count is explicable.

clear_delisted had no route, which made "safe to automate because it is
reversible" false — reversal needed SQL. POST/DELETE /tickers/{symbol}/delisting
now mark and un-mark, giving an operator a non-destructive alternative to the
cascading DELETE that was the only option.

Shared-CIK siblings (GOOG/GOOGL) stay safe by construction: the probe is
per-symbol and gated on that symbol's own staleness, so a class that still
trades is never probed.

Not addressed: pruning a symbol merely dropped from the index still destroys its
history — the same survivorship problem in a different costume, needing a
tracked/membership state separate from delisting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:58:51 +02:00
dennisthiessenandClaude Opus 5 d950fcf70e fix(tickers): let an SEC confirmation upgrade a manual delisting mark
mark_delisted returned early on any already-delisted row, so the sequence an
operator actually hits — mark EA by hand today, Form 25-NSE surfaces three days
later dated 2026-08-04 — left the estimated date and "manual" reason in place
permanently. Form 25 carries the real effective date, so it now replaces an
operator's estimate; a confirmed row is never downgraded or re-probed.

Also cover _get_ohlcv_priority_tickers, the one place active_only wraps a
compound select rather than a bare one — the unit suite reached none of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:58:51 +02:00
dennisthiessenandClaude Opus 5 6501b7e9a0 feat(tickers): record delisting instead of deleting the symbol
Retiring a symbol meant delete_ticker or bootstrap_universe(prune_missing),
both of which cascade through OHLCV, setups and scores. That destroys exactly
the history four research documents already apologise for: today's tracked
universe projected backward is survivorship-biased, and hard-deleting every
delisted name is what causes it. Keeping the rows preserves the option to fix
that — it does not fix it, which needs the replay to model a delisting as an
exit event.

tickers gains delisted_on / delisted_reason (migration 032). NULL means
actively traded.

The filter is opt-in via ticker_service.active_only rather than folded into a
shared getter: the registry and admin views deliberately keep delisted rows so
the delisting is visible, and a silent default would undo that. Applied to the
live path only — scanner, momentum ranking, scoring, breadth, fundamentals
candidates, SEC universe, earnings import, ingestion loops. run_backtest keeps
them on purpose.

Detection runs off OHLCV staleness, not off the SEC fundamentals import: that
importer stalls for days on unrelated Company-Facts gaps and would take
detection down with it. On a stale symbol the scheduler asks SEC for a Form
25/25-NSE/15 and retires it only on a hit, so a halt or a rename (SATS->ECHO)
keeps the existing warning. The probe waits 3 stale days so a market-data
outage cannot turn into one SEC request per symbol per run.

Safe to automate because it is reversible: clear_delisted un-retires a false
positive, where a delete had already taken the history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:58:51 +02:00
dennisthiessenandClaude Opus 5 486fb500d1 test(sec): cover the ceiling's promote/queue/alert path end to end
The ceiling tests asserted validate()'s verdict but nothing proved the claim
the design rests on: that a forced promotion actually queues the filings it
released and says that it did. Drive it through run_import with a filing that
stays inside the per-filing window, so only the aggregate ceiling can release
it, and assert the SecFilingGap row and the promotion_ceiling_forced event.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:58:46 +02:00
dennisthiessenandClaude Opus 5 77570557db feat(sec): cap how long the fundamentals import can stay deferred
MISSING_XBRL_RETRY_DAYS bounds how long ONE filing blocks promotion. It does
not bound the import as a whole, and the two come apart because a blocking
filing is only queued by promote(), which a deferred run never reaches. During
a rolling supply of unresolvable filings — earnings season, when SEC's
Company-Facts aggregation lags furthest — each new arrival restarts the 3-day
clock before the previous one clears, and nothing is written at all: not the
good rows, not the gap rows that would stop those filings blocking again.

Add an aggregate ceiling. Once promotions have been stale for
PROMOTION_CEILING_DAYS (7), every unresolved filing is aged past the retry
window in place, so promote() queues them all through the path that already
exists, source_max_date advances, and _missing() keeps queued rows aged-out on
later runs. The import self-heals instead of compounding.

Deliberately not the alternative of queueing gap rows on a deferred run: that
would drop the grace period to a single run for every filing, including the
common case of a Company-Facts lag that resolves in a day, and it needs a write
on a run that failed validation.

The per-filing window is untouched, a never-promoted source never trips (that
is initial setup, not a wedge), and affected symbols stay barred from setups
either way since setup_blocked_ciks ignores the window. A forced promotion
raises promotion_ceiling_forced so the safety valve is never silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:58:46 +02:00
dennisthiessenandClaude Opus 5 fbca38e144 fix(backtest): roll back the portfolio-sim DB failures too
The first pass guarded the replay loop but not the portfolio-simulation block,
which re-fetches price columns and loads the benchmark and the live exit policy
from the same session much later. A failure in any of those swallows the
exception without clearing the transaction — the identical failure mode, with
the identical symptom: the report write is the first unguarded statement and
takes the blame.

The outer handler is the backstop for the price_columns loop, which has no
handler of its own; rolling back a session an inner handler already cleared is
a no-op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 16:10:57 +02:00
dennisthiessenandClaude Opus 5 6ca7f13779 fix(backtest): roll back the session after a swallowed DB failure
Every DB call in run_backtest is best-effort so one unreadable ticker cannot
abort the whole replay, but the handlers swallowed the exception without
clearing the transaction. asyncpg then reports "current transaction is
aborted" for every later statement, and the first unguarded one — the report
write — surfaced it as the job error, long after the real cause.

Add _rollback_quietly at the three swallowing sites (benchmark load, parallel
fetch, sequential replay), matching the guard price_service already uses.

Load plain symbols instead of Ticker instances: a rollback expires ORM objects
held across it, and touching an expired attribute afterwards triggers sync
lazy-loading, which raises on an AsyncSession. rr_scanner_service hit this
same trap. Only .symbol was ever used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 14:20:19 +02:00
dennisthiessenandClaude Opus 5 d02fd82ced docs(research): regenerate the artifact under the v2-mandatory harness
Deploy / lint (push) Successful in 8s
Deploy / test (push) Successful in 1m8s
Deploy / deploy (push) Successful in 35s
Supersedes the previous artifact, which predates v2_reconstruction becoming a
required variant and the P1-cap conditional reporting. Generated from a clean
tree (git_dirty false at rev 43ee619), all hard gates passing, so the recorded
source hashes actually identify the code that produced it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:16:52 +02:00
dennisthiessenandClaude Opus 5 43ee619412 fix(research): require the v2 reproduction, and correct the P1-cap denominator
Two review findings, plus a lost-edit repair.

v2_reconstruction is now a required variant. It carries every published figure
the reproduction rests on (avg, p80, max, P3-pegged, W1-live), so a run without
it could emit a confident, non-provisional recommendation having checked nothing
against v2 at all -- while the methodology doc claims v2 and v3 are reproduced
first. The default invocation is now derived from REQUIRED_VARIANTS so the two
cannot drift, and a test asserts the default satisfies its own requirement.

The doc and the P1_TREND_BREAK_ANCHORS comment still justified skipping the
P1_SCORE_CAP with 17/408 = 4.2%, which is the all-session share and does not
evaluate the rule. The rule names sessions with State >= 40: 47 of them, P1 sole
argmax on 17 = 36.2%, against P2's 16 and P3's 14. Conclusion unchanged -- well
under the 80% trigger -- but the published rationale now states the metric that
actually decided it.

Root cause of that survival: the earlier correction WAS made, but in a script
that applied several substitutions and wrote the file once at the end. A later
substitution raised, so the successful edits were discarded with it. The
"Unlike P3 and V1 ... P3's do not" fix was lost the same way and is restored.

Also adds tests for the refusal paths themselves -- missing required variant,
unknown variant, custom window with no calendar anchor. They were verified by
hand last round but left unpinned, which is the same shape of problem as the
optional gates they exist to enforce. All return before any network call.

Deliberately not done, as not load-bearing: recording the oas400 variant's
missing-credit session count (the truncation conclusion rests on the
distribution mismatch, which is already recorded), and generalising
_pipeline_gates for arbitrary --end/--sessions windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:16:03 +02:00
dennisthiessenandClaude Opus 5 87224a1451 docs(research): regenerate the calibration artifact from a clean tree
The previous artifact was produced from a dirty working tree while HEAD still
pointed at the harness commit, so its recorded revision could not reproduce it.
This one records git_dirty false alongside sha256 of the three source files it
depends on, so the claim "checking out this revision reproduces this artifact"
is now checkable rather than implied.

All hard gates pass, including the first-scored-date anchor and the row-wise
state_v4 <= state_v3 invariant, so it carries a non-null recommendation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:46:41 +02:00
dennisthiessenandClaude Opus 5 ec1b0acfad fix(research): make the calibration artifact live up to its refusal guarantees
Review of the v4 evidence path. The shipped sensors, bands, methodology bump and
categorical allowlist were found sound; these are gaps in the harness that
produced the evidence for them.

The recommendation gates were optional, so they were not gates. The calendar
anchor lived behind --expected-first-session, which defaulted to None -- so the
committed artifact had no first-date check at all, leaving only a session COUNT
that is tautological (the harness slices the tail of the price series to whatever
was asked for). And the state_v4 <= state_v3 invariant was appended only when
both variants were present, so `--methodology v3` alone could still emit a v4
recommendation having never evaluated v4. The anchor is now a published constant
asserted unconditionally, required explicitly whenever --end/--sessions are
overridden, and v3+v4 are mandatory. Both refusals exit 2.

The P1_SCORE_CAP decision was taken on the wrong population. The agreed rule was
"sole price argmax on >80% of sessions with State >= 40"; the harness reported
only all-session counts and the doc concluded from 17/408 = 4.2%. Measured on the
actual population: 47 qualifying sessions, P1 sole argmax on 17 = **36.2%** (P2
16, P3 14). Still well under 80, so the conclusion holds -- but it was reached
from a denominator that did not test the rule, and 36.2% is a materially
different number to have on the page.

Provenance did not identify the code that produced the artifact. It recorded
git_rev c3ae5ad while the live v4 variant depended on app changes that were still
uncommitted, so checking out that revision would not reproduce it. Now records
git_dirty plus sha256 of regime_monitor_service, breadth_service and the script
itself, and this artifact is regenerated from a clean tree.

The 400- vs 700-day OAS question was described as settled but was not
reproducible: the artifact carried only oas_fetch_days 4748, and
v2_reconstruction patches the per-session window to 3653 regardless, so
--oas-window-days 400 could not simulate it. Patching a window cannot stand in
for data that was simply absent, so v2_reconstruction_oas400 truncates the OAS
SOURCE series instead: avg 26.54, p80 42.52, max 100.00 against published
22.6 / 35.1 / 91.2. Full coverage reproduces all three, so the published figures
predate the truncation. Now recorded in the doc.

v4-vix-only and v4-p1-only had become no-ops: after the cutover the shipped
sensors ARE v4, so patching one candidate in left the other shipped and both
variants evaluated full v4. Each now restores the other sensor to its v3 formula,
and they separate properly (v3 18.13, v4-vix-only 16.64, v4-p1-only 16.28,
v4 14.78 -- each fix contributing about half the move).

Docs: the copy-paste invocation was mangled by a backslash-escaping bug and is
now a fenced, forward-slash command; "Unlike P3 and V1 ... P3's do not" corrected
to "Unlike P1 and V1"; the point-in-time section updated from 400 sessions to the
672-calendar-day / ~464-session window production actually replays; the exercised
52.33 VIX print recorded so the top anchors are not merely asserted.

Tests: band_for now pinned at 64.9/65 from both sides so a silent revert to 80
cannot pass, and the categorical carry-forward test stores locked=True and
asserts it survives -- losing it is half the failure mode, since
update_regime_monitor only auto-refreshes when locked is false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:46:10 +02:00
dennisthiessenandClaude Opus 5 3143477a62 feat(regime): cut the risk monitor to v4 — desaturate VIX and the trend break
Two sensors saturated in exactly the range where resolution matters, and the top
State band had no headroom. Calibrated with scripts/run_regime_monitor_calibration.py
over the 408 sessions ending 2026-07-24; the shipped code reproduces that run's
band shares exactly (78.9 / 13.0 / 4.7 / 3.4).

V1 read VIX 30, 50 and 82 as an identical 100 — the same defect v3 had just
removed from P3, left in place one sensor over. In the window it flattened five
distinct April-2025 prints (52.33, 46.98, 45.31, 40.72, 38.57) into one value.
Now an anchor table reaching full scale at 55, not at 2020's ~82: anchoring the
top at a once-in-a-generation print would make VIX 50 read only ~70. Pegged on
14 of 408 sessions before; none now.

_under_200 returned a bare 0/100, so P1 printed 100 the moment SMH and QQQ were
both under their average — and since the price pillar takes max(P1, P2, P3),
that pinned the pillar and stopped P3's ladder resolving for the whole of a
selloff. Now graded by depth below the 200-DMA, with a deliberate floor of 20 at
the crossing: the break is a genuine binary event, only its depth is graded.
Pegged on 46 of 408 sessions before; none now. A 2% break reads ~30, not 100.

max() was KEPT — the defect was the step function feeding it, not the vote, and
v3's "one capped vote for correlated reads" rationale still holds. P1 is the sole
price argmax on 17 of 408 sessions (4.2%), so the P1_SCORE_CAP fallback drafted
during design was measured as unnecessary and not shipped.

STATE_BANDS breaking 80 -> 65, and only that threshold. Credit returns 0.0 (not
None) when calm, so it holds its 20 points pinned at zero and price + breadth +
volatility at literal maximum summed to exactly 80.0 — v3's threshold to the
decimal, with nothing above it. The sensor is deliberately unchanged: a
calm-credit selloff genuinely is less stressed. What was stale is the band, fit
on v2 while credit's since-removed percentile leg still contributed. A
2022-style AI/tech drawdown with calm credit computes to 70.3 (no death cross) or
74.0 (with one); 70 would have left 0.33 points of headroom, reproducing the
defect. Chosen by scenario arithmetic, and the realized breaking share then lands
on 3.4% — the same as v3's, arrived at independently.

"v4" added to CATEGORICAL_FUNDAMENTAL_METHODOLOGIES in this same commit, which is
load-bearing: that set is checked against the STORED blob, so bumping without it
discards the collected observation on first write, leaving fetched_at null and
locked false — and update_regime_monitor then fires a paid LLM refresh on every
run, forever. Now guarded by a test parametrised over v2 and v3 stored blobs.

SENSOR_REVISION deliberately stays 2: a METHODOLOGY change already forces a full
reseed via _parse_snapshot, and bumping both would imply the reseed was
revision-driven.

QUADRANT_STATE_DIVIDER stays 50 because only breaking moved, so alert_service,
RegimeChart and the quadrant tests need no change. A new test enforces
divider == band boundary on both axes, which nothing did before.

Doc renamed to regime-monitor-v4.md with a tombstone at the old path (commit
messages cite it), the three open questions converted to resolved with the
reasoning that closed them, and indexed in docs/research/README.md for the first
time. The P2 limit is stated honestly: _death_cross pegs at a -5% MA gap, so a
deep selloff still reaches 100 via P2 — v4 repairs the shallow-to-moderate break,
not "the price pillar no longer pegs".

DEPLOY: the first run reseeds ~464 sessions. Expect one phantom quadrant alert
(the dedup key carries basket_hash, not methodology) and re-run the Event Study
manually — its cached report self-invalidates but does not self-regenerate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:34:25 +02:00
dennisthiessenandClaude Opus 5 c3ae5ad949 feat(research): commit the regime-monitor replay harness, and reproduce v2/v3
v3 was calibrated by replaying the series offline, but that harness was never
committed -- so its published numbers could not be re-derived, and a v4 cut would
have had to choose anchors by argument rather than measurement. This is that
harness, and it reproduces the published figures.

scripts/run_regime_monitor_calibration.py replays State/Warning session by
session from the same inputs the live job uses (Alpaca for all 33 symbols, FRED
for VIX and HY OAS), with no database: breadth and divergence come from
breadth_service's pure helpers. It never reimplements an unchanged live sensor --
_compute_index, _score_pillars, P2, P4 and the Warning sensors are imported and
called. Only candidate formulas (proposed v4) and retired ones (v2, gone from the
codebase) are defined here and patched onto the module for a variant's duration.

Reproduction of the 408 sessions ending 2026-07-24, against the figures in
docs/research/regime-monitor-v3.md:

  v2 State avg      22.6   ->  22.68
  v2 State p80      35.1   ->  35.1     exact
  v2 State max      91.2   ->  91.2     exact
  v2 P3 pegged        39   ->    39     exact
  v2 W1 live         108   ->   108     exact
  v3 State max      87.4   ->  87.4     exact
  v3 band shares  73.3/15.0/8.3/3.4 -> 73.0/15.4/8.1/3.4

Three things the harness had to get right to reach that, each of which was
initially wrong and caught by a gate rather than by inspection:

  - "W1 live 108" counts NONZERO sessions, not non-null ones. v2's divergence
    gate returned 0.0 during any decline (v3 tapers instead), so the retired
    divergence formula had to be reconstructed too.
  - v2 sliced HY_OAS_REFERENCE_YEARS = 10.0 per session, not v3's 700 days. The
    percentile leg ranks against that window, so replaying it short shifted the
    middle of the distribution while leaving the max exact.
  - The published v2 numbers correspond to FULL OAS coverage. Replaying v2 with
    the 400-calendar-day fetch it shipped with yields max 100.0 and 133
    credit-less sessions -- so that truncation was not in force when the figures
    were taken. Recorded rather than assumed.

The script refuses to emit a band recommendation unless every hard gate passes
(33 symbols fetched, per-symbol warm-up and final bar, full basket on every
session, calendar anchors, 100% coverage, and a row-wise state_v4 <= state_v3
invariant), and exits non-zero. It is meant to be structurally impossible to read
a calibration result out of a run whose pipeline did not validate. No v4 code
ships in this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:05:12 +02:00
dennisthiessenandClaude Opus 5 f22313deaf chore: remove dead frontend code and one unused service helper
Deploy / lint (push) Successful in 8s
Deploy / test (push) Successful in 1m13s
Deploy / deploy (push) Successful in 37s
Scan of every module and exported symbol, with each candidate verified by hand
rather than trusted from the scan.

Deleted outright:
  frontend/src/lib/fundamentals.ts  (112 lines, 12 exports) — imported by
    nothing, including FundamentalsPanel, which reads backend values. It mirrors
    scoring_service._compute_fundamental_score, so it is the same *kind* of
    thing as lib/qualification.ts — but nothing consumes it, so it mirrored
    nothing and could drift out of sync unnoticed.
  Skeleton.SkeletonLine, paperTrades.getEquityCurve, regime.regimeColor
  breadth_service.compute_breadth_today — self-described "thin wrapper, for
    future live use"; that future did not arrive.

Kept, but unexported — used inside their own module, so the dead part was the
public surface, not the code: Button.Spinner, exitPlan.SETUP_STOP_ATR_MULTIPLIER,
client.ApiError.

Three things the scan flagged that are NOT dead, recorded so the next sweep does
not re-raise them:
  RegimeChart.tsx — lazy(() => import(...)) in RegimePage, so it looks orphaned
    to any importer-graph scan. Deleting it would break the risk page.
  qualification.ts MIN_TARGET_PROBABILITY / liveRiskReward — that file is a live
    mirror of app/services/qualification.py used in five places, and the
    constant is exported to document the backend value it tracks.
  ssl_bootstrap.ssl_status — called from an inline python snippet inside
    scripts/run_tier1_macbook.sh, invisible to a .py-only search.

No orphaned backend modules across app/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 17:24:04 +02:00
dennisthiessenandClaude Opus 5 d116fcc146 chore(frontend): drop the dead FundamentalsPanel dev harness
Deploy / lint (push) Successful in 9s
Deploy / test (push) Successful in 1m12s
Deploy / deploy (push) Successful in 37s
Left over from the fundamentals work. Verified orphaned before removing: not in
vite.config.ts (default build input is index.html alone, and dist/ only ever
contained index.html, so it never shipped), not referenced by any script,
config or module, and imported by nothing. Its only mention anywhere was its own
header comment — the many other "harness" hits in the repo are the factor /
backtest IC harness, which is unrelated and stays.

Removes frontend/src/dev entirely.

The pattern it embodied — a fixture-seeded page for eyeballing one component
without a backend — is still the only way to see a component render, since the
frontend has no test runner. But that is worth recreating per component on the
spot, not preserving as a stale file pinned to one panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 17:13:11 +02:00
dennisthiessenandClaude Opus 5 94baa89423 fix(jobs): make last-run writes atomic and drain them properly on shutdown
Deploy / lint (push) Successful in 11s
Deploy / test (push) Successful in 1m18s
Deploy / deploy (push) Successful in 42s
Two review findings on 22bee28, both reproduced before fixing.

[1] Concurrent writes could lose or rewind a row. record_finish did
select-then-insert-or-update, and pipelines are separate scheduler jobs that can
overlap while sharing step ids -- data_collector belongs to all four. Reproduced
both halves: two sessions that SELECT before either INSERTs make the second
commit raise IntegrityError, which _persist_job_run swallows, so the run
silently vanishes; and a later write carrying an OLDER finished_at rewound the
row from 12:00 back to 09:00, dragging the status with it, so the panel would
report a stale outcome as the latest.

Now a single atomic INSERT ... ON CONFLICT (job_name) DO UPDATE, guarded by
WHERE job_run_state.finished_at < excluded.finished_at so an older completion
can never overwrite a newer one. Dialect-specific because prod is Postgres and
tests are SQLite; both support it (SQLite >= 3.24, ours is 3.45). updated_at is
set explicitly, as the model's onupdate hook does not fire for a core upsert,
and get_map now uses populate_existing since core writes leave any
previously-loaded ORM instance stale in the identity map.

[2] Shutdown could dispose the engine underneath a pending write.
scheduler.shutdown(wait=False) returns before APScheduler dispatches its
completion events, and those events are what create persist tasks -- so a single
snapshot of the task set missed writes still to be queued. flush_job_run_persists
now settles briefly for pending callbacks, then drains in a loop until the set
stays empty, with the deadline still bounding total shutdown time. Left
shutdown(wait=False) alone deliberately: waiting would block a deploy restart
behind a long-running scan.

Six regression tests: interleaved first writes, older-never-rewinds,
newer-still-wins, a task queued mid-drain, prompt return when idle, and giving
up rather than hanging shutdown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:15:05 +02:00
dennisthiessenandClaude Opus 5 4b4a1084cb refactor(jobs): group Admin -> Jobs into sections instead of one flat list
Nineteen jobs rendered as one alphabetical list in which a pipeline, one of its
steps, a standalone cron job and a manual-only job were indistinguishable.

Four sections, ordered by the API (category rank, then trading-day order within
it) so the client does not re-derive ordering: Pipelines, Pipeline steps,
Standalone scheduled, Manual only. An unrecognised category still renders, under
"Other" -- a stray section beats a job silently vanishing from the admin page.

Sections rather than nesting steps under their parent, which is what the flat
"runs via pipeline" label invited. Membership is many-to-many -- data_collector
runs in all four pipelines, alerts and outcome_evaluator in two each -- so
nesting means duplicating those rows, and the duplicates would each carry a
Trigger button despite not being distinct actions: only plain collect_ohlcv is
registered, while the near-close and after-close variants are different
coroutines that are not individually triggerable. Instead each pipeline card
lists its step sequence and each step says which pipelines run it, which is the
same information without a button that lies.

Every job now answers "when does this next run" the same way: its own timer, its
soonest enabled parent's ("Next via Intraday Pipeline in 42m"), or "manual
only". Jobs with no recorded run say so explicitly rather than showing nothing.

The status chip and the rate-limit banner still read runtime_* only, so a
persisted failure cannot pin either to a stale state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 12:25:05 +02:00
dennisthiessenandClaude Opus 5 22bee28ac7 feat(jobs): persist each job's last run so it survives a restart
Job outcomes lived only in scheduler._job_runtime, an in-memory dict. Every
deploy wiped it, so Admin -> Jobs could report "Active" with no indication a job
had ever run or how it ended -- which is the main thing that page is for.

New job_run_state table (migration 031): one row per job, upserted on job_name.
Deliberately not history -- system_events already grows unbounded with no
retention job, and a second append-only operational table would repeat that
debt. Adding history later is purely additive.

Written from two hooks, NOT from _runtime_finish. That looked cheapest (one
function, ~40 call sites) but unit tests invoke job coroutines directly, so it
would fire detached DB writes at the real session factory throughout the suite,
and there is no testing flag to guard on.

  - An APScheduler EVENT_JOB_EXECUTED/ERROR listener covers everything the
    scheduler fires, including manual triggers. Its detached task is held in a
    module-level set (a bare create_task result can be collected mid-flight) and
    drained in the app lifespan before engine.dispose().
  - _run_pipeline persists directly, and must: pipeline steps are plain
    coroutine calls that emit no scheduler events, so the listener cannot see
    them. The step persist sits AFTER the except that swallows step errors --
    inside it, exactly the failed runs worth seeing would be skipped. The
    orchestrator persists in the finally, and the disabled early-return persists
    too, or "skipped" is silently dropped.

_persist_job_run never raises: a persistence failure must not break an otherwise
successful pipeline.

The API reports this as last_run_* and leaves runtime_* meaning strictly live
in-memory state. Reusing runtime_status would have been a regression, not a
no-op: JobControls drives the status chip from it (a job that errored eight days
ago would read "Last run error" forever instead of "Active") and picks the
rate-limit banner from it (a week-old rate limit would pin the banner
permanently). Tests pin the split.

The table starts empty; each job fills its row the next time it finishes. No
backfill from system_events, which records only warning/error outcomes under a
different status vocabulary and would invent successes that never happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 12:19:29 +02:00
dennisthiessenandClaude Opus 5 083c9dbf7c refactor(jobs): derive job topology from one catalog, make next-run coherent
Groundwork for the Admin -> Jobs cleanup. Three sources of truth collapse into
app/job_catalog.py, which imports nothing from app so both the scheduler and
admin_service can import it at module level (admin_service otherwise has to
import the scheduler inside functions to dodge a cycle).

PIPELINE_MEMBERS is now DERIVED from the four pipeline step lists instead of
being a literal set in admin_service duplicating four lists in scheduler.py with
nothing asserting they agreed. A test pins that the derivation reproduces the
previous hand-maintained 9 names exactly, so this is behaviour-preserving.

Deletes the private _JOB_NAMES list, which held 16 of the 19 jobs:
benchmark_collector, outcome_evaluator and shadow_book had no runtime row, and
so no "last run" line in the panel, until their first run in a given process.
_job_runtime is now seeded from the catalog, and a test pins the invariant.

Next-run is decided by category rather than by reading a timestamp. A pipeline
step has no schedule of its own, so it reports its parent's ("next via Morning
Pipeline in 3h") instead of nothing; a manual job says manual_only rather than
rendering a date. This also fixes a real bug: triggering a paused job set
next_run_time=now, APScheduler re-armed the 520-week backstop behind it, and the
panel displayed "next run in ~87600h". Two independent guards -- the category
rule, plus _visible_next_run dropping anything past a year -- and an APScheduler
listener that re-pauses steps and manual jobs once their run finishes. The
listener is registered at module level because configure_scheduler is called
more than once and add_listener does not deduplicate.

Migrates backtest and ticker_universe_sync from interval to cron (Sun 03:00 ET
and 01:00 ET). configure_scheduler calls remove_all_jobs() on every startup, so
an interval countdown restarts each deploy -- a 168h backtest needed a week of
uninterrupted uptime to fire even once. The codebase already documented this
pitfall as the reason cron was adopted; these two were never migrated. Both are
now editable in Admin -> Schedule.

Also: list_jobs went from one settings query per job (19) to one for all of
them, and data_backfill is hidden from the listing while staying registered and
API-triggerable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 12:08:01 +02:00
84 changed files with 12392 additions and 1942 deletions
+3
View File
@@ -54,3 +54,6 @@ reports/.cache/
# Runtime A5 parity bundles are generated on the production server. Research
# conclusions belong in docs/research, not as an ever-growing artifact archive.
reports/fundamentals-parity/
# Calibration harness raw-pull cache (Alpaca/FRED); regenerable, not a record.
.calib-cache/
+141 -39
View File
@@ -2,7 +2,7 @@
Investing-signal platform for US equities. It runs one strategy, and it is a boring one:
> **A long-only cross-sectional momentum book.** Buy the top quintile by beta-adjusted 12-1 month momentum, tilt toward higher volatility, hold at most 10 names, cut at 1.5× ATR, then trail at 3× ATR for up to 30 trading days. After an initial-stop exit, re-enter only after the gate has failed and subsequently qualified again.
> **A long-only cross-sectional momentum book.** Buy the top quintile by beta-adjusted 12-1 month momentum, tilt toward higher volatility, hold at most 15 names, cut at 1.5× ATR, then trail at 3× ATR for up to 30 trading days. After an initial-stop exit, re-enter only after the gate has failed and subsequently qualified again.
**Philosophy:** don't predict price — rank it. The edge is *relative* strength across the universe, and the discipline is in the exit: cut losers fast, let winners run until the trail catches them.
@@ -31,7 +31,7 @@ flowchart TD
Q -->|no| SKIP
Q -->|yes| RANK["Rank by production score<br/>80% momentum %ile<br/>+ 20% volatility %ile"]
RANK --> BOOK{"Room in the book?<br/>max 10 positions"}
RANK --> BOOK{"Room in the book?<br/>max 15 positions"}
BOOK -->|no| WAIT["Wait for a slot"]
BOOK -->|yes| OPEN["OPEN — size at 1% account risk"]
@@ -131,16 +131,17 @@ indicators.
**Morning** (~02:00 ET) — data and display only, **no** qualifying R:R scan:
1. **OHLCV** — latest daily bars (Alpaca); new tickers backfill ~5 years.
1. **OHLCV** — latest daily bars (Alpaca) plus the SPY benchmark; new tickers backfill ~5 years. A symbol whose bars have been stale for 3 days is probed against SEC for a Form 25/25-NSE/15 and **retired** on a hit (history kept — see *Delisting*).
2. **Sentiment** — stale names that matter (top-pick feeders, watchlist, open paper, discovery net). Display context only; the activation gate is price-only.
3. **Market Trend (SPY)** + **AI/Tech Risk Monitor** — the SPY trend guard and the v3 risk thermometer; feed no trades.
3. **Market Trend (SPY)** + **AI/Tech Risk Monitor** — the SPY trend guard and the v4 risk thermometer; feed no trades.
4. **Telegram alerts** — change-driven (risk-quadrant etc.); quiet days stay quiet. Setup alerts still fire on the near-close pipeline after the scan.
**Near-close** (~15:30 ET MonFri) — the only full-universe qualifying observation:
1. **OHLCV fetch** — refresh the in-progress day-t bar (same path as intraday).
2. **R:R Scan** — Structural S/R, scores, Gate Target Ladder setups, residual 121 + 80/20 rank. Advances post-stop gate-reset transitions; failed scans never count.
3. **Telegram alerts** — chained immediately so manual MOC fills can still hit ~15:50/15:55.
3. **Shadow book** — opt-in automated book; opens top-ranked qualified setups up to capacity at the same near-close prices. Only accepts a scan from this same pipeline run.
4. **Telegram alerts** — chained immediately so manual MOC fills can still hit ~15:50/15:55.
**After close** (~16:45 ET MonFri):
@@ -157,6 +158,38 @@ Hourly mid-session (MonFri ~10:0015:00 ET): only **OHLCV → Outcome Eval*
Dolt earnings import (daily 02:30 ET) · SEC fundamentals import (daily 04:00 ET, also refreshes the fundamentals cache scoring reads) · Backtest (weekly) · Ticker-universe sync (daily). Alerts auto-fire only via the near-close pipeline (still manually triggerable). Deep history backfill and event study are manual-only (Admin → Jobs).
The SEC import defers a run rather than writing partial data when a filing's XBRL
hasn't landed. Two bounds keep that from compounding: `MISSING_XBRL_RETRY_DAYS`
caps how long *one* filing blocks promotion, and `PROMOTION_CEILING_DAYS` (7)
caps how long the import as a whole can stay deferred — past the ceiling every
unresolved filing is aged out in place so `promote()` queues it as a gap row,
`source_max_date` advances, and the import self-heals. A `deferred_stale` alert
inside that window is normal and clears on its own; check `source_max_date` in
`data_import_runs` before diagnosing a wedge.
### Delisting, not deletion
Retiring a symbol used to mean `delete_ticker` or a pruning universe bootstrap,
both of which cascade through OHLCV, setups and scores. That destroys exactly the
history four research documents apologise for: today's tracked universe projected
backward is survivorship-biased, and hard-deleting every delisted name is what
causes it. Keeping the rows preserves the option to fix that later (it does not
fix it — the replay still has to model a delisting as an exit event).
`tickers` therefore carries `delisted_on` / `delisted_reason` (migration 032);
`NULL` means actively traded. The filter is **opt-in** via
`ticker_service.active_only`, applied to the live path only — scanner, momentum
ranking, scoring, breadth, fundamentals candidates, SEC universe, earnings import,
ingestion. The registry and admin views deliberately keep delisted rows visible,
and `run_backtest` keeps them on purpose. Detection runs off OHLCV staleness
(not the SEC fundamentals import, which stalls for days on unrelated Company-Facts
gaps) and retires only on a Form 25/25-NSE/15 hit, so a halt or a rename keeps the
existing warning instead. `delisted_on` is the *effective* date — Rule 12d2-2
makes a Form 25 removal take effect ten days after filing, so a symbol filed today
keeps trading (and keeps qualifying) until that date. It is safe to automate
because it is reversible: `clear_delisted` un-retires a false positive, where a
delete had already taken the history.
### From score to "top pick"
1. **Composite score** — technical, S/R-quality, sentiment, fundamental and momentum sub-scores (0100) combine into a weighted composite (weights configurable; missing dimensions re-normalize). **Display and ranking only — it does not select trades.**
@@ -166,6 +199,33 @@ Dolt earnings import (daily 02:30 ET) · SEC fundamentals import (daily 04:00 ET
**What the R:R and reach-probability in step 3 actually are.** They are *gate inputs*, computed from a Gate Target Ladder proposal the trade will never exit at — they exist to filter setups, not to forecast the trade you're about to take. A setup with "R:R 2.4:1, 34% reach probability" is not a claim that you'll make 2.4R with 34% probability; it's a claim that this setup cleared the screen. What actually happens to a trade is in the exit box of the diagram above, and on the "what usually happens" panel in the UI. Conflating the two is the single easiest way to misread this app.
### Two books: shadow (automated) and discretionary (manual)
The platform keeps **two** paper books, and the difference between them is the
whole point.
| Book | Who selects | What it measures |
|---|---|---|
| **Shadow book** (`app/services/shadow_book_service.py`) | The machine — top-ranked qualified setups up to capacity, every near-close scan | The **strategy**, faithfully |
| **Discretionary book** | You, by clicking "paper trade" on a setup | The strategy **plus** your discretion and availability |
The manual book only ever contains trades the user chose to take, inside a ~20
minute window, on days they were around. The backtest that validated this
strategy does none of that, which makes the manual record unusable on its own as
out-of-sample evidence. The shadow book closes that gap: it mirrors
`_simulate_portfolio`'s selection rule exactly, orders on the *stored*
`strategy_rank` the scanner already wrote (so the two cannot drift apart) and
shares the manual book's exit policy — the only difference between the books is
*which* qualified setups get taken.
It runs as a step of the near-close pipeline, straight after the scan so entries
mark at the same near-close prices, and it only accepts a scan from the same
pipeline run. It is **opt-in** (`shadow_book_enabled`, with capacity, risk % and
starting equity under **Admin → Settings → Performance & Shadow Book**) because it
writes live trades. The **Dashboard**'s performance chart plots shadow vs
discretionary vs SPY; *Signals → Paper Trades* still shows the discretionary book
only.
## Strategy Status — What's Validated and What Isn't
**Read this before touching scoring, gating, or setup logic.** The platform measures itself — a weekly-replay backtest plus a factor rank-IC harness (`app/services/backtest_service.py`) — and the verdicts below come from those reports (latest run July 2026, ~5 years of OHLCV), not from opinion.
@@ -176,7 +236,8 @@ Dolt earnings import (daily 02:30 ET) · SEC fundamentals import (daily 04:00 ET
|---|---|---|
| **Residual 12-1 cross-sectional momentum** (the activation gate, long-only) | **Production gate — in-sample edge** | Promoted July 2026 after the portfolio variant beat raw 80 on CAGR, Sharpe and drawdown. Raw 12-1 remains a fallback only when benchmark data is unavailable |
| **3× ATR trailing exit** (+ 1.5× ATR initial stop, 30-day max hold) | **Production exit — best Sharpe of every exit tested** | Beat hold / SMA50 / 20-day-low / technical-40 and both take-profit variants (July 2026) |
| **Post-stop gate reset** | **Production re-entry policy** | The initial stop always closes; the ticker must later fail the daily gate and subsequently qualify again. At the production capacity of 10: Sharpe 1.67 → 1.77, CAGR 45.2% → 48.3%, DD 24.3% → 21.6% versus immediate re-entry. [Full study](docs/research/post-stop-reentry.md) |
| **Post-stop gate reset** | **Production re-entry policy** | The initial stop always closes; the ticker must later fail the daily gate and subsequently qualify again. At the then-production capacity of 10: Sharpe 1.67 → 1.77, CAGR 45.2% → 48.3%, DD 24.3% → 21.6% versus immediate re-entry. Capacity has since been raised to 15 — see the open question under the re-entry section. [Full study](docs/research/post-stop-reentry.md) |
| **Book capacity 15** (raised from 10, 2026-08-05) | **Production sizing** | The focused daily capacity bracket found the count cap was binding and cost real compounding: +1.075pp CAGR paired, 51 paths better / 2 worse, drawdown unchanged. Cash plus the 20% notional cap saturates the book near 12, so the cap no longer binds. [Findings](docs/research/portfolio-capacity-bracket-findings.md#correction-2026-08-05-ev-per-trade-was-the-wrong-lens) |
| **Structural S/R** | **Human-facing context only — not a gate and not an exit** | Clean, capped zones are persisted for charts and alerts. The scanner deliberately does not read them. |
| **Gate Target Ladder** | **Gate input only — not market structure and not an exit** | Volume-free range grid + pivots preserves the useful legacy screening behavior exactly: 1,086/1,086 qualified setups retained and identical Sharpe 2.03 / CAGR 50.0% / DD 21.4% / 321 trades. The exit never reads its target. [Full write-up](docs/research/sr-levels-and-exits.md#explicit-gate-target-ladder) |
| Composite score + 5 dimensions | **Display/ranking only** | Sub-scores are hand-built heuristics; none has a measured IC. Note: the "momentum" *dimension* is 5/20-day ROC — NOT the validated 12-1 factor (that lives in `momentum_service`) |
@@ -187,7 +248,7 @@ Dolt earnings import (daily 02:30 ET) · SEC fundamentals import (daily 04:00 ET
| Gate target as a take-profit (tested July 2026) | **Rejected** | Sharpe 2.04 → 1.47, CAGR halved. Win rate *rose* — it truncates the right tail where the edge lives |
| "Clear-air" gate relaxation (tested July 2026) | **Rejected — failed out-of-sample** | Strictly better in-sample (Sharpe 2.07 / CAGR 62.3% / DD 20.1%), then lost on a real train/test split (Sharpe 2.78 → 2.45). A cautionary tale: nested lookbacks are not OOS |
Caveats on the momentum result: in-sample, roughly one market regime, costs/slippage approximated at 0.1% per side, and residual momentum still needs SPY benchmark history to compute. The **out-of-sample proof is the forward paper-trade record**: Signals → Track Record compares live qualified expectancy against the backtest.
Caveats on the momentum result: in-sample, roughly one market regime, costs/slippage approximated at 0.1% per side, and residual momentum still needs SPY benchmark history to compute. The **out-of-sample proof is the forward record of the shadow book** — the automated twin that takes every top-ranked qualified setup, with no discretion or availability mixed in. The Dashboard chart tracks it against the discretionary book and SPY; *Signals → Backtest* is what it is being compared against.
### Daily post-stop re-entry decision (2026-07-17)
@@ -200,7 +261,9 @@ The production policy is **normal gate reset**, evaluated with daily setup oppor
| Strict gate reset (live timing analogue) | 342.7% | 44.8% | -23.4% | 1.68 | 471 |
| Fixed five-session cooldown | 250.8% | 36.6% | -22.2% | 1.47 | 473 |
In the disjoint 2025+ book, gate reset also beat immediate re-entry (Sharpe 1.66 vs 1.55; CAGR 41.8% vs 39.3%) and the fixed five-session rule (Sharpe 1.43; CAGR 32.7%). Its lead over both survived costs of 0.2% and 0.3% per side. The result is capacity-specific: cooldown 5 won at capacity 5, while immediate had slightly higher return and Sharpe at capacity 15. Production uses capacity 10, so that is the portfolio for which this decision is valid.
In the disjoint 2025+ book, gate reset also beat immediate re-entry (Sharpe 1.66 vs 1.55; CAGR 41.8% vs 39.3%) and the fixed five-session rule (Sharpe 1.43; CAGR 32.7%). Its lead over both survived costs of 0.2% and 0.3% per side. The result is capacity-specific: cooldown 5 won at capacity 5, while immediate had slightly higher return and Sharpe at capacity 15.
> **Open question (since 2026-08-05).** This study was run — and gate reset promoted — at capacity 10. Production capacity was subsequently raised to 15, which is the one capacity in the matrix where *immediate* re-entry edged ahead. The re-entry policy is therefore currently running outside the portfolio it was validated on. Nothing else changed, and the two arms differed only modestly, but the matrix should be rerun at capacity 15 before treating gate reset as settled. Until then, keep gate reset (the incumbent) rather than switching on an untested read.
Those promotion numbers belong to the selected normal-reset study arm. Under the **pre-cutover** morning-scan scheduler (scan always before any outcome eval), live first-observation timing matched the stricter `strict_gate_reset` analogue (full-period Sharpe 1.68 / CAGR 44.8% / DD 23.4%). After the **near-close cutover** (2026-07), stops closed by earlier same-day intraday evals can receive a same-day fail observation at ~15:30 ET — moving live behavior **toward** the promoted `gate_reset` arm. Requalification still requires a later America/New_York trading date than the failure (`trade_policy` distinct-day guard). Full definitions and all nine policy arms: [docs/research/post-stop-reentry.md](docs/research/post-stop-reentry.md); execution evidence: [docs/research/execution-recovery.md](docs/research/execution-recovery.md).
@@ -208,7 +271,7 @@ Those promotion numbers belong to the selected normal-reset study arm. Under the
### Historical weekly production baseline (pre gate-reset)
Use this as the historical ranking/exit regression guardrail, not as a return promise or the current re-entry-policy result. This run predates the post-stop gate reset and uses weekly entry replay, so its portfolio headline is not directly comparable with the daily matrix above. Backtest run: local production SQLite snapshot, 506 tickers, weekly cadence, 30-trading-day horizon, 2022-06-28 → 2026-07-02, 0.1% per-side costs, price-only SPY benchmark. Numbers below are the 2026-07-11 run (`reports/backtest-20260711-prod-baseline.json`) — measured *after* the primary-target probability floor shipped, which pruned lottery-target setups (1,428 → 1,089 qualified) and lifted Sharpe on all three promotion contenders.
Use this as the historical ranking/exit regression guardrail, not as a return promise or the current re-entry-policy result. This run predates the post-stop gate reset **and the 2026-08-05 capacity raise to 15**, and uses weekly entry replay, so its portfolio headline is not directly comparable with the daily matrix above. Backtest run: local production SQLite snapshot, 506 tickers, weekly cadence, 30-trading-day horizon, 2022-06-28 → 2026-07-02, 0.1% per-side costs, price-only SPY benchmark. Numbers below are the 2026-07-11 run (`reports/backtest-20260711-prod-baseline.json`) — measured *after* the primary-target probability floor shipped, which pruned lottery-target setups (1,428 → 1,089 qualified) and lifted Sharpe on all three promotion contenders.
| Item | Historical weekly baseline |
|---|---|
@@ -248,16 +311,16 @@ Parity guard (July 2026): the portfolio monitor's **Production** row replays the
### Tuned and confirmed — do not retest without new data (July 2026)
A systematic single-variable sweep (offline prod snapshot, production gate/rank/exit, 2022-06 → 2026-07 plus disjoint 202223 / 202426 folds) confirmed **every** production setting. Retesting these against the same ~4-year snapshot is wasted compute and invites overfitting; revisit only with meaningfully new data (longer history or broader universe).
A systematic single-variable sweep (offline prod snapshot, production gate/rank/exit, 2022-06 → 2026-07 plus disjoint 202223 / 202426 folds) confirmed every production setting **except book size**, which a later focused bracket reversed (see the row below). Retesting these against the same ~4-year snapshot is wasted compute and invites overfitting; revisit only with meaningfully new data (longer history or broader universe) — or, as with capacity, a demonstrably better measurement lens.
| Knob tested | Verdict | Evidence |
|---|---|---|
| ATR trail multiple {1.54.0} | **Keep 3.0** | Return+Sharpe peak; ≤2.0 whipsaws out the momentum right tail; ≥2.5 is a plateau |
| SPY 200d-MA regime overlay (block entries / go flat) | **Reject** | Halves return (315%→138%) with zero drawdown benefit — the ATR trail already manages downside, and the filter blocks the recovery-phase entries that make the money |
| Momentum lookback: 6-1, 3-1, 12-7 (Novy-Marx), composites | **Keep residual 12-1** | 6-1/3-1 rank-IC ≈ 0; 12-7 IC 0.045 / t 1.58 — weaker than residual 12-1 (0.055 / t 1.98) |
| Selection cutoff {70, 75, 85, 90} × book size {10, 15, 20} | **Keep 80 × 10** | Monotonically worse in both directions from 80; the 10-slot cap never binds (<10 concurrent) |
| Selection cutoff {70, 75, 85, 90} × book size {10, 15, 20} | **Keep cutoff 80; book size raised to 15 (2026-08-05)** | The cutoff is monotonically worse in both directions from 80. The book-size half of this row was **reversed**: the weekly replay's "the 10-slot cap never binds" read came from EV per trade, which is the wrong lens for anything that changes trade *count*. The focused daily bracket found cap 10 *was* binding and cost +1.075pp CAGR; at 15 the cap never bound in any cell (max observed 12 concurrent, zero full-book skips) |
| Position sizing: equal-weight, inverse-vol, risk-% sweep | **Keep 1% fixed-fractional** | See the inverse-vol warning below |
| Post-stop re-entry: immediate, fixed 25 sessions, gate resets, confirmation filters | **Keep normal gate reset for the 10-position production book** | Sharpe 1.77 vs 1.67 immediate and 1.47 cooldown 5; rerun before changing portfolio capacity |
| Post-stop re-entry: immediate, fixed 25 sessions, gate resets, confirmation filters | **Keep normal gate reset** — but measured at capacity 10, and capacity is now 15 | Sharpe 1.77 vs 1.67 immediate and 1.47 cooldown 5. The "rerun before changing portfolio capacity" caveat is now outstanding — see the open question above |
| FIP path-smoothness as an in-book tie-breaker/filter | **Reject** (but see the lead below) | Non-monotonic across FIP quintiles within the qualified set; either half of a median split underperforms the full book — thinning the entry stream costs more compounding than the tilt returns |
Two findings future sessions must not re-litigate:
@@ -270,23 +333,24 @@ Two findings future sessions must not re-litigate:
A signal earns its way into selection **only** through the factor harness:
1. Add it as a point-in-time function of past bars in `_signal_values()` (`backtest_service.py`).
2. Run the backtest (Admin → Jobs, or the weekly run) and read the **Signal edge** table (Signals → Track Record).
2. Run the backtest (Admin → Jobs, or the weekly run) and read the report's `signal_eval` section. This one is **local-report only** — the deployed Backtest tab does not render it (see *Reading a local backtest report* below).
3. Wire it into the gate or ranking **only if** |mean IC| ≳ 0.03 with a consistent sign and `reliable: true` (≥ 12 non-overlapping windows).
Corollaries: never let an unvalidated score gate setups; the outcome evaluator must keep scoring **all** setups (unqualified ones are the control group); LLM output stays display-only in the quant path.
### Highest-value next experiments (in order)
> Check **[docs/research/](docs/research/README.md)** first — 12 strategy ideas have already been tested and rejected, including the obvious ones (take-profit exits, regime overlays, inverse-vol sizing, shorts).
> Check **[docs/research/](docs/research/README.md)** first — 13 strategy ideas have already been tested and rejected, including the obvious ones (take-profit exits, regime overlays, inverse-vol sizing, shorts, sector-residual momentum).
1. **Forward monitor the promoted strategy**the production UI now behaves like a portfolio monitor for the current strategy, with selectable lookbacks and SPY comparison. Forward paper-trade months are the only evidence the snapshot cannot provide; the July 2026 tuning pass closed every in-sample lead. (Trailing-stop sensitivity and the max-15 capacity check are done — see the tuning table above.)
1. **Forward monitor the promoted strategy***Signals → Backtest* behaves like a portfolio monitor for the current strategy, with selectable lookbacks and SPY comparison, and the Dashboard chart carries the forward record. Forward months of the **shadow book** are the only evidence the snapshot cannot provide; the July 2026 tuning pass closed every in-sample lead. (Trailing-stop sensitivity and the capacity bracket are done — capacity was raised to 15.)
2. **Signal context snapshots** — accumulate point-in-time composite/sentiment/fundamental context for every new setup so the discretionary overlay can be tested forward-only.
3. **Breadth is no longer free leverage** — Phase B found residual-mom t-stat *fell* on liquid-1500 vs the 505-name fingerprint (0.055/1.98 → 0.029/1.33). Any breadth book must clear a pre-registered baseline arm before fip tilts mean anything. (Deeper history was considered and declined.)
## Key Use Cases
- **Find today's best long setup.** On the **Dashboard**, the *Top Setups* table lists residual-gated qualified setups ranked by the production 80/20 residual/high-vol score, with the #1 flagged "Top pick". Each row opens the ticker page for its chart, Structural S/R, Gate Target Ladder targets and entry/stop.
- **Track a trade you took.** Mark a setup as a **paper trade**: it's marked-to-market against the latest close, auto-closed by the active exit policy (default: 3x ATR trail with a 30-trading-day max hold), and its sentiment stays fresh while open. *Signals → Track Record* shows the realized edge.
- **Track a trade you took.** Mark a setup as a **paper trade**: it's marked-to-market against the latest close, auto-closed by the active exit policy (default: 3x ATR trail with a 30-trading-day max hold), and its sentiment stays fresh while open. *Signals → Paper Trades* shows the realized edge of your discretionary book; the Dashboard chart puts it next to the automated shadow book and SPY.
- **Ask whether the strategy is worth trading at all.** *Signals → Backtest* replays the promoted strategy over history — portfolio monitor vs SPY over selectable lookbacks, headline risk-adjusted metrics (Sharpe, Sortino, Gain-to-Pain, dollar profit factor) and the report's own recommendation — with the live-outcome evaluation panel underneath it.
## Stack
@@ -306,7 +370,7 @@ Corollaries: never let an unvalidated score gate setups; the outcome evaluator m
## Features
### Backend
- Ticker registry with full cascade delete
- Ticker registry with reversible delisting (history preserved) plus an explicit cascade delete
- Universe bootstrap for `sp500`, `nasdaq100`, `nasdaq_all` via admin endpoint — free public sources (Wikipedia / NASDAQ Trader), then the cached snapshot, then a built-in seed list. The seeds are representative, not complete, so a *fresh* install bootstrapped while the public source is unreachable gets a partial universe; a warm instance falls through to its cache.
- OHLCV price storage with upsert and validation
- Technical indicators: ADX, EMA, RSI, ATR, Volume Profile, Pivot Points, EMA Cross
@@ -319,6 +383,8 @@ Corollaries: never let an unvalidated score gate setups; the outcome evaluator m
- Activation gate — qualifies setups on a residual-momentum percentile floor (the actual selection), a headline gate-target R:R floor (prod: 2.0) and a 20% primary-target reach-probability floor (validated long-only edge)
- Recommendation layer — directional confidence, conflict detection, per-target reach-probability
- Paper trading — take a setup, mark-to-market vs. latest close, auto-close per the exit policy (default: 3x ATR trail with a 30-trading-day max hold; time / percent-trailing / target-stop selectable), realized track record + outcome evaluation
- Shadow book — opt-in automated twin of the backtest's selection rule (top-ranked qualified setups up to capacity, every near-close scan), sharing the manual book's exit policy; the honest forward out-of-sample record
- System events — structured job/import/data warnings with acknowledgement, surfaced in Admin and deduplicated for alerting
- Market-regime guard + observational State/Warning monitor (fixed-basket breadth, VIX, credit level + impulse) with a manual chronological correction study
- Telegram alerts (e.g. regime-quadrant changes)
- User-curated watchlist (cap: 20), enriched with composite score, R:R and S/R summary
@@ -337,7 +403,10 @@ Corollaries: never let an unvalidated score gate setups; the outcome evaluator m
- Ticker detail page: chart, scores, sentiment breakdown, fundamentals, technical indicators, S/R table
- Rankings table with configurable dimension weights
- Trade scanner showing detected R:R setups
- Admin page: user management, job status with live indicators, enable/disable toggles, data cleanup, system settings
- Backtest tab: portfolio monitor vs SPY over selectable lookbacks, headline risk-adjusted tiles (Sharpe, Sortino, Gain-to-Pain, dollar profit factor), the report's recommendation card, and a live-outcome evaluation panel
- Dashboard performance chart: cumulative shadow book vs discretionary book vs SPY since the configured start date
- Paper Trades tab: open/closed discretionary trades with realized R and P&L tiles
- Admin page: user management, job status with live indicators, enable/disable toggles, pipeline readiness, system-event log, ticker management, data cleanup, system settings
- Protected routes with JWT auth, admin-only sections
- Responsive layout with mobile navigation
- Toast notifications for async operations
@@ -348,14 +417,14 @@ Corollaries: never let an unvalidated score gate setups; the outcome evaluator m
|---|---|---|
| `/login` | Login | Public |
| `/register` | Register | Public (when enabled) |
| `/` | Dashboard — top setups, open trades, regime (default) | Authenticated |
| `/` | Dashboard — top setups, open trades, regime, shadow-vs-manual-vs-SPY performance chart (default) | Authenticated |
| `/market` | Market — watchlist + rankings tabs | Authenticated |
| `/signals` | Signals — scanner + track record tabs | Authenticated |
| `/signals` | Signals — Setups / Paper Trades / Backtest tabs | Authenticated |
| `/regime` | AI/Tech Risk Monitor | Authenticated |
| `/ticker/:symbol` | Ticker Detail | Authenticated |
| `/admin` | Admin Panel | Admin only |
Legacy routes redirect: `/watchlist``/market`, `/rankings``/market?tab=rankings`, `/scanner``/signals`, `/performance``/signals?tab=track`.
Legacy routes redirect: `/watchlist``/market`, `/rankings``/market?tab=rankings`, `/scanner``/signals`, `/performance``/signals?tab=track` (the Paper Trades tab — `track` stays its slug so the old link keeps working).
## API Endpoints
@@ -365,7 +434,7 @@ All under `/api/v1/`. Interactive docs at `/docs` (Swagger) and `/redoc`.
|---|---|
| Health | `GET /health` |
| Auth | `POST /auth/register`, `POST /auth/login` |
| Tickers | `POST /tickers`, `GET /tickers`, `DELETE /tickers/{symbol}` |
| Tickers | `POST /tickers`, `GET /tickers`, `DELETE /tickers/{symbol}`, `POST /tickers/{symbol}/delisting`, `DELETE /tickers/{symbol}/delisting` |
| OHLCV | `POST /ohlcv`, `GET /ohlcv/{symbol}` |
| Ingestion | `POST /ingestion/fetch/{symbol}` |
| Indicators | `GET /indicators/{symbol}/{type}`, `GET /indicators/{symbol}/ema-cross` |
@@ -375,11 +444,11 @@ All under `/api/v1/`. Interactive docs at `/docs` (Swagger) and `/redoc`.
| Fundamentals | `GET /fundamentals/{symbol}` |
| Scores | `GET /scores/{symbol}`, `GET /rankings`, `PUT /scores/weights` |
| Trades | `GET /trades`, `GET /trades/{symbol}`, `GET /trades/{symbol}/history`, `GET /trades/activation`, `GET /trades/performance` |
| Paper Trades | `GET /paper-trades`, `POST /paper-trades`, `POST /paper-trades/{id}/close` |
| Market / Regime | `GET /market/regime`, `GET /regime/monitor`, `GET/PUT /regime/config`, `GET /regime/history`, `GET /regime/event-study`, `GET/PUT /regime/fundamentals`, `GET /backtest/report` |
| Paper Trades | `GET /paper-trades`, `POST /paper-trades`, `POST /paper-trades/{id}/close`, `GET /paper-trades/equity-curve`, `GET /paper-trades/performance` (shadow vs manual vs SPY), `GET/PUT /paper-trades/exit-policy` |
| Market / Regime | `GET /market/regime`, `GET /regime/monitor`, `GET/PUT /regime/config`, `GET /regime/history`, `GET /regime/event-study`, `GET/PUT /regime/fundamentals`, `POST /regime/fundamentals/refresh`, `GET /backtest/report` |
| Jobs | `GET /jobs/running` |
| Watchlist | `GET /watchlist`, `POST /watchlist/{symbol}`, `DELETE /watchlist/{symbol}` |
| Admin | `GET /admin/users`, `POST /admin/users`, `PUT /admin/users/{id}/access`, `PUT /admin/users/{id}/password`, `PUT /admin/settings/registration`, `GET /admin/settings`, `PUT /admin/settings/{key}`, `GET/PUT /admin/settings/recommendations`, `GET/PUT /admin/settings/ticker-universe`, `POST /admin/tickers/bootstrap`, `POST /admin/data/cleanup`, `GET /admin/jobs`, `POST /admin/jobs/{name}/trigger`, `PUT /admin/jobs/{name}/toggle`, `GET /admin/pipeline/readiness` |
| Admin | `GET /admin/users`, `POST /admin/users`, `PUT /admin/users/{id}/access`, `PUT /admin/users/{id}/password`, `PUT /admin/settings/registration`, `GET /admin/settings`, `PUT /admin/settings/{key}`, `GET/PUT /admin/settings/{recommendations,activation,schedule,performance,shadow-book,sentiment,alerts,ticker-universe}`, `POST /admin/settings/{sentiment,alerts}/test`, `POST /admin/tickers/bootstrap`, `POST /admin/tickers/backfill-names`, `POST /admin/data/cleanup`, `POST /admin/track-record/reset`, `GET /admin/jobs`, `POST /admin/jobs/{name}/trigger`, `PUT /admin/jobs/{name}/toggle`, `GET /admin/pipeline/readiness`, `GET /admin/system-events`, `GET /admin/system-events/summary`, `POST /admin/system-events/acknowledge` |
## Development Setup
@@ -439,8 +508,8 @@ npm run preview # Preview the production build locally
# Backend tests (in-memory SQLite — no PostgreSQL needed)
pytest tests/ -v
# Frontend: there is no test suite — `npm test` calls vitest, which is not
# installed. The frontend check is the full TypeScript build:
# Frontend: there is no test suite and no `test` script at all. The frontend
# check is the full TypeScript build:
cd frontend
npm run build
```
@@ -524,10 +593,10 @@ the [full research record](docs/research/sr-levels-and-exits.md#gtl-tuning-matri
### Reading a local backtest report
The deployed **Signals → Track Record** page is deliberately trimmed to validation
(portfolio monitor vs SPY, realized paper trades) and how-to-trade. The
strategy-tuning tables that used to live there now live **only** in the local
report — inspect these `reports/backtest-<timestamp>.json` sections and produce the
The deployed **Signals → Backtest** tab is deliberately trimmed to validation
(portfolio monitor vs SPY, headline metrics, the report's recommendation, and the
live-outcome evaluation panel). The strategy-tuning tables that used to live there
now live **only** in the local report — inspect these `reports/backtest-<timestamp>.json` sections and produce the
matching decision. Every change still goes through the factor harness first (see
**The iron rule for strategy changes** above).
@@ -564,8 +633,11 @@ Research-only flags, all off by default (the default report is byte-identical to
| `BACKTEST_ATR_TARGET_FALLBACK=k` | Synthesizes a k×ATR target where S/R offers none |
| `BACKTEST_FALLBACK_CLEAR_AIR_ONLY=1` | Restricts that fallback to setups with genuinely no structure ahead |
`recommendation` is the one section surfaced on the deployed page ("What this
backtest recommends"); everything else in this table is intentionally local-only.
`portfolio_monitor` and `recommendation` are the sections surfaced on the deployed
Backtest tab (the monitor chart/tiles and "What this backtest recommends"; the
recommendation is rebuilt on read, so it always matches the lookback on screen and
flags one it was not computed on). Everything else in this table is intentionally
local-only.
## Environment Variables
@@ -583,6 +655,14 @@ Configure in `.env` (copy from `.env.example`):
| `OPENAI_API_KEY` | For sentiment (OpenAI path) | — | OpenAI API key |
| `OPENAI_MODEL` | No | `gpt-4o-mini` | OpenAI model name |
| `OPENAI_SENTIMENT_BATCH_SIZE` | No | `5` | Micro-batch size for sentiment collector |
| `DEEPSEEK_API_KEY` / `XAI_API_KEY` | For sentiment (those paths) | — | Alternative pluggable sentiment providers |
| `SEC_USER_AGENT` | **For fundamentals** | placeholder | SEC EDGAR requires a real `name (contact: email)` UA — the shipped default is a placeholder and SEC will throttle/refuse it |
| `SEC_REQUEST_SPACING_SECONDS` | No | `0.2` | Politeness delay between SEC requests |
| `SEC_MAX_RETRIES` / `SEC_REQUEST_TIMEOUT_SECONDS` | No | `4` / `30` | SEC client retry and timeout budget |
| `DOLT_BINARY` | For earnings import | `dolt` | Path to the `dolt` executable |
| `DOLT_DATA_DIR` / `DOLT_EARNINGS_SUBDIR` | No | `dolt-data` / `earnings` | Local Dolt clone location |
| `DOLT_MIN_FREE_DISK_GB` | No | `5.0` | Refuse to clone/pull below this free space |
| `DOLT_COMMAND_TIMEOUT_SECONDS` | No | `600` | Per-command Dolt timeout |
| `FRED_API_KEY` | Optional (risk monitor) | — | FRED key for the AI/Tech risk monitor (VIX, credit spreads) |
| `TELEGRAM_BOT_TOKEN` | Optional (alerts) | — | Telegram bot token for alerts (can also be set in Admin) |
| `TELEGRAM_CHAT_ID` | Optional (alerts) | — | Telegram chat id for alerts |
@@ -591,6 +671,9 @@ Configure in `.env` (copy from `.env.example`):
| `RR_SCAN_FREQUENCY` | No | `daily` | R:R scanner schedule |
| `DEFAULT_WATCHLIST_AUTO_SIZE` | No | `10` | Auto-watchlist size |
| `DEFAULT_RR_THRESHOLD` | No | `1.5` | Minimum R:R ratio for setups |
| `OHLCV_HISTORY_DAYS` | No | `1825` | Backfill depth for new tickers (~5 years) |
| `OUTCOME_EVALUATION_MAX_BARS` | No | `30` | Bars the outcome evaluator resolves a setup over |
| `BACKTEST_WORKERS` | No | `4` | Worker processes for the scheduled backtest |
| `DB_POOL_SIZE` | No | `5` | Database connection pool size |
| `LOG_LEVEL` | No | `INFO` | Logging level |
@@ -683,7 +766,9 @@ app/
├── exceptions.py # Exception hierarchy
├── middleware.py # Global error handler → JSON envelope
├── cache.py # LRU cache with per-ticker invalidation
├── ssl_bootstrap.py # TLS trust-store bootstrap for outbound calls
├── scheduler.py # APScheduler job definitions
├── job_catalog.py # Single source of truth for job names + pipeline step lists
├── models/ # SQLAlchemy ORM models
├── schemas/ # Pydantic request/response schemas
├── services/ # Business logic layer
@@ -703,9 +788,11 @@ frontend/
│ ├── admin/ # User table, job controls, settings, data cleanup
│ ├── auth/ # Protected route wrapper
│ ├── charts/ # Canvas candlestick chart
│ ├── dashboard/ # Top setups, open trades, shadow-vs-manual performance chart
│ ├── layout/ # App shell, sidebar, mobile nav
│ ├── rankings/ # Rankings table, weights form
│ ├── scanner/ # Trade table
│ ├── signals/ # Setups / Paper Trades / Backtest panels
│ ├── ticker/ # Sentiment panel, fundamentals, indicators, S/R overlay
│ ├── ui/ # Badge, toast, skeleton, score card, confirm dialog
│ └── watchlist/ # Watchlist table, add ticker form
@@ -716,16 +803,26 @@ frontend/
└── styles/ # Global CSS with glassmorphism classes
docs/
├── dolt-integration-plan.md # Design record for the Dolt/SEC fundamentals workstream
├── dolt-sec-a3-design.md
├── fundamentals-deployment.md
└── research/ # Experiment log: what was tested, the result, the decision
├── README.md # Overview — start here before proposing a strategy change
── sr-levels-and-exits.md
── sr-levels-and-exits.md
├── post-stop-reentry.md
├── portfolio-capacity-bracket*.md
├── execution-recovery.md
├── fip-breadth-ic.md
├── regime-monitor-v3.md / -v4.md
└── … # 16 documents total
reports/ # Committed backtest reports (JSON) + compare_reports.py
deploy/
├── nginx.conf # Reverse proxy + static file serving
├── setup_db.sh # Idempotent DB setup script
── stock-data-backend.service # systemd unit
── provision_fundamentals.sh # Server-side Dolt/SEC fundamentals provisioning
└── signalplatform.service # systemd unit
tests/
├── conftest.py # Fixtures, strategies, test DB
@@ -743,9 +840,11 @@ Context for whoever — human or AI — continues this work. The owner pushes st
- **Live scan and backtest share the same pure functions.** The backtest replays production logic through DB-free functions (`compute_technical_from_arrays`, `compute_momentum_from_closes`, `detect_sr_levels`, `detect_gate_target_ladder`, the recommendation helpers). New strategy logic must stay in pure functions consumed by both paths, or the backtest stops measuring what production actually does.
- **Keep the two price-level models separate.** `detect_sr_levels` produces persisted Structural S/R for charts and alerts. `detect_gate_target_ladder` produces transient screening proposals and must never be persisted or presented as market structure. The scanner must not read `SRLevel` rows for target generation.
- **The Gate Target Ladder target is a gate input, never an exit.** `_atr_trailing_close()` does not take it as a parameter, and it must stay that way — take-profit exits were tested and halve CAGR. Any UI or alert that implies the trade exits at the target is a bug ([research](docs/research/sr-levels-and-exits.md#explicit-gate-target-ladder)).
- **The outcome evaluator evaluates ALL setups**, not just qualified ones — unqualified setups are the control group that makes the Track Record meaningful.
- **The outcome evaluator evaluates ALL setups**, not just qualified ones — unqualified setups are the control group that makes the realized-outcome record meaningful.
- **`SystemSetting` access goes through `app/services/settings_store.py`** — don't query the model directly.
- **Time-series data gets a real table** (see `benchmark_prices`, `regime_snapshots`); `SystemSetting` JSON is only for config and cached reports.
- **The shadow book must stay parity-clean.** It orders on the *stored* `strategy_rank` the scanner wrote and mirrors `_simulate_portfolio`'s selection rule; it accepts only a scan from its own pipeline run. Recomputing its ranking, or letting it consume a stale/manual scan, turns the forward OOS record back into an approximation.
- **Delisted tickers are retired, never deleted.** Live paths opt into `ticker_service.active_only`; the registry, admin views and `run_backtest` deliberately still see them. Deleting a symbol takes the history that a survivorship-bias fix would need.
- **Discretionary overlay data is forward-only.** `signal_context_snapshots` captures composite/dimension/sentiment/fundamental context for new setups. Do not approximate historical sentiment/fundamental snapshots from today's data.
- Style: surgical changes, minimal new files; extend existing services rather than adding parallel ones.
@@ -762,11 +861,14 @@ Context for whoever — human or AI — continues this work. The owner pushes st
| Backtest + factor rank-IC harness ("Signal edge") | `app/services/backtest_service.py` |
| Outcome resolution (target/stop/expired/ambiguous) | `app/services/outcome_service.py` |
| Paper trades + time/trailing/target auto-exit | `app/services/paper_trade_service.py` |
| Shadow book (automated twin of the backtest's selection) | `app/services/shadow_book_service.py` |
| Re-entry locks / distinct-day guard / book identities | `app/services/trade_policy.py` |
| Ticker registry, delisting + `active_only` filter | `app/services/ticker_service.py` |
| Point-in-time setup context snapshots | `app/models/signal_context_snapshot.py` + `app/services/rr_scanner_service.py` |
| Structural S/R detection, Gate Target Ladder & zone clustering | `app/services/sr_service.py` |
| **Research log — what's been tested and rejected** | **`docs/research/`** |
| SPY benchmark for residual momentum + paper-trade alpha | `app/services/benchmark_service.py` |
| Pipelines & job registration | `app/scheduler.py` |
| Pipelines & job registration | `app/scheduler.py` (step lists and job names in `app/job_catalog.py`) |
### Verifying changes
@@ -775,7 +877,7 @@ pytest tests/ -q # backend; in-memory SQLite, no Postgres needed
cd frontend && npm run build # full tsc check — this IS the frontend "test"
```
- `npm test` in `frontend/` is dead (vitest isn't installed; there are no frontend test files). Use `npm run build`.
- There is no `npm test` in `frontend/` — no test script, no test files. `npm run build` (`tsc -b && vite build`) is the frontend check.
- Backend tests that exercise services which `commit()` need a plain session fixture, not the rolling-back `db_session` — copy the pattern in `tests/unit/test_rr_scanner_integration.py`.
- `ruff` reports ~11 pre-existing errors in old test files; those are not regressions.
@@ -792,6 +894,6 @@ Practical consequences:
### Roadmap (agreed June 2026)
1. **Forward paper-test the momentum book** — the out-of-sample proof the backtest can't give. Watch Signals → Track Record (live vs backtest).
1. **Forward paper-test the momentum book** — the out-of-sample proof the backtest can't give. Watch the Dashboard chart (shadow book vs discretionary vs SPY) against Signals → Backtest.
2. **Full IBKR integration** — read real positions, overlay entries/stops on charts, alert on holdings' score deterioration. (Paper trading, the lighter alternative, is done.)
3. Strategy experiments in the order listed under **Strategy Status** above — each one goes through the factor harness first.
+51
View File
@@ -0,0 +1,51 @@
"""Durable last-run state per scheduled job
Revision ID: 031
Revises: 030
Create Date: 2026-08-08 00:00:00.000000
Job run state lived only in an in-memory dict in ``app.scheduler``, so every
process restart wiped it. Admin → Jobs could then only report "Active" with no
indication of whether a job had ever run, or how it ended — which is exactly
the information an operator opens that page for.
One row per job, upserted on ``job_name``. Not history: ``system_events``
already grows unbounded with no retention job, and a second append-only
operational table would repeat that debt.
The table starts empty; each job populates its row the next time it finishes.
No backfill from ``system_events`` — that table only records warning/error
outcomes and uses a different status vocabulary, so seeding from it would
invent successful runs that never happened.
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
revision: str = "031"
down_revision: Union[str, None] = "030"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"job_run_state",
sa.Column("id", sa.Integer(), nullable=False),
sa.Column("job_name", sa.String(length=64), nullable=False),
sa.Column("status", sa.String(length=32), nullable=False),
sa.Column("started_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("finished_at", sa.DateTime(timezone=True), nullable=False),
sa.Column("processed", sa.Integer(), nullable=True),
sa.Column("total", sa.Integer(), nullable=True),
sa.Column("message", sa.Text(), nullable=True),
sa.Column("updated_at", sa.DateTime(timezone=True), nullable=False),
sa.PrimaryKeyConstraint("id"),
sa.UniqueConstraint("job_name", name="uq_job_run_state_job_name"),
)
def downgrade() -> None:
op.drop_table("job_run_state")
+49
View File
@@ -0,0 +1,49 @@
"""Record delisting on tickers instead of deleting them
Revision ID: 032
Revises: 031
Create Date: 2026-08-11 00:00:00.000000
Until now the only way to retire a symbol was ``delete_ticker`` (or
``bootstrap_universe(prune_missing=True)``), both of which cascade through
OHLCV, setups and scores. That destroys exactly the history four research
documents already apologise for: today's tracked universe projected backward
is survivorship-biased, and hard-deleting every delisted name is what causes
it. Keeping the rows preserves the option to fix that later — it does not fix
it by itself, which needs the replay to model a delisting as an exit event.
``delisted_on`` is the effective date (from SEC Form 25/25-NSE/15 where we can
confirm it, else the day it was marked); ``delisted_reason`` is a short code
for how we learned. NULL in both means actively traded — the live signal path
filters on that, while list and admin views keep showing the row so the
delisting is visible rather than silently absent.
Nullable and reversible by design: clearing ``delisted_on`` un-retires a
symbol, which is what makes automatic marking safe where a delete would not be.
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
revision: str = "032"
down_revision: Union[str, None] = "031"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.add_column("tickers", sa.Column("delisted_on", sa.Date(), nullable=True))
op.add_column(
"tickers", sa.Column("delisted_reason", sa.String(length=32), nullable=True)
)
# The live path filters "actively traded" on every universe scan; the index
# keeps that predicate cheap as delisted rows accumulate.
op.create_index("ix_tickers_delisted_on", "tickers", ["delisted_on"])
def downgrade() -> None:
op.drop_index("ix_tickers_delisted_on", table_name="tickers")
op.drop_column("tickers", "delisted_reason")
op.drop_column("tickers", "delisted_on")
@@ -0,0 +1,70 @@
"""Point-in-time history for the sourced fundamental observation
Revision ID: 033
Revises: 032
Create Date: 2026-08-12 00:00:00.000000
The hyperscaler capex / "good news, stock down" read lived in a single
``SystemSetting`` slot, so each refresh overwrote the last and no history
existed. The read is now a categorical channel reported alongside State and
Warning (never a term in either), and a channel with no history cannot be
replayed: a snapshot rebuild would record every historical session as if nothing
had ever been observed, and the event study could not measure the channel at all.
Keyed on ``effective_date`` (the session the observation becomes usable on,
normally the next weekday) rather than ``fetched_at``, because that is the gate
that stops a rebuild stamping today's reading onto historical rows.
The table starts empty. ``update_regime_monitor`` records the currently stored
observation on its next run, so a deployment does not lose the live reading —
but genuine history does not exist and cannot be invented here. Backfilling it
from the SEC capex line and earnings-date reactions is separate work; until then
every historical session reads ``unknown``, which is the honest value rather than
a guessed one.
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
revision: str = "033"
down_revision: Union[str, None] = "032"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"regime_fundamental_observations",
sa.Column("id", sa.Integer(), nullable=False),
sa.Column("effective_date", sa.Date(), nullable=False),
sa.Column("f1_score", sa.Float(), nullable=True),
sa.Column("f3_score", sa.Float(), nullable=True),
sa.Column("capex_json", sa.Text(), nullable=False),
sa.Column("good_news_stock_down", sa.String(length=10), nullable=False),
sa.Column("reasoning", sa.Text(), nullable=True),
sa.Column("source", sa.String(length=30), nullable=False),
sa.Column("fetched_at", sa.DateTime(timezone=True), nullable=False),
sa.Column("created_at", sa.DateTime(timezone=True), nullable=False),
sa.PrimaryKeyConstraint("id"),
)
# One unique index, not a unique constraint plus a plain index: the model
# declares `unique=True, index=True`, which SQLAlchemy renders as exactly
# this. The constraint-plus-index pairing worked but left a redundant second
# index on the column and a permanent metadata diff for autogenerate to keep
# trying to reconcile. Matches RegimeSnapshot.date, the sibling table.
op.create_index(
"ix_regime_fundamental_observations_effective_date",
"regime_fundamental_observations",
["effective_date"],
unique=True,
)
def downgrade() -> None:
op.drop_index(
"ix_regime_fundamental_observations_effective_date",
table_name="regime_fundamental_observations",
)
op.drop_table("regime_fundamental_observations")
@@ -0,0 +1,41 @@
"""Track when a filing gap stops pausing setups
Revision ID: 034
Revises: 033
Create Date: 2026-08-21 00:00:00.000000
An escalated gap stops pausing setups while the issuer's own fundamentals are
still recent (``GAP_GATE_RECENT_FILING_DAYS``). That reprieve is not permanent:
the stored filings age out, or a newer gap appears, and the pause returns —
silently, because ``filing_gap_aged`` only escalates gaps whose ``escalated_at``
is NULL and so never fires twice for the same gap.
``exempted_at`` is the state marker that makes the transition observable. It is
set (quietly) while the issuer is exempt and cleared when the exemption lapses,
which is when ``filing_gap_repaused`` fires — once per lapse, re-arming if the
issuer's data recovers and ages out again.
Nullable, and carrying no meaning of its own beyond that state: an existing gap
starts NULL and is stamped on the next import that finds it exempt.
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
revision: str = "034"
down_revision: Union[str, None] = "033"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.add_column(
"sec_filing_gaps",
sa.Column("exempted_at", sa.DateTime(timezone=True), nullable=True),
)
def downgrade() -> None:
op.drop_column("sec_filing_gaps", "exempted_at")
+207
View File
@@ -0,0 +1,207 @@
"""Job topology: names, labels, pipeline membership, categories, ordering.
The single source of truth for *what the jobs are*, as opposed to how they run.
It deliberately imports nothing from ``app`` so both ``app.scheduler`` and
``app.services.admin_service`` can import it at module level -- admin_service
otherwise has to do ``from app.scheduler import ...`` inside functions to dodge a
cycle.
The pipeline step lists live here rather than in the scheduler because three
separate things need them and used to keep private copies: the runner, the
``PIPELINE_MEMBERS`` set the admin API reports, and the UI's grouping. Steps are
``(step_name, coroutine_name)``; ``_run_pipeline`` resolves the coroutine late
out of the scheduler's own globals, so nothing here depends on those functions
existing.
"""
from __future__ import annotations
# ---------------------------------------------------------------------------
# Pipelines
# ---------------------------------------------------------------------------
_DAILY_PIPELINE_STEPS = [
("data_collector", "collect_ohlcv"),
("benchmark_collector", "collect_benchmark"),
("sentiment_collector", "collect_sentiment"),
("market_regime", "compute_market_regime"),
# Observational only — display/alerts; not trade selection.
("regime_monitor", "compute_regime_monitor"),
# Alerts after regime so quadrant changes reach Telegram in the morning.
# Dispatcher is change-driven; quiet days stay quiet. Setup alerts still
# fire on the near-close pipeline after the qualifying scan.
("alerts", "dispatch_alerts_job"),
]
# Near-close (~15:30 ET MonFri): refresh in-progress day-t bars (incremental
# ingestion overlaps the latest stored session), then the only daily
# qualifying R:R scan, then Telegram immediately so manual fills can still hit
# MOC cutoffs (~15:50/15:55). Under a 15-minute delayed SIP feed a 15:30 scan
# may see ~15:15 prices — immaterial for a 12-1 momentum signal.
#
# US early-close days (~3/year, 13:00 ET close): this job runs post-close and
# entries behave like stale_close (still acceptable per execution-recovery matrix).
# No exchange calendar dependency.
_NEAR_CLOSE_PIPELINE_STEPS = [
# Must land today's in-progress bar (~20 min behind live), or the scan falls
# back to the previous close and execution degrades to the stale_close floor.
("data_collector", "collect_ohlcv_for_scan"),
("rr_scanner", "scan_rr"),
# Straight after the scan so shadow entries mark at the same near-close
# prices the discretionary book is looking at.
("shadow_book", "run_shadow_book"),
("alerts", "dispatch_alerts_job"),
]
# After close (~16:45 ET MonFri): fresh OHLCV fetch so outcomes resolve on the
# final bar, not the near-close partial bar, then outcome/paper close.
_AFTER_CLOSE_PIPELINE_STEPS = [
("data_collector", "collect_ohlcv_final"),
("outcome_evaluator", "evaluate_outcomes"),
]
# Intraday (light): keep prices current and resolve outcomes through the day,
# without the expensive scan/sentiment. The dashboard recomputes live R:R from
# the latest price, so refreshing OHLCV is enough to stop prices lagging; the
# outcome step also closes paper trades that hit their stop/target intraday.
_INTRADAY_PIPELINE_STEPS = [
("data_collector", "collect_ohlcv"),
("outcome_evaluator", "evaluate_outcomes"),
]
# Ordered by trading day, not alphabetically: this is the sequence an operator
# reads down the page, and it drives the UI's ordering too.
PIPELINE_STEPS: dict[str, list[tuple[str, str]]] = {
"daily_pipeline": _DAILY_PIPELINE_STEPS,
"intraday_pipeline": _INTRADAY_PIPELINE_STEPS,
"near_close_pipeline": _NEAR_CLOSE_PIPELINE_STEPS,
"after_close_pipeline": _AFTER_CLOSE_PIPELINE_STEPS,
}
# Derived, never hand-maintained: this used to be a literal set in admin_service
# duplicating the four lists above from another module, with nothing asserting
# the two agreed.
PIPELINE_MEMBERS: frozenset[str] = frozenset(
step for steps in PIPELINE_STEPS.values() for step, _ in steps
)
def _pipelines_by_member() -> dict[str, tuple[str, ...]]:
"""Member -> the orchestrators that run it, in trading-day order.
Membership is many-to-many: data_collector runs in all four pipelines (via
three different coroutines), alerts and outcome_evaluator in two each.
"""
out: dict[str, list[str]] = {}
for pipeline, steps in PIPELINE_STEPS.items():
for step, _ in steps:
bucket = out.setdefault(step, [])
if pipeline not in bucket:
bucket.append(pipeline)
return {member: tuple(pipelines) for member, pipelines in out.items()}
PIPELINES_BY_MEMBER: dict[str, tuple[str, ...]] = _pipelines_by_member()
# ---------------------------------------------------------------------------
# Job identity
# ---------------------------------------------------------------------------
# Orchestrators, in trading-day order.
PIPELINE_JOBS: tuple[str, ...] = tuple(PIPELINE_STEPS)
# Own timer, independent of any pipeline.
SCHEDULED_JOBS: tuple[str, ...] = (
"dolt_earnings_import",
"sec_fundamentals_import",
"ticker_universe_sync",
"backtest",
)
# Registered but never auto-fired; run only when a human asks.
MANUAL_JOBS: tuple[str, ...] = ("event_study", "data_backfill")
# Steps in the order an operator meets them across the trading day, so the UI
# reads as a sequence rather than an alphabetical jumble.
PIPELINE_STEP_JOBS: tuple[str, ...] = tuple(
dict.fromkeys(step for steps in PIPELINE_STEPS.values() for step, _ in steps)
)
VALID_JOB_NAMES: frozenset[str] = frozenset(
PIPELINE_JOBS + PIPELINE_STEP_JOBS + SCHEDULED_JOBS + MANUAL_JOBS
)
JOB_LABELS: dict[str, str] = {
"data_collector": "Data Collector (OHLCV)",
"data_backfill": "Data Backfill (deep history)",
"benchmark_collector": "Benchmark Collector",
"sentiment_collector": "Sentiment Collector",
"dolt_earnings_import": "Dolt Earnings Import",
"sec_fundamentals_import": "SEC Fundamentals Import",
"rr_scanner": "R:R Scanner",
"ticker_universe_sync": "Ticker Universe Sync",
"outcome_evaluator": "Outcome Evaluator",
"alerts": "Alerts Dispatcher",
# Keys are persisted job ids and must not change; these are display only.
"market_regime": "Market Trend (SPY)",
"regime_monitor": "AI/Tech Risk Monitor",
"event_study": "Event Study",
"backtest": "Backtest",
"daily_pipeline": "Morning Pipeline",
"near_close_pipeline": "Near-Close Pipeline (scan+alert)",
"after_close_pipeline": "After-Close Pipeline (outcome)",
"intraday_pipeline": "Intraday Pipeline",
"shadow_book": "Shadow Book (auto-traded strategy)",
}
CATEGORY_PIPELINE = "pipeline"
CATEGORY_STEP = "pipeline_step"
CATEGORY_SCHEDULED = "scheduled"
CATEGORY_MANUAL = "manual"
# Order the sections appear in.
CATEGORY_ORDER: tuple[str, ...] = (
CATEGORY_PIPELINE,
CATEGORY_STEP,
CATEGORY_SCHEDULED,
CATEGORY_MANUAL,
)
CATEGORY_LABELS: dict[str, str] = {
CATEGORY_PIPELINE: "Pipelines",
CATEGORY_STEP: "Pipeline steps",
CATEGORY_SCHEDULED: "Standalone scheduled",
CATEGORY_MANUAL: "Manual only",
}
_CATEGORY_MEMBERS: dict[str, tuple[str, ...]] = {
CATEGORY_PIPELINE: PIPELINE_JOBS,
CATEGORY_STEP: PIPELINE_STEP_JOBS,
CATEGORY_SCHEDULED: SCHEDULED_JOBS,
CATEGORY_MANUAL: MANUAL_JOBS,
}
JOB_CATEGORY: dict[str, str] = {
name: category
for category, names in _CATEGORY_MEMBERS.items()
for name in names
}
# Registered and triggerable through the API, but kept out of Admin → Jobs.
# data_backfill's only capability beyond collect_ohlcv (which already backfills
# full history for *new* tickers) is re-deepening *existing* ones after
# ohlcv_history_days is raised -- a rare one-off, not something to scan past
# every time you open the page.
HIDDEN_JOBS: frozenset[str] = frozenset({"data_backfill"})
_SORT_INDEX: dict[str, tuple[int, int]] = {
name: (CATEGORY_ORDER.index(category), position)
for category, names in _CATEGORY_MEMBERS.items()
for position, name in enumerate(names)
}
def sort_order(job_name: str) -> tuple[int, int]:
"""(category rank, position within category). Unknown jobs sort last."""
return _SORT_INDEX.get(job_name, (len(CATEGORY_ORDER), 0))
+9 -1
View File
@@ -21,7 +21,12 @@ from app.config import settings
from app.database import async_session_factory, engine
from app.middleware import register_exception_handlers
from app.models.user import User
from app.scheduler import configure_scheduler, load_schedule_config, scheduler
from app.scheduler import (
configure_scheduler,
flush_job_run_persists,
load_schedule_config,
scheduler,
)
from app.routers.admin import router as admin_router
from app.routers.auth import router as auth_router
from app.routers.health import router as health_router
@@ -91,6 +96,9 @@ async def lifespan(_app: FastAPI) -> AsyncGenerator[None, None]:
scheduler.shutdown(wait=False)
logger.info("Scheduler stopped")
# Drain detached last-run writes before the engine goes away, or a job that
# finished during shutdown loses the row it just wrote.
await flush_job_run_persists()
await engine.dispose()
logger.info("Shutting down")
+4
View File
@@ -14,10 +14,12 @@ from app.models.settings import SystemSetting, IngestionProgress
from app.models.alert import AlertLog
from app.models.paper_trade import PaperTrade
from app.models.regime_snapshot import RegimeSnapshot
from app.models.regime_fundamental_observation import RegimeFundamentalObservation
from app.models.benchmark_price import BenchmarkPrice
from app.models.signal_context_snapshot import SignalContextSnapshot
from app.models.system_event import SystemEvent
from app.models.sec_filing_gap import SecFilingGap
from app.models.job_run_state import JobRunState
__all__ = [
"Ticker",
@@ -38,8 +40,10 @@ __all__ = [
"AlertLog",
"PaperTrade",
"RegimeSnapshot",
"RegimeFundamentalObservation",
"BenchmarkPrice",
"SignalContextSnapshot",
"SystemEvent",
"SecFilingGap",
"JobRunState",
]
+37
View File
@@ -0,0 +1,37 @@
from datetime import datetime
from sqlalchemy import DateTime, Integer, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from app.database import Base
class JobRunState(Base):
"""How each scheduled job last finished. One row per job, overwritten.
The scheduler's ``_job_runtime`` dict is the live view and is deliberately
in-memory, but it is also wiped by every process restart -- so after a deploy
Admin → Jobs could only say "Active" with no indication of whether a job had
ever run. This is the durable half.
Deliberately not history: ``system_events`` already grows without a reaper,
and a second append-only operational table would repeat that. Rows are
upserted on ``job_name``; adding history later is purely additive.
"""
__tablename__ = "job_run_state"
id: Mapped[int] = mapped_column(primary_key=True)
job_name: Mapped[str] = mapped_column(String(64), unique=True, nullable=False)
# Scheduler vocabulary: completed | skipped | error | rate_limited | deferred.
# Distinct from data_import_runs' statuses, which is one reason this is its
# own table rather than a widened column there.
status: Mapped[str] = mapped_column(String(32), nullable=False)
started_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
finished_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
processed: Mapped[int | None] = mapped_column(Integer, nullable=True)
total: Mapped[int | None] = mapped_column(Integer, nullable=True)
message: Mapped[str | None] = mapped_column(Text, nullable=True)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), default=datetime.utcnow, onupdate=datetime.utcnow, nullable=False
)
@@ -0,0 +1,44 @@
from datetime import date as date_type
from datetime import datetime
from sqlalchemy import Date, DateTime, Float, String, Text
from sqlalchemy.orm import Mapped, mapped_column
from app.database import Base
class RegimeFundamentalObservation(Base):
"""Point-in-time record of the sourced hyperscaler capex / earnings read.
One row per ``effective_date`` (unique, upserted). Before this table the
observation lived in a single ``SystemSetting`` slot, so every refresh
overwrote the previous one and no history existed at all — which made the
read impossible to replay, impossible to backtest, and meant a snapshot
rebuild could only ever score historical sessions as if nothing had been
observed.
The read is a categorical channel reported beside State and Warning, never a
term in either, so this series is not a scoring input. It is the record that
makes the channel replayable at all -- and the only route to eventually
testing whether it improves prediction conditional on Warning, which is the
one thing that could justify combining the channels later.
``effective_date`` rather than ``fetched_at`` is the key: it is the session
the observation becomes usable on (normally the next weekday), and the gate
that stops a rebuild stamping today's reading onto historical rows.
"""
__tablename__ = "regime_fundamental_observations"
id: Mapped[int] = mapped_column(primary_key=True)
effective_date: Mapped[date_type] = mapped_column(
Date, nullable=False, unique=True, index=True
)
f1_score: Mapped[float | None] = mapped_column(Float, nullable=True)
f3_score: Mapped[float | None] = mapped_column(Float, nullable=True)
capex_json: Mapped[str] = mapped_column(Text, nullable=False)
good_news_stock_down: Mapped[str] = mapped_column(String(10), nullable=False)
reasoning: Mapped[str | None] = mapped_column(Text, nullable=True)
source: Mapped[str] = mapped_column(String(30), nullable=False)
fetched_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
+5
View File
@@ -30,3 +30,8 @@ class SecFilingGap(Base):
first_seen_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
last_attempted_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
escalated_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
# Set while this gap's issuer is exempt from the setup pause (escalated, and
# its own fundamentals still recent — see fundamentals_quality_service).
# Cleared when the exemption lapses, which is the moment the pause silently
# comes back and the only moment worth alerting on.
exempted_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
+9 -2
View File
@@ -1,6 +1,6 @@
from datetime import datetime
from datetime import date, datetime
from sqlalchemy import String, DateTime
from sqlalchemy import Date, String, DateTime
from sqlalchemy.orm import Mapped, mapped_column, relationship
from app.database import Base
@@ -21,6 +21,13 @@ class Ticker(Base):
cik: Mapped[str | None] = mapped_column(String(10), nullable=True)
sic: Mapped[str | None] = mapped_column(String(4), nullable=True)
sic_description: Mapped[str | None] = mapped_column(String(160), nullable=True)
# Delisting is recorded, never deleted: the rows carry the price history that
# makes a backtest less survivorship-biased, and a delete cascades it away.
# NULL == actively traded. The live signal path filters on this (see
# ticker_service.active_only); list/admin views keep the row and show it.
delisted_on: Mapped[date | None] = mapped_column(Date, nullable=True, index=True)
# How we learned: "form_25" (SEC confirmed), "manual" (operator).
delisted_reason: Mapped[str | None] = mapped_column(String(32), nullable=True)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), default=datetime.utcnow, nullable=False
)
+33 -1
View File
@@ -6,7 +6,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.dependencies import get_db, require_access
from app.models.user import User
from app.schemas.common import APIEnvelope
from app.schemas.ticker import TickerCreate, TickerResponse
from app.schemas.ticker import TickerCreate, TickerDelistingUpdate, TickerResponse
from app.services import ticker_service
router = APIRouter(tags=["tickers"])
@@ -51,3 +51,35 @@ async def delete_ticker(
"""Delete a ticker and all associated data."""
await ticker_service.delete_ticker(db, symbol)
return APIEnvelope(status="success", data=None)
@router.post("/tickers/{symbol}/delisting", response_model=APIEnvelope)
async def mark_ticker_delisted(
symbol: str,
body: TickerDelistingUpdate,
_user: User = Depends(require_access),
db: AsyncSession = Depends(get_db),
):
"""Retire a symbol: excluded from signals, price history kept.
The non-destructive alternative to DELETE, which cascades the history away.
"""
changed = await ticker_service.mark_delisted(
db, symbol, delisted_on=body.delisted_on, reason=ticker_service.REASON_MANUAL
)
return APIEnvelope(status="success", data={"changed": changed})
@router.delete("/tickers/{symbol}/delisting", response_model=APIEnvelope)
async def clear_ticker_delisting(
symbol: str,
_user: User = Depends(require_access),
db: AsyncSession = Depends(get_db),
):
"""Un-retire a symbol wrongly marked delisted.
Automatic marking is only defensible because this exists: a false positive
costs one row update rather than the price history a delete would take.
"""
changed = await ticker_service.clear_delisted(db, symbol)
return APIEnvelope(status="success", data={"changed": changed})
+177 -76
View File
@@ -18,11 +18,13 @@ import logging
import asyncio
from datetime import date, datetime, timedelta, timezone
from apscheduler.events import EVENT_JOB_ERROR, EVENT_JOB_EXECUTED
from apscheduler.schedulers.asyncio import AsyncIOScheduler
from apscheduler.triggers.cron import CronTrigger
from sqlalchemy import and_, case, func, or_, select
from sqlalchemy.ext.asyncio import AsyncSession
from app import job_catalog
from app.config import settings
from app.database import async_session_factory
from app.models.ohlcv import OHLCVRecord
@@ -31,6 +33,7 @@ from app.models.ticker import Ticker
from app.exceptions import ProviderError
from app.providers.alpaca import AlpacaOHLCVProvider
from app.providers.protocol import SentimentData
from app.services import job_run_store
from app.services import (
ingestion_service,
pipeline_run,
@@ -63,6 +66,7 @@ from app.services.event_study_service import run_and_store as run_event_study_an
from app.services.outcome_service import evaluate_pending_setups
from app.services.rr_scanner_service import scan_all_tickers
from app.services.sentiment_provider_service import build_sentiment_provider
from app.services import ticker_service
from app.services.ticker_universe_service import bootstrap_universe
logger = logging.getLogger(__name__)
@@ -84,6 +88,47 @@ scheduler = AsyncIOScheduler(
}
)
def _on_job_finished(event: object) -> None:
"""Persist the run, then re-pause the job if it only runs on demand.
Covers every job APScheduler fires itself, including manual triggers.
Pipeline *steps* are invoked as plain coroutines and emit no events, so
``_run_pipeline`` persists those directly.
"""
job_id = getattr(event, "job_id", None)
if job_id:
_schedule_persist(job_id)
_repause_after_manual_run(event)
def _repause_after_manual_run(event: object) -> None:
"""Re-pause a job that only ever runs on demand, once its run finishes.
Pipeline steps and manual jobs are registered with a 520-week interval and
``next_run_time=None`` as a backstop. Triggering one sets next_run_time=now,
and APScheduler then re-arms that backstop -- so Admin → Jobs would show a
"next run" ten years out. Guarding on category means the six cron jobs and
the real interval jobs are never touched.
Registered at module level, not inside ``configure_scheduler``: that function
is called more than once (idempotency test) and ``add_listener`` does not
deduplicate.
"""
job_id = getattr(event, "job_id", None)
if job_catalog.JOB_CATEGORY.get(job_id) not in (
job_catalog.CATEGORY_STEP,
job_catalog.CATEGORY_MANUAL,
):
return
try:
scheduler.modify_job(job_id, next_run_time=None)
except Exception: # job gone, scheduler stopped — nothing to re-pause
logger.debug("Could not re-pause %s after its run", job_id, exc_info=True)
scheduler.add_listener(_on_job_finished, EVENT_JOB_EXECUTED | EVENT_JOB_ERROR)
# Track last successful ticker per job for rate-limit resume
_last_successful: dict[str, str | None] = {
"data_collector": None,
@@ -91,26 +136,10 @@ _last_successful: dict[str, str | None] = {
"sentiment_collector": None,
}
# Jobs whose per-run progress is surfaced to Admin → Jobs. (outcome_evaluator is
# created lazily on first run via _runtime_start.)
_JOB_NAMES = [
"data_collector",
"data_backfill",
"sentiment_collector",
"dolt_earnings_import",
"sec_fundamentals_import",
"rr_scanner",
"ticker_universe_sync",
"alerts",
"market_regime",
"regime_monitor",
"event_study",
"backtest",
"daily_pipeline", # morning: OHLCV/sentiment/regime — no qualifying scan
"near_close_pipeline", # OHLCV fetch → R:R scan → Telegram alerts
"after_close_pipeline", # OHLCV fetch → outcome eval (final bar)
"intraday_pipeline",
]
# Seeded from the catalog rather than a private list. The old literal held 16 of
# the 19 jobs -- benchmark_collector, outcome_evaluator and shadow_book were
# missing, so they had no runtime row (and so no "last run" line in Admin → Jobs)
# until their first run in a given process.
def _idle_runtime() -> dict[str, object]:
@@ -127,7 +156,9 @@ def _idle_runtime() -> dict[str, object]:
}
_job_runtime: dict[str, dict[str, object]] = {name: _idle_runtime() for name in _JOB_NAMES}
_job_runtime: dict[str, dict[str, object]] = {
name: _idle_runtime() for name in sorted(job_catalog.VALID_JOB_NAMES)
}
_next_backtest_target_model = PRODUCTION_GTL_TARGET_MODEL
_next_backtest_cadence = DEFAULT_BACKTEST_CADENCE
@@ -292,6 +323,67 @@ def _runtime_finish(
pass
async def _persist_job_run(job_name: str) -> None:
"""Write a job's finished runtime row to the durable last-run table.
Never raises: a persistence failure must not break the pipeline that was
otherwise successful. The in-memory row stays authoritative for live state.
"""
runtime = _job_runtime.get(job_name)
if not runtime or runtime.get("running") or not runtime.get("finished_at"):
return
try:
async with async_session_factory() as db:
await job_run_store.record_finish(db, job_name, runtime)
await db.commit()
except Exception:
logger.exception("Could not persist last-run state for %s", job_name)
# Detached persists are kept referenced: a bare create_task result can be
# garbage-collected mid-flight, and the shutdown drain needs something to await.
_persist_tasks: set[asyncio.Task] = set()
def _schedule_persist(job_name: str) -> None:
try:
task = asyncio.get_running_loop().create_task(_persist_job_run(job_name))
except RuntimeError: # no loop (sync context / tests) — nothing to persist
return
_persist_tasks.add(task)
task.add_done_callback(_persist_tasks.discard)
async def flush_job_run_persists(timeout: float = 5.0, settle: float = 0.05) -> None:
"""Drain last-run writes, including ones queued while we are draining.
``scheduler.shutdown(wait=False)`` returns before APScheduler has dispatched
its job-completion events, and those events are what create persist tasks. A
single snapshot of the set therefore misses writes still to be queued, and
``engine.dispose()`` could then close the pool underneath them. So: give the
loop a moment for pending callbacks to land, then keep draining until the
set stays empty or the deadline passes.
"""
loop = asyncio.get_running_loop()
deadline = loop.time() + timeout
# Bounded settle so callbacks dispatched by shutdown get to queue their work
# before the first emptiness check decides there is nothing to wait for.
await asyncio.sleep(min(settle, timeout))
while True:
pending = {task for task in _persist_tasks if not task.done()}
if not pending:
return
remaining = deadline - loop.time()
if remaining <= 0:
logger.warning(
"Timed out draining %d last-run write(s); some may be lost", len(pending)
)
return
await asyncio.wait(pending, timeout=remaining)
# Loop rather than return: a completion callback may have queued another.
await asyncio.sleep(0)
def get_job_runtime_snapshot(job_name: str | None = None) -> dict[str, dict[str, object]] | dict[str, object]:
if job_name is not None:
return dict(_job_runtime.get(job_name, {}))
@@ -305,8 +397,10 @@ async def _is_job_enabled(db: AsyncSession, job_name: str) -> bool:
async def _get_all_tickers(db: AsyncSession) -> list[str]:
"""Return all tracked ticker symbols sorted alphabetically."""
result = await db.execute(select(Ticker.symbol).order_by(Ticker.symbol))
"""Return all actively-traded ticker symbols sorted alphabetically."""
result = await db.execute(
ticker_service.active_only(select(Ticker.symbol).order_by(Ticker.symbol))
)
return list(result.scalars().all())
@@ -321,8 +415,10 @@ async def _get_ohlcv_priority_tickers(db: AsyncSession) -> list[str]:
latest_date = func.max(OHLCVRecord.date)
missing_first = case((latest_date.is_(None), 0), else_=1)
result = await db.execute(
ticker_service.active_only(
select(Ticker.symbol)
.outerjoin(OHLCVRecord, OHLCVRecord.ticker_id == Ticker.id)
)
.group_by(Ticker.id, Ticker.symbol)
.order_by(missing_first.asc(), latest_date.asc(), Ticker.symbol.asc())
)
@@ -571,6 +667,26 @@ async def collect_ohlcv(
_runtime_progress(job_name, processed=processed, total=total, current_ticker=symbol)
_log_event(logging.INFO, "ticker_collected", job=job_name, ticker=symbol, status=result.status, records=result.records_ingested)
if result.status == "stale":
# "No new bars" cannot distinguish a delisting from a halt
# or a rename, so ask SEC before warning again. A confirmed
# delisting retires the symbol (keeping its history) and
# ends the alert; anything unproven keeps warning.
delisted_on = await ticker_service.confirm_delisting(
db, symbol, last_bar=result.last_date
)
if delisted_on is not None:
await _record_system_event(
severity="info",
source=job_name,
code="ticker_delisted",
message=(
f"{symbol} delisted on {delisted_on} (SEC Form 25/15). "
"Retired from signals; price history retained."
),
symbol=symbol,
dedup_key=f"ticker_delisted:{symbol}",
)
else:
await _record_system_event(
severity="warning",
source=job_name,
@@ -1235,8 +1351,11 @@ async def run_event_study_job() -> None:
report = await run_event_study_and_store(db)
_runtime_progress(job_name, processed=1, total=1)
shipped = report.get("shipped") or {}
if report.get("available"):
metrics = report.get("metrics") or {}
# The shipped quadrant rule is the headline; the fitted-threshold
# variant lives under report["fitted"] and is not what fires.
metrics = shipped.get("metrics") or {}
msg = (
f"{metrics.get('events_warned', 0)}/{metrics.get('events', 0)} warned, "
f"{metrics.get('false_alarms_per_year', 0)} false alarms/year"
@@ -1244,7 +1363,10 @@ async def run_event_study_job() -> None:
else:
msg = report.get("reason", "no data")
_runtime_finish(job_name, "completed", processed=1, total=1, message=msg)
_log_event(logging.INFO, "job_complete", job=job_name, events=len(report.get("events", [])))
_log_event(
logging.INFO, "job_complete", job=job_name,
events=len(shipped.get("events") or []),
)
except Exception as exc:
_runtime_finish(job_name, "error", processed=0, total=1, message=str(exc))
_log_event(logging.ERROR, "job_error", job=job_name, error_type=type(exc).__name__, message=str(exc))
@@ -1300,54 +1422,14 @@ async def sync_ticker_universe() -> None:
# the intraday partial one (covers a long weekend / holiday gap).
_FINAL_REFETCH_DAYS = 5
_DAILY_PIPELINE_STEPS = [
("data_collector", "collect_ohlcv"),
("benchmark_collector", "collect_benchmark"),
("sentiment_collector", "collect_sentiment"),
("market_regime", "compute_market_regime"),
# Observational only — display/alerts; not trade selection.
("regime_monitor", "compute_regime_monitor"),
# Alerts after regime so quadrant changes reach Telegram in the morning.
# Dispatcher is change-driven; quiet days stay quiet. Setup alerts still
# fire on the near-close pipeline after the qualifying scan.
("alerts", "dispatch_alerts_job"),
]
# Near-close (~15:30 ET MonFri): refresh in-progress day-t bars (incremental
# ingestion overlaps the latest stored session), then the only daily
# qualifying R:R scan, then Telegram immediately so manual fills can still hit
# MOC cutoffs (~15:50/15:55). Under a 15-minute delayed SIP feed a 15:30 scan
# may see ~15:15 prices — immaterial for a 12-1 momentum signal.
#
# US early-close days (~3/year, 13:00 ET close): this job runs post-close and
# entries behave like stale_close (still acceptable per execution-recovery matrix).
# No exchange calendar dependency.
_NEAR_CLOSE_PIPELINE_STEPS = [
# Must land today's in-progress bar (~20 min behind live), or the scan falls
# back to the previous close and execution degrades to the stale_close floor.
("data_collector", "collect_ohlcv_for_scan"),
("rr_scanner", "scan_rr"),
# Straight after the scan so shadow entries mark at the same near-close
# prices the discretionary book is looking at.
("shadow_book", "run_shadow_book"),
("alerts", "dispatch_alerts_job"),
]
# After close (~16:45 ET MonFri): fresh OHLCV fetch so outcomes resolve on the
# final bar, not the near-close partial bar, then outcome/paper close.
_AFTER_CLOSE_PIPELINE_STEPS = [
("data_collector", "collect_ohlcv_final"),
("outcome_evaluator", "evaluate_outcomes"),
]
# Intraday (light): keep prices current and resolve outcomes through the day,
# without the expensive scan/sentiment. The dashboard recomputes live R:R from
# the latest price, so refreshing OHLCV is enough to stop prices lagging; the
# outcome step also closes paper trades that hit their stop/target intraday.
_INTRADAY_PIPELINE_STEPS = [
("data_collector", "collect_ohlcv"),
("outcome_evaluator", "evaluate_outcomes"),
]
# Step lists live in app.job_catalog so the runner, the admin API's pipeline
# membership and the UI's grouping all read one definition. Re-exported here
# under their original names: _run_pipeline and the scheduler_configured log
# payload refer to them directly.
_DAILY_PIPELINE_STEPS = job_catalog._DAILY_PIPELINE_STEPS
_NEAR_CLOSE_PIPELINE_STEPS = job_catalog._NEAR_CLOSE_PIPELINE_STEPS
_AFTER_CLOSE_PIPELINE_STEPS = job_catalog._AFTER_CLOSE_PIPELINE_STEPS
_INTRADAY_PIPELINE_STEPS = job_catalog._INTRADAY_PIPELINE_STEPS
# Warn if near-close fetch+scan+alert drifts past this — entries leave the close
# and the stale_close floor quietly becomes the ceiling.
@@ -1370,6 +1452,7 @@ async def _run_pipeline(job_name: str, steps: list[tuple[str, str]]) -> None:
if not await _is_job_enabled(db, job_name):
_log_event(logging.INFO, "job_skipped", job=job_name, reason="disabled")
_runtime_finish(job_name, "skipped", processed=0, total=0, message="Disabled")
await _persist_job_run(job_name)
return
total = len(steps)
@@ -1385,6 +1468,11 @@ async def _run_pipeline(job_name: str, steps: list[tuple[str, str]]) -> None:
await funcs[func_name]()
except Exception:
logger.exception("%s step %s failed", job_name, step_name)
# Outside the except on purpose: the step's own _runtime_finish has
# already recorded its outcome, so persisting here captures failures
# too. Steps are plain coroutine calls and fire no scheduler events,
# so the listener cannot see them -- this is their only write path.
await _persist_job_run(step_name)
done += 1
_runtime_finish(job_name, "completed", processed=done, total=total, message="Pipeline complete")
_log_event(logging.INFO, "job_complete", job=job_name)
@@ -1393,6 +1481,7 @@ async def _run_pipeline(job_name: str, steps: list[tuple[str, str]]) -> None:
_log_event(logging.ERROR, "job_error", job=job_name, error_type=type(exc).__name__, message=str(exc))
finally:
pipeline_run.release(token)
await _persist_job_run(job_name)
async def run_daily_pipeline() -> None:
@@ -1486,6 +1575,12 @@ SCHEDULE_DEFAULTS: dict[str, str] = {
"schedule_after_close_pipeline_cron": "45 16 * * mon-fri",
# Hourly mid-session price + outcome (10:0015:00 ET MonFri).
"schedule_intraday_pipeline_cron": "0 10-15 * * mon-fri",
# Both were interval jobs until 2026-08-08 and hit exactly the pitfall
# described above: configure_scheduler calls remove_all_jobs() on every
# startup, so an interval countdown restarts from zero each deploy. A 168h
# backtest needed a week of uninterrupted uptime to fire even once.
"schedule_backtest_cron": "0 3 * * sun",
"schedule_ticker_universe_cron": "0 1 * * *",
}
# job id -> schedule setting key
@@ -1496,6 +1591,8 @@ _CRON_JOBS: dict[str, str] = {
"near_close_pipeline": "schedule_near_close_pipeline_cron",
"after_close_pipeline": "schedule_after_close_pipeline_cron",
"intraday_pipeline": "schedule_intraday_pipeline_cron",
"backtest": "schedule_backtest_cron",
"ticker_universe_sync": "schedule_ticker_universe_cron",
}
@@ -1630,9 +1727,12 @@ def configure_scheduler(schedule_config: dict[str, str] | None = None) -> None:
id="intraday_pipeline", name="Intraday Pipeline", replace_existing=True,
)
# Independent interval jobs (own cadence, no ordering dependency)
# Independent jobs (own cadence, no ordering dependency). Cron, not interval,
# for the reason documented at SCHEDULE_DEFAULTS: an interval countdown
# restarts on every deploy, so these could be deferred indefinitely.
scheduler.add_job(
sync_ticker_universe, "interval", hours=24,
sync_ticker_universe,
_cron_trigger(cfg["schedule_ticker_universe_cron"], tz, "schedule_ticker_universe_cron"),
id="ticker_universe_sync", name="Ticker Universe Sync", replace_existing=True,
)
# Alerts auto-fire only via near_close_pipeline (scan → alert before MOC).
@@ -1643,7 +1743,8 @@ def configure_scheduler(schedule_config: dict[str, str] | None = None) -> None:
replace_existing=True, next_run_time=None,
)
scheduler.add_job(
run_backtest_job, "interval", hours=168,
run_backtest_job,
_cron_trigger(cfg["schedule_backtest_cron"], tz, "schedule_backtest_cron"),
id="backtest", name="Backtest", replace_existing=True,
)
# Deep history backfill: manual only (never auto-fires); triggered from
+2
View File
@@ -83,6 +83,8 @@ class ScheduleConfigUpdate(BaseModel):
schedule_near_close_pipeline_cron: str | None = Field(default=None, max_length=120)
schedule_after_close_pipeline_cron: str | None = Field(default=None, max_length=120)
schedule_intraday_pipeline_cron: str | None = Field(default=None, max_length=120)
schedule_backtest_cron: str | None = Field(default=None, max_length=120)
schedule_ticker_universe_cron: str | None = Field(default=None, max_length=120)
class PerformanceConfigUpdate(BaseModel):
+12 -1
View File
@@ -1,6 +1,6 @@
"""Ticker request/response schemas."""
from datetime import datetime
from datetime import date, datetime
from pydantic import BaseModel, Field
@@ -14,5 +14,16 @@ class TickerResponse(BaseModel):
symbol: str
name: str | None = None
created_at: datetime
# NULL == actively traded. Delisted symbols stay in the registry with their
# history and are excluded from signals — the date is what makes that
# visible instead of the row silently disappearing.
delisted_on: date | None = None
delisted_reason: str | None = None
model_config = {"from_attributes": True}
class TickerDelistingUpdate(BaseModel):
delisted_on: date = Field(
..., description="Effective date the symbol stopped trading"
)
+98 -69
View File
@@ -7,6 +7,7 @@ from passlib.hash import bcrypt
from sqlalchemy import delete, func, select
from sqlalchemy.ext.asyncio import AsyncSession
from app import job_catalog
from app.exceptions import DuplicateError, NotFoundError, ValidationError
from app.models.fundamental import FundamentalData
from app.models.ohlcv import OHLCVRecord
@@ -17,7 +18,7 @@ from app.models.settings import SystemSetting
from app.models.ticker import Ticker
from app.models.trade_setup import TradeSetup
from app.models.user import User
from app.services import settings_store
from app.services import job_run_store, settings_store
logger = logging.getLogger(__name__)
@@ -606,91 +607,110 @@ async def get_pipeline_readiness(db: AsyncSession) -> list[dict]:
# Job control (placeholder — scheduler is Task 12.1)
# ---------------------------------------------------------------------------
VALID_JOB_NAMES = {
"data_collector",
"data_backfill",
"benchmark_collector",
"sentiment_collector",
"dolt_earnings_import",
"sec_fundamentals_import",
"rr_scanner",
"ticker_universe_sync",
"outcome_evaluator",
"alerts",
"market_regime",
"regime_monitor",
"event_study",
"backtest",
"daily_pipeline",
"near_close_pipeline",
"after_close_pipeline",
"intraday_pipeline",
"shadow_book",
}
# Job identity, labels and pipeline membership now live in app.job_catalog, which
# derives PIPELINE_MEMBERS from the pipeline step lists instead of restating them.
# Re-exported here because callers (routers, tests) import them from this module.
VALID_JOB_NAMES = job_catalog.VALID_JOB_NAMES
JOB_LABELS = job_catalog.JOB_LABELS
PIPELINE_MEMBERS = job_catalog.PIPELINE_MEMBERS
JOB_LABELS = {
"data_collector": "Data Collector (OHLCV)",
"data_backfill": "Data Backfill (deep history)",
"benchmark_collector": "Benchmark Collector",
"sentiment_collector": "Sentiment Collector",
"dolt_earnings_import": "Dolt Earnings Import",
"sec_fundamentals_import": "SEC Fundamentals Import",
"rr_scanner": "R:R Scanner",
"ticker_universe_sync": "Ticker Universe Sync",
"outcome_evaluator": "Outcome Evaluator",
"alerts": "Alerts Dispatcher",
# Keys are persisted job ids and must not change; these are display only.
"market_regime": "Market Trend (SPY)",
"regime_monitor": "AI/Tech Risk Monitor",
"event_study": "Event Study",
"backtest": "Backtest",
"daily_pipeline": "Morning Pipeline",
"near_close_pipeline": "Near-Close Pipeline (scan+alert)",
"after_close_pipeline": "After-Close Pipeline (outcome)",
"intraday_pipeline": "Intraday Pipeline",
"shadow_book": "Shadow Book (auto-traded strategy)",
}
# Anything further out than this is a parked backstop, not a schedule: pipeline
# steps and manual jobs are registered on a 520-week interval, and triggering one
# re-arms it. Belt-and-braces behind the category rule in _next_run_fields.
_NEXT_RUN_HORIZON_DAYS = 365
# Jobs driven by a pipeline (in order) rather than their own auto timer.
PIPELINE_MEMBERS = {
"data_collector",
"benchmark_collector",
"sentiment_collector",
"rr_scanner",
"outcome_evaluator",
"alerts",
"market_regime",
"regime_monitor",
"shadow_book",
def _visible_next_run(next_run: datetime | None) -> datetime | None:
"""Drop a next-run that is really the parked backstop."""
if next_run is None:
return None
horizon = datetime.now(next_run.tzinfo) + timedelta(days=_NEXT_RUN_HORIZON_DAYS)
return None if next_run > horizon else next_run
def _own_next_run(scheduler, name: str) -> datetime | None:
# getattr: APScheduler only sets next_run_time once the scheduler is running,
# so a job registered but not yet started has no such attribute at all.
job = scheduler.get_job(name)
return _visible_next_run(getattr(job, "next_run_time", None)) if job else None
def _next_run_fields(scheduler, name: str, enabled_map: dict[str, bool]) -> dict:
"""Where this job's next run comes from, decided by category not by clock.
A pipeline step has no meaningful schedule of its own, so reporting one is
the bug: its parent's timer is the answer. Manual jobs have no answer at all,
and saying so beats rendering a parked backstop as a date.
"""
category = job_catalog.JOB_CATEGORY.get(name)
if category == job_catalog.CATEGORY_STEP:
parents = job_catalog.PIPELINES_BY_MEMBER.get(name, ())
soonest: datetime | None = None
via: str | None = None
for parent in parents:
if not enabled_map.get(parent, True):
continue
candidate = _own_next_run(scheduler, parent)
if candidate is not None and (soonest is None or candidate < soonest):
soonest, via = candidate, parent
return {
"next_run_at": None,
"next_run_source": "via_pipeline",
"via_next_run_at": soonest.isoformat() if soonest else None,
"via_next_run_job": via,
}
if category == job_catalog.CATEGORY_MANUAL:
return {
"next_run_at": None,
"next_run_source": "manual_only",
"via_next_run_at": None,
"via_next_run_job": None,
}
own = _own_next_run(scheduler, name)
return {
"next_run_at": own.isoformat() if own else None,
"next_run_source": "own_schedule",
"via_next_run_at": None,
"via_next_run_job": None,
}
async def list_jobs(db: AsyncSession) -> list[dict]:
"""Return status of all scheduled jobs."""
"""Return status of all scheduled jobs, grouped and ordered by category."""
from app.scheduler import get_job_runtime_snapshot, scheduler
visible = sorted(VALID_JOB_NAMES - job_catalog.HIDDEN_JOBS, key=job_catalog.sort_order)
# One query for every flag instead of one per job. Parents are read too, since
# a step reports its parent's next run only while that parent is enabled.
flags = await settings_store.get_map(
db, [f"job_{name}_enabled" for name in VALID_JOB_NAMES]
)
enabled_map = {
name: flags.get(f"job_{name}_enabled", "true") == "true"
for name in VALID_JOB_NAMES
}
last_runs = await job_run_store.get_map(db, visible)
jobs_out = []
for name in sorted(VALID_JOB_NAMES):
# Check enabled setting
setting = await settings_store.get_setting(db, f"job_{name}_enabled")
enabled = setting.value == "true" if setting else True # default enabled
# Get scheduler job info
for name in visible:
job = scheduler.get_job(name)
next_run = None
if job and job.next_run_time:
next_run = job.next_run_time.isoformat()
runtime = get_job_runtime_snapshot(name)
last = last_runs.get(name)
jobs_out.append({
"name": name,
"label": JOB_LABELS.get(name, name),
"enabled": enabled,
"next_run_at": next_run,
"via_pipeline": name in PIPELINE_MEMBERS,
"enabled": enabled_map.get(name, True),
"category": job_catalog.JOB_CATEGORY.get(name),
"sort_order": job_catalog.sort_order(name),
# Parent pipelines for a step; the steps themselves for a pipeline.
"pipelines": list(job_catalog.PIPELINES_BY_MEMBER.get(name, ())),
"steps": [step for step, _ in job_catalog.PIPELINE_STEPS.get(name, ())],
"registered": job is not None,
"running": bool(runtime.get("running", False)),
# runtime_* are strictly live in-memory state. Persisted history is
# reported separately as last_run_*, so a stale error cannot pin the
# status chip or the rate-limit banner.
"runtime_status": runtime.get("status"),
"runtime_processed": runtime.get("processed"),
"runtime_total": runtime.get("total"),
@@ -699,6 +719,15 @@ async def list_jobs(db: AsyncSession) -> list[dict]:
"runtime_started_at": runtime.get("started_at"),
"runtime_finished_at": runtime.get("finished_at"),
"runtime_message": runtime.get("message"),
# Survives restarts, unlike runtime_*. Reported separately so the
# status chip keeps meaning "state now" rather than "last outcome,
# forever" -- an error a week ago must not read as Inactive today.
"last_run_at": last.finished_at.isoformat() if last else None,
"last_run_status": last.status if last else None,
"last_run_message": last.message if last else None,
"last_run_processed": last.processed if last else None,
"last_run_total": last.total if last else None,
**_next_run_fields(scheduler, name, enabled_map),
})
return jobs_out
+108
View File
@@ -97,6 +97,14 @@ SIGNAL_BUNDLE_MAX_CHARS = 3900 # Telegram limit is 4096; keep room for HTML par
# Hysteresis (a deadband around each divider) stops a point sitting on a boundary
# from flip-flopping; the cooldown caps how often a genuine change can re-alert.
QUAD_TYPE = "regime_quadrant"
# The fundamental channel gets its own alerts rather than shifting a score:
# "the context changed" and "both channels are elevated" are different facts from
# "the market axes moved", and fusing them into one number would destroy exactly
# the information an operator uses to decide how much the alert is worth.
FUND_TYPE = "regime_fundamental"
CONFLUENCE_TYPE = "regime_confluence"
# States that count as fundamental risk for the confluence test.
FUND_ADVERSE = "adverse"
QUAD_X_DIV = 50.0 # v3 State divider (backend response is authoritative)
QUAD_Y_DIV = 40.0 # v3 Warning divider; the axes have different ranges
QUAD_MARGIN = 5.0 # half-width of the hysteresis deadband around each divider
@@ -859,16 +867,111 @@ async def _collect_regime_quadrant(db: AsyncSession) -> list[tuple[str, str]]:
)
else:
metrics = f"State {x:.0f} · Warning {y:.0f}"
# The fundamental channel is reported, never added in: this alert is about
# the two market axes, and the context is stated beside them so a reader can
# judge confluence themselves rather than being handed a fused number.
context = data.get("fundamental_context") or {}
context_line = (
f"fundamentals: {context.get('state', 'unknown')} "
f"({context.get('evidence_quality', 'unavailable')})\n"
)
text = (
f"🧭 <b>AI/Tech risk quadrant change</b>\n"
f"{QUAD_LABELS.get(prev, prev)}{QUAD_LABELS.get(new_q, new_q)}\n"
f"{metrics}\n"
f"{context_line}"
f"coverage: state {state.get('coverage'):.0f}% / warning {warning.get('coverage'):.0f}%\n"
f"<i>Risk thermometer - not a trade signal.</i>"
)
return [(_quadrant_log_key(new_q, x, y, basket_hash), text)]
async def _last_logged_key(db: AsyncSession, alert_type: str) -> str | None:
"""Most recent logged key for a type, our baseline for change detection."""
result = await db.execute(
select(AlertLog.dedup_key)
.where(AlertLog.alert_type == alert_type)
.order_by(AlertLog.created_at.desc())
.limit(1)
)
row = result.first()
return row[0] if row else None
async def _collect_regime_fundamental(db: AsyncSession) -> list[tuple[str, str, str]]:
"""Fundamental-context changes and market/fundamental confluence.
Two triggers, deliberately separate from the quadrant alert and from each
other, because they answer different questions: *what the evidence says* and
*whether both channels agree*. Neither is derived by moving a score.
``unknown`` never alerts. An absence of evidence is not a change in the
evidence, and alerting on it would train the reader to ignore the channel.
Both seed silently on first run, exactly as the quadrant alert does.
"""
from app.services.regime_monitor_service import get_regime_monitor
data = await get_regime_monitor(db)
if not data.get("available"):
return []
warning = data.get("warning") or {}
context = data.get("fundamental_context") or {}
state = str(context.get("state") or "unknown")
# `usable`, not `available`: the state is deliberately preserved past its
# staleness horizon so the card can keep showing the last thing observed, and
# an observation whose extraction failed is fresh but knows nothing. Neither
# may confirm anything — without this gate a months-old adverse read silently
# corroborates every new Warning crossing forever, which is the strongest
# claim this channel makes and the one it has least right to make.
usable = bool(context.get("usable"))
score = warning.get("score")
quality = data.get("data_quality") or {}
if not quality.get("is_fresh") or float(warning.get("coverage") or 0) < 75:
return []
quadrant_cfg = data.get("quadrant_config") or {}
y_div = float(quadrant_cfg.get("warning_divider", QUAD_Y_DIV))
warning_elevated = score is not None and float(score) >= y_div
out: list[tuple[str, str, str]] = []
previous_state = await _last_logged_key(db, FUND_TYPE)
if previous_state is None:
_log_alert(db, FUND_TYPE, state) # seed
elif previous_state != state and state != "unknown" and usable:
effective = context.get("effective_date")
out.append((
FUND_TYPE,
state,
f"📋 <b>Fundamental context changed</b>\n"
f"{previous_state}{state}\n"
f"evidence: {context.get('evidence_quality', 'unavailable')}"
+ (f" · effective {effective}" if effective else "")
+ "\n<i>Context channel — not a score, not a trade signal.</i>",
))
confluence = "yes" if (warning_elevated and state == FUND_ADVERSE and usable) else "no"
previous_confluence = await _last_logged_key(db, CONFLUENCE_TYPE)
if previous_confluence is None:
_log_alert(db, CONFLUENCE_TYPE, confluence) # seed
elif previous_confluence != confluence and confluence == "yes":
out.append((
CONFLUENCE_TYPE,
confluence,
f"⚠️ <b>Confluence: market and fundamental risk both elevated</b>\n"
f"Warning {float(score):.0f} (≥ {y_div:.0f}) with fundamentals {state}\n"
f"evidence: {context.get('evidence_quality', 'unavailable')}\n"
f"<i>Highest attention. Still a thermometer — not a trade signal.</i>",
))
elif previous_confluence != confluence:
# Falling out of confluence is a state change worth recording as the new
# baseline, but not worth a message.
_log_alert(db, CONFLUENCE_TYPE, confluence)
return out
# ---------------------------------------------------------------------------
# Dispatch
# ---------------------------------------------------------------------------
@@ -961,6 +1064,11 @@ async def dispatch_alerts(db: AsyncSession) -> dict:
# cooldown/hysteresis handled in the collector (like score drops)
for key, text in await _collect_regime_quadrant(db):
outgoing.append((QUAD_TYPE, key, text))
# Deliberately three separate messages off one toggle, not one fused
# signal: the market axes and the fundamental channel are different kinds
# of evidence, and an operator needs to know which one moved.
for alert_type, key, text in await _collect_regime_fundamental(db):
outgoing.append((alert_type, key, text))
if cfg["trade_closed"]:
for key, text, pnl_usd in await _collect_closed_trades(db):
+154 -78
View File
@@ -1460,6 +1460,21 @@ def _mp_context():
return None
async def _rollback_quietly(db: AsyncSession, context: str) -> None:
"""Discard a failed unit of work so later statements on this session survive.
Every DB call in ``run_backtest`` is best-effort one unreadable ticker must
not abort the whole replay. But swallowing the exception alone leaves asyncpg
in "current transaction is aborted": every later statement then fails the same
way until the first unguarded one (the report write) surfaces it as the job
error, long after the real cause. Same guard as ``price_service``.
"""
try:
await db.rollback()
except Exception:
logger.exception("Session rollback after %s also failed", context)
async def _fetch_columns(db: AsyncSession, symbol: str) -> tuple | None:
"""Read one ticker's OHLCV and detach it to primitive column arrays in the
event loop (safe ORM access), ready to hand to a worker. None if no data."""
@@ -2610,6 +2625,19 @@ def _simulate_portfolio(
diag = sharpe_diagnostics(rets)
sharpe = diag["sharpe"]
# Sortino: the same numerator as Sharpe over downside deviation about a zero
# target. The denominator divides by len(rets) — the full-sample lower partial
# moment — NOT by the count of down days, which would shrink the denominator
# and inflate the ratio. n >= 3 matches sharpe_diagnostics so the two appear
# together or not at all. No down days is +inf, reported as None.
sortino = None
downside = [r for r in rets if r < 0.0]
if len(rets) >= 3 and downside:
mean_ret = sum(rets) / len(rets)
dd = math.sqrt(sum(r * r for r in downside) / len(rets))
if dd > 0:
sortino = round(mean_ret / dd * math.sqrt(252.0), 2)
# Per-calendar-year returns off the equity curve — shows whether every year
# contributed or one exceptional stretch carried the result.
yearly: list[dict] = []
@@ -2635,8 +2663,40 @@ def _simulate_portfolio(
),
})
# Gain-to-Pain off the same curve, on MONTHLY returns: Schwager's ratio is
# defined monthly and the daily variant is not comparable to published
# figures. Distinct loop variables from the yearly pass above — that one exits
# with last_eq at final equity, so reusing its names silently corrupts the
# first month. The monthly series itself is not emitted: 36-120 floats per
# strategy per lookback would bloat the single stored report blob.
monthly: list[float] = []
month_start_eq = curve[0][1]
month_last_eq = curve[0][1]
cur_month = date.fromordinal(curve[0][0]).replace(day=1)
for o, eq in curve:
m = date.fromordinal(o).replace(day=1)
if m != cur_month:
if month_start_eq > 0:
monthly.append(month_last_eq / month_start_eq - 1.0)
cur_month = m
month_start_eq = month_last_eq
month_last_eq = eq
if month_start_eq > 0:
monthly.append(month_last_eq / month_start_eq - 1.0)
# Schwager: SUM OF ALL monthly returns over the absolute sum of the negative
# ones. Not sum(positive)/|sum(negative)| — that is profit-factor-shaped and
# sits exactly 1.0 higher for every input, since sum(all) = sum(pos) - |sum(neg)|.
monthly_pain = -sum(r for r in monthly if r < 0.0)
gain_to_pain = round(sum(monthly) / monthly_pain, 2) if monthly_pain > 0 else None
pnls = [t["pnl"] for t in trades]
wins = sum(1 for p in pnls if p > 0)
# Dollar-based, over closed-trade P&L. Distinct from the R-based profit_factor
# in _robustness_stats; the two never share an object.
gross_win = sum(p for p in pnls if p > 0)
gross_loss = -sum(p for p in pnls if p < 0)
profit_factor = round(gross_win / gross_loss, 2) if gross_loss > 0 else None
reason_counts = {
reason: sum(1 for t in trades if t["reason"] == reason)
for reason in sorted({t["reason"] for t in trades})
@@ -2691,7 +2751,13 @@ def _simulate_portfolio(
"total_return_pct": round(total_return_pct, 1),
"cagr_pct": round(cagr_pct, 1) if cagr_pct is not None else None,
"max_drawdown_pct": round(max_dd_pct, 1),
# calmar IS MAR here (CAGR / max drawdown) — one field, two names.
"calmar": round(calmar, 2) if calmar is not None else None,
# Emitted unconditionally even when None: the UI treats an ABSENT key as
# "report predates these metrics", so presence is a contract.
"sortino": sortino,
"gain_to_pain": gain_to_pain,
"profit_factor": profit_factor,
"sharpe": sharpe,
"sharpe_se": diag["sharpe_se"],
"psr": diag["psr"],
@@ -3878,40 +3944,11 @@ def _build_recommendation(report: dict) -> dict:
})
q = report.get("overall_qualified") or {}
target_net = q.get("net_avg_r")
# Legacy diagnostic: target/stop race vs the best fixed hold.
time_rows = [r for r in report.get("time_exit_sweep") or [] if r.get("net_avg_r") is not None]
best_hold = max(time_rows, key=lambda r: r["net_avg_r"], default=None)
sim_rows = {
p.get("policy"): p
for p in (report.get("portfolio_sim") or {}).get("policies", [])
}
hold_sim = sim_rows.get("hold")
if best_hold is not None and target_net is not None:
if best_hold["net_avg_r"] > target_net + _EXIT_SWITCH_THRESHOLD:
text = (
f"Legacy exit diagnostic: hold {best_hold['hold_days']} trading days with the initial stop "
f"({best_hold['net_avg_r']:+.2f}R net/trade vs {target_net:+.2f}R for the S/R target exit)."
)
target_sim = sim_rows.get("target")
if (
hold_sim is not None and target_sim is not None
and hold_sim.get("cagr_pct") is not None and target_sim.get("cagr_pct") is not None
):
text += (
f" The simulated book agrees: {hold_sim['cagr_pct']:+.1f}% vs "
f"{target_sim['cagr_pct']:+.1f}% CAGR at similar drawdown."
)
items.append({"topic": "exit", "text": text})
else:
items.append({
"topic": "exit",
"text": (
f"Legacy exit diagnostic: keep the S/R target exit ({target_net:+.2f}R net/trade) — "
"no fixed hold beats it by a meaningful margin."
),
})
# Nothing here reads time_exit_sweep any more. The hold-vs-target comparison
# is not reported (both are exits the production book replaced, so choosing
# between them cannot lead to an action), and the robustness check below no
# longer picks its basis from them either.
# Gate floors, judged under the hold exit (the ablation's Hold column).
ablation = {r["variant"]: r for r in report.get("gate_ablation") or []}
@@ -3959,33 +3996,32 @@ def _build_recommendation(report: dict) -> dict:
),
})
# Book vs benchmark.
book = hold_sim or sim_rows.get("target")
if book is not None and book.get("spy_return_pct") is not None:
edge = book["total_return_pct"] - book["spy_return_pct"]
# Book vs benchmark — read from the SAME production monitor row the page
# shows in its tiles. It used to read the hold/target policy sim, so the
# recommendation quoted a different portfolio return than the tile directly
# above it, against an identical SPY figure. Those policies are legacy
# diagnostics; the production book is the ATR trail.
if production_row is not None and production_row.get("spy_return_pct") is not None:
edge = production_row["total_return_pct"] - production_row["spy_return_pct"]
verdict = "beats" if edge > 0 else "LAGS"
items.append({
"topic": "benchmark",
"text": (
f"Book vs SPY: {verdict} buy-and-hold by {edge:+.1f} points "
f"({book['total_return_pct']:+.1f}% vs {book['spy_return_pct']:+.1f}%), "
f"max drawdown {book['max_drawdown_pct']:.1f}%."
f"({production_row['total_return_pct']:+.1f}% vs "
f"{production_row['spy_return_pct']:+.1f}%)."
),
})
# Robustness: does the edge survive without the biggest winners? Judged on
# the RECOMMENDED exit — outlier dependence under an exit we'd abandon
# would be the wrong warning.
hold_recommended = (
best_hold is not None and target_net is not None
and best_hold["net_avg_r"] > target_net + _EXIT_SWITCH_THRESHOLD
)
if hold_recommended and best_hold.get("net_avg_r_ex_top5") is not None:
trimmed = best_hold["net_avg_r_ex_top5"]
basis = f"under the recommended {best_hold['hold_days']}d hold"
else:
# Robustness: does the edge survive without the biggest winners?
#
# There is no ATR-trail equivalent of this number in the report — the only
# ex-top-5% figure is the gate-level target/stop grading. So it is reported
# on that basis and SAYS SO, rather than being dressed up as a verdict on the
# production book. It used to pick between "the recommended Nd hold" and "the
# S/R target exit", naming a rejected exit as recommended.
trimmed = q.get("net_avg_r_ex_top5")
basis = "under the S/R target exit"
basis = "gate-level grading, not the production ATR-trail book"
if trimmed is not None:
if trimmed > 0:
items.append({
@@ -4006,20 +4042,20 @@ def _build_recommendation(report: dict) -> dict:
),
})
if headline is None and hold_recommended:
cagr_note = (
f" (~{hold_sim['cagr_pct']:.0f}% CAGR simulated)"
if hold_sim is not None and hold_sim.get("cagr_pct") is not None
else ""
)
headline = (
f"Trade the qualified list long-only; hold {best_hold['hold_days']} trading days "
f"with the initial ATR stop{cagr_note}."
)
# No fallback headline. It used to recommend the fixed-hold exit whenever the
# portfolio monitor was missing, which meant a report without a production
# row advised an exit the production book had already replaced. A report that
# cannot describe the production baseline states no baseline.
return {
"headline": headline,
"items": items,
# Which monitor row every production/benchmark figure above was read
# from. The page defaults its lookback selector to this, so the tiles and
# the recommendation cannot open on different windows — they used to,
# because this preferred "all" while the UI defaulted to "3y".
"basis_lookback": (production_row or {}).get("lookback"),
"basis_lookback_label": (production_row or {}).get("lookback_label"),
"note": "Derived from this report's numbers on every run — the advice flips if the data does.",
}
@@ -4037,9 +4073,12 @@ async def run_backtest(
config = await get_recommendation_config(db)
activation = await get_activation_config(db)
result = await db.execute(select(Ticker).order_by(Ticker.symbol))
tickers = list(result.scalars().all())
total = len(tickers)
# Plain strings, not Ticker instances: the rollbacks below expire any ORM
# objects held across them, and touching an expired attribute afterwards
# triggers sync lazy-loading, which raises on an AsyncSession.
result = await db.execute(select(Ticker.symbol).order_by(Ticker.symbol))
symbols = list(result.scalars().all())
total = len(symbols)
rank_only_symbols = await _load_research_rank_only_symbols(db)
if rank_only_symbols:
logger.info(json.dumps({
@@ -4063,6 +4102,7 @@ async def run_backtest(
)
except Exception:
logger.exception("Benchmark load for residual momentum failed")
await _rollback_quietly(db, "benchmark load")
def _merge(result: tuple[list[dict], dict]) -> None:
cands, series = result
@@ -4094,26 +4134,27 @@ async def run_backtest(
done = 0
with pool:
for start in range(0, total, chunk):
batch = tickers[start : start + chunk]
batch = symbols[start : start + chunk]
futures = []
for ticker in batch:
for symbol in batch:
try:
columns = await _fetch_columns(db, ticker.symbol)
columns = await _fetch_columns(db, symbol)
except Exception:
logger.exception("Backtest fetch failed for %s", ticker.symbol)
logger.exception("Backtest fetch failed for %s", symbol)
await _rollback_quietly(db, f"fetch for {symbol}")
continue
if columns is not None:
futures.append(loop.run_in_executor(
pool,
_replay_and_signals,
ticker.symbol,
symbol,
columns,
config,
activation,
benchmark_closes,
target_model,
cadence,
ticker.symbol in rank_only_symbols,
symbol in rank_only_symbols,
))
for result in await asyncio.gather(*futures, return_exceptions=True):
if isinstance(result, Exception):
@@ -4126,25 +4167,26 @@ async def run_backtest(
else:
# Sequential fallback (Windows / 1 worker): run each replay in a worker
# thread so the event loop — and the API server — stays responsive.
for index, ticker in enumerate(tickers):
for index, symbol in enumerate(symbols):
if progress_cb is not None:
progress_cb(index, total, ticker.symbol)
progress_cb(index, total, symbol)
try:
columns = await _fetch_columns(db, ticker.symbol)
columns = await _fetch_columns(db, symbol)
if columns is not None:
_merge(await asyncio.to_thread(
_replay_and_signals,
ticker.symbol,
symbol,
columns,
config,
activation,
benchmark_closes,
target_model,
cadence,
ticker.symbol in rank_only_symbols,
symbol in rank_only_symbols,
))
except Exception:
logger.exception("Backtest replay failed for %s", ticker.symbol)
logger.exception("Backtest replay failed for %s", symbol)
await _rollback_quietly(db, f"replay for {symbol}")
if progress_cb is not None and total:
progress_cb(total, total, "")
@@ -4209,6 +4251,7 @@ async def run_backtest(
)
except Exception:
logger.exception("Benchmark load for the portfolio sim failed")
await _rollback_quietly(db, "portfolio-sim benchmark load")
for policy in ("target", "hold"):
sim = _simulate_portfolio(
@@ -4229,6 +4272,7 @@ async def run_backtest(
live_exit_policy = await get_exit_policy(db)
except Exception:
logger.exception("Live exit policy load failed; monitor uses defaults")
await _rollback_quietly(db, "exit policy load")
portfolio_monitor_report = _portfolio_monitor(
candidates, price_columns, spy_closes, hold_horizon,
live_exit_policy=live_exit_policy,
@@ -4248,6 +4292,11 @@ async def run_backtest(
)
except Exception:
logger.exception("Portfolio simulation failed")
# Catches the price_columns fetch loop, which has no handler of its
# own. The inner handlers above may already have rolled back; a
# rollback on a clean session is a no-op, so this stays safe as the
# backstop for whichever DB call actually failed.
await _rollback_quietly(db, "portfolio simulation")
report = {
"generated_at": datetime.now(timezone.utc).isoformat(),
@@ -4387,11 +4436,38 @@ async def run_and_store(
async def get_backtest_report(db: AsyncSession) -> dict | None:
"""Return the last cached backtest report, or None if never run."""
"""Return the last cached backtest report, or None if never run.
The recommendation is **re-derived from the cached report** rather than
served as stored. It is a pure function of the numbers already in the
report the payload's own note says it is derived from them on every run —
so recomputing costs nothing and keeps one class of bug out:
A report cached by an older build carries that build's recommendation. After
a change to how the recommendation is sourced, the page would keep showing
the old one quoting the legacy policy book, naming a rejected exit as
"recommended", and omitting ``basis_lookback``, which in turn let the
lookback selector default somewhere else. The result was the exact
tiles-disagree-with-recommendation contradiction this rebuild exists to
prevent, silently, until the next scheduled run happened to overwrite it.
Re-deriving means a corrected recommendation appears on the first page load
after deploy instead of after the next backtest.
"""
setting = await settings_store.get_setting(db, KEY_REPORT)
if setting is None:
return None
try:
return json.loads(setting.value)
report = json.loads(setting.value)
except (TypeError, ValueError):
return None
if not isinstance(report, dict):
return None
try:
report["recommendation"] = _build_recommendation(report)
except Exception:
# Fail closed: drop it rather than fall back to the stored one, which is
# precisely the stale derivation this rebuild is here to replace.
logger.exception("Could not rebuild the backtest recommendation; omitting it")
report.pop("recommendation", None)
return report
+2 -9
View File
@@ -25,6 +25,7 @@ from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from app.models.ticker import Ticker
from app.services import ticker_service
from app.services.price_service import query_ohlcv
logger = logging.getLogger(__name__)
@@ -112,7 +113,7 @@ def compute_divergence_series(
async def _load_universe_closes(
db: AsyncSession, symbols: list[str] | None = None
) -> dict[str, Series]:
stmt = select(Ticker).order_by(Ticker.symbol)
stmt = ticker_service.active_only(select(Ticker).order_by(Ticker.symbol))
if symbols is not None:
stmt = stmt.where(Ticker.symbol.in_(symbols))
result = await db.execute(stmt)
@@ -148,11 +149,3 @@ async def compute_breadth_details(
"""Breadth values plus the qualifying-member count for snapshot metadata."""
closes_by_symbol = await _load_universe_closes(db, symbols)
return _breadth_with_counts(closes_by_symbol, window, min_tickers)
async def compute_breadth_today(db: AsyncSession) -> float | None:
"""Latest breadth reading (thin wrapper, for future live use)."""
series = await compute_breadth_series(db)
if not series:
return None
return series[max(series)]
+6 -2
View File
@@ -32,7 +32,7 @@ from app.config import settings
from app.database import insert_for_session
from app.models.earnings_event import EarningsEvent
from app.models.ticker import Ticker
from app.services import dolt_client, earnings_alignment
from app.services import dolt_client, earnings_alignment, ticker_service
from app.services.data_import import ValidationResult
logger = logging.getLogger(__name__)
@@ -289,7 +289,11 @@ class DoltEarningsImporter:
# -- helpers -----------------------------------------------------------
async def _load_universe(self, db) -> dict[str, int]:
rows = (await db.execute(select(Ticker.id, Ticker.symbol))).all()
rows = (
await db.execute(
ticker_service.active_only(select(Ticker.id, Ticker.symbol))
)
).all()
return {
earnings_alignment.normalise_symbol(symbol): tid
for tid, symbol in rows
+770 -71
View File
@@ -1,15 +1,48 @@
"""Compact chronological validation for the AI/Tech Risk Monitor warning score.
"""Chronological validation for the AI/Tech Risk Monitor warning score.
The study calls its outcome a 10% correction, uses the first 70% of sessions to
freeze an 80th-percentile warning threshold, and reports alarm episodes only on
the final 30%. It is still labelled exploratory while the fixed breadth basket
is reconstructed before its freeze date.
The outcome is a 10% correction in the leader, never a regime break. Two rules
are measured against it, and they answer different questions:
* **shipped** -- the quadrant-change rule that actually reaches Telegram
(``alert_service._collect_regime_quadrant``). Its thresholds are fixed
constants chosen by scenario arithmetic, so nothing is fitted, so there is no
training set to protect and the whole sample is evaluable. This is the
headline.
* **fitted** -- the original study: an 80th-percentile Warning threshold frozen
on the first 70% of sessions and measured on the last 30%. Kept because it is
what the methodology document reports, and because a fitted threshold is a
genuinely different question -- but it is measured on the four corrections that
happen to fall in the holdout, which is too few to read as a property of the
score.
Both are scored by the same ``evaluate_alarms`` harness, alongside ablations
(does the quadrant machinery earn its place?), external baselines (does the
score earn its complexity?), and a random-alarm null (is any of this better than
chance?). Without those rows a bare "2 of 4" is unreadable in either direction.
The fundamental channel is compared, never fused. It appears as its own rule
(transitions into an adverse state), as a confluence gate (a market crossing kept
only when the state agrees), and as a market-only comparator over the identical
window -- because with ~10 correction events and almost no fundamental history,
any weight that combined it with the market axes would be a policy preference
presented as a measurement.
Those three rows are **coverage-matched**: scored only on the sessions where the
channel had usable context and on the corrections whose warning horizon fell
inside it, and marked ``measurable: false`` until enough corrections are covered.
A fundamental rule scores zero whether it is wrong or merely absent, so scoring
it over the market rows' full sample would turn a fortnight of observations into
a 0/10 that reads as a failed test.
Still labelled exploratory while the fixed breadth basket is reconstructed
before its freeze date.
"""
from __future__ import annotations
import json
import logging
import random
from datetime import date, datetime, timedelta, timezone
from sqlalchemy.ext.asyncio import AsyncSession
@@ -17,11 +50,26 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.services import breadth_service, settings_store
from app.services import regime_monitor_service as rms
from app.services.admin_service import update_setting
from app.services.alert_service import (
QUAD_COOLDOWN_DAYS,
QUAD_MARGIN,
QUAD_X_DIV,
QUAD_Y_DIV,
_classify_quadrant,
)
logger = logging.getLogger(__name__)
KEY_REPORT = "regime_event_study"
# Report shape, independent of METHODOLOGY. A cached report from an older shape
# parses fine and reports the current methodology, so without this check the
# panel would render a report missing half its blocks. Bumping discards the cache
# the way a methodology change does -- and it is the *only* thing that does so
# here, because the fundamental-channel rework left METHODOLOGY on v4 (the scores
# did not change), so the methodology check cannot catch a stale report.
STUDY_SCHEMA = 3
EVENT_THRESHOLD_PCT = 10.0
EVENT_COOLDOWN_DAYS = 40
DRAWDOWN_LOOKBACK = 252
@@ -33,6 +81,21 @@ TRAIN_FRACTION = 0.70
MIN_EVENTS_FOR_CONFIDENCE = 8
SENSOR_MISMATCH_TOLERANCE = 0.10
# _collect_regime_quadrant confirms against get_regime_history(db, days=14), so a
# prior session older than that window is not available to confirm with.
QUAD_HISTORY_DAYS = 14
# Quadrants with Warning above its divider: "1" early warning, "2" active stress.
WARNING_QUADRANTS = ("1", "2")
STRESS_QUADRANT = ("2",)
# Draws for the random-alarm null. Seeded, because a cached report that moves
# on re-run for RNG reasons is worse than no report.
NULL_DRAWS = 2000
NULL_SEED = 20260812
BASELINE_SMA_WINDOW = 50
BASELINE_VIX_LEVEL = 20.0
def _median(values: list[float]) -> float | None:
if not values:
@@ -148,41 +211,355 @@ def evaluate_alarms(
}
def _warning_series(
def _score_rule(
alarm_indices: list[int],
event_indices: list[int],
dates: list[date],
horizon: int,
sessions: int,
) -> dict:
"""``evaluate_alarms`` plus the annualised false-alarm rate for one rule.
The rate is ``None`` when the rule had no eligible sessions. Dividing by a
tiny floor instead produced 5e9 alarms/year for a coverage-matched rule with
an empty window -- a number that means "undefined" while looking like a
measurement, which is the failure mode this whole panel is built to avoid.
"""
metrics = evaluate_alarms(alarm_indices, event_indices, dates, horizon)
metrics["false_alarms_per_year"] = (
round(metrics["false_alarms"] / (sessions / 252.0), 2) if sessions > 0 else None
)
return metrics
# ---------------------------------------------------------------------------
# The shipped rule
# ---------------------------------------------------------------------------
def _axis_rows(
prices: dict[str, rms.Series],
breadth_divergence: dict[date, float],
vix_series: rms.Series | None,
oas_series: rms.Series | None,
breadth_series: rms.Series | None,
divergence_series: rms.Series | None,
dates: list[date],
config: dict,
oas_series: rms.Series | None = None,
) -> tuple[dict[date, float], dict[date, int]]:
"""Warning score per session plus how many sensors backed it.
observations: list[dict] | None = None,
) -> dict[date, dict]:
"""State and Warning per session, from the function that writes snapshots.
v2 re-derived this by hand from ``WARNING_WEIGHTS`` and so would have kept
measuring the old construct after a scoring change. Since v3 dropped
fundamentals from the score, this is now exactly the live Warning score
rather than a technical-only approximation of it.
Calling ``_compute_index`` rather than re-deriving the two axes is the same
anti-drift argument that produced ``warning_sensor_scores``: the v2 study
re-derived Warning by hand and would have kept measuring the old construct
through a scoring change. State has no such shared helper, so the whole
snapshot builder is the shared definition.
The sensor count matters because the score renormalises over whatever is
available: a session backed by two sensors is not drawn from the same
distribution as one backed by three, and the frozen threshold assumes it is.
``observations`` is the point-in-time fundamental series. It does not enter
either score -- the fundamental channel is categorical and read by confluence
-- but the per-session ``fundamental_state`` it produces is what the
confluence rule below is measured on, so it has to be the same series
production reports from. Every variant in this module reads its Warning from
these rows, so there is no second derivation to fall out of step.
"""
tickers = config["tickers"]
smh_full = prices.get(tickers["leaders"][0], [])
spy_full = prices.get(tickers["market"], [])
out: dict[date, float] = {}
backing: dict[date, int] = {}
rows: dict[date, dict] = {}
for session in dates:
sensors = rms.warning_sensor_scores(
breadth_divergence.get(session),
rms._closes_asof(smh_full, session),
rms._closes_asof(spy_full, session),
rms._window_asof(oas_series, session, rms.HY_OAS_WINDOW_DAYS),
snapshot = rms._compute_index(
prices,
vix_series,
oas_series,
{},
config,
session,
breadth_series=breadth_series,
divergence_series=divergence_series,
observations=observations or [],
)
score = rms.score_warning_sensors(sensors)
if score is not None:
out[session] = round(score, 2)
backing[session] = sum(1 for value in sensors.values() if value is not None)
return out, backing
state = snapshot["state"]
warning = snapshot["warning"]
rows[session] = {
"state": state.get("score"),
"warning": warning.get("score"),
"fundamental_state": (snapshot.get("fundamental_context") or {}).get("state"),
# `usable`, not `available`: a stale observation keeps its state for
# display but stops counting as evidence, and an observation whose
# extraction failed on everything is fresh but knows nothing. Either
# one counted here would inflate the covered window with sessions the
# channel could not have contributed to.
"fundamental_usable": bool(
(snapshot.get("fundamental_context") or {}).get("usable")
),
"state_coverage": state.get("coverage") or 0.0,
"warning_coverage": warning.get("coverage") or 0.0,
# The score renormalises over available sensors, so a session backed
# by two is not drawn from the same distribution as one backed by
# three, and a frozen threshold assumes it is.
"warning_sensors": len(warning.get("available_pillars") or []),
"inputs_fresh": bool((snapshot.get("data_quality") or {}).get("inputs_fresh")),
}
return rows
def _publishable(row: dict | None) -> bool:
"""What ``get_regime_history`` leaves for the alert to confirm against.
Deliberately not freshness-gated: ``_collect_regime_quadrant`` checks
``is_fresh`` on today's live reading only, while the prior session comes from
stored history where the only filter is a published band on both axes.
"""
return (
row is not None
and row["state"] is not None
and row["warning"] is not None
and row["state_coverage"] >= rms.MIN_COVERAGE
and row["warning_coverage"] >= rms.MIN_COVERAGE
)
def _prior_publishable(
rows: dict[date, dict], dates: list[date], index: int, history_days: int
) -> dict | None:
"""``valid[-2]``: the previous published session inside the 14-day window.
The monitor writes today's snapshot before the alert step runs
(``job_catalog._DAILY_PIPELINE_STEPS``), so ``valid[-1]`` is today and this
is genuinely the prior session rather than t-2.
"""
cutoff = dates[index] - timedelta(days=history_days)
for position in range(index - 1, -1, -1):
if dates[position] < cutoff:
return None
candidate = rows.get(dates[position])
if _publishable(candidate):
return candidate
return None
def replay_quadrant_changes(
rows: dict[date, dict],
dates: list[date],
state_divider: float = QUAD_X_DIV,
warning_divider: float = QUAD_Y_DIV,
margin: float = QUAD_MARGIN,
cooldown_days: int = QUAD_COOLDOWN_DAYS,
history_days: int = QUAD_HISTORY_DAYS,
) -> list[dict]:
"""Every quadrant change the shipped alert would have sent, in order.
A faithful replay of ``_collect_regime_quadrant``, including three details a
state machine written from first principles gets wrong:
* the prior session is classified against the *current baseline*, not against
its own predecessor, so confirmation asks "did yesterday already look like
this change" rather than "did yesterday change too";
* the baseline advances only when an alert actually fires, so a change that
fails confirmation or cooldown is re-evaluated against the old quadrant on
the next session rather than being forgotten;
* one cooldown is shared by every quadrant change, so a 3->4 alert can
swallow a 4->2 alert three days later.
Returns the fires themselves rather than alarm indices, because which
transitions count as a *warning* is the caller's question: entering
Warning-high territory and entering both-high territory are different rules
over the same replay.
"""
fires: list[dict] = []
baseline: str | None = None
baseline_date: date | None = None
for index, session in enumerate(dates):
row = rows.get(session)
if not _publishable(row) or not row["inputs_fresh"]:
continue
x, y = float(row["state"]), float(row["warning"])
if baseline is None: # seeds silently, exactly as a fresh install does
baseline = _classify_quadrant(x, y, None, margin, state_divider, warning_divider)
baseline_date = session
continue
new_quadrant = _classify_quadrant(x, y, baseline, margin, state_divider, warning_divider)
if new_quadrant == baseline:
continue
prior = _prior_publishable(rows, dates, index, history_days)
if prior is None:
continue
prior_quadrant = _classify_quadrant(
float(prior["state"]), float(prior["warning"]),
baseline, margin, state_divider, warning_divider,
)
if prior_quadrant != new_quadrant:
continue
if baseline_date is not None and (session - baseline_date).days < cooldown_days:
continue
fires.append({
"index": index,
"date": session.isoformat(),
"from": baseline,
"to": new_quadrant,
"state": x,
"warning": y,
})
baseline, baseline_date = new_quadrant, session
return fires
def entry_alarms(fires: list[dict], entry: tuple[str, ...]) -> list[int]:
"""Fires that *enter* the given quadrant set from outside it."""
return [f["index"] for f in fires if f["to"] in entry and f["from"] not in entry]
# ---------------------------------------------------------------------------
# Ablations, baselines, null
# ---------------------------------------------------------------------------
def _usable_adverse(rows: dict[date, dict], session: date) -> bool:
"""Adverse *and* still within its staleness horizon.
Both callers need this pair, and neither may use the state alone: the state
survives going stale so the card can show it, which would otherwise let a
months-old read confirm crossings indefinitely.
"""
row = rows.get(session) or {}
return row.get("fundamental_state") == "adverse" and bool(row.get("fundamental_usable"))
def adverse_episodes(
rows: dict[date, dict], dates: list[date], start_index: int
) -> list[int]:
"""Sessions where the fundamental state *becomes* usably adverse.
The market rules alarm on a rising-edge crossing; a categorical state has no
crossing, so its analogue is the transition into ``adverse``. That keeps the
row comparable with every other row in the table rather than counting every
day the state happens to sit there.
"""
alarms: list[int] = []
was_adverse = start_index > 0 and _usable_adverse(rows, dates[start_index - 1])
for index in range(start_index, len(dates)):
if dates[index] not in rows:
continue
adverse = _usable_adverse(rows, dates[index])
if adverse and not was_adverse:
alarms.append(index)
was_adverse = adverse
return alarms
def confluence_episodes(
warning_alarms: list[int], rows: dict[date, dict], dates: list[date]
) -> list[int]:
"""Warning crossings that happen while the fundamental state is usably adverse.
Deliberately gated on the market crossing rather than on either channel
moving: it preserves the rising-edge semantics every other row uses, so the
column measures "does requiring fundamental agreement help?" instead of a
differently-shaped rule that cannot be compared with the others.
"""
return [index for index in warning_alarms if _usable_adverse(rows, dates[index])]
def covered_events(
event_indices: list[int],
rows: dict[date, dict],
dates: list[date],
horizon: int,
) -> list[int]:
"""Corrections a fundamental rule actually had a chance to warn about.
An alarm counts only if it fires in ``[event - horizon, event - 1]``, so a
correction is *coverable* only if the channel had usable context somewhere in
that window. Scoring these rules against every correction instead would make
one day of observation render as 0/10 -- an untested rule reported as a
failed one, which is the exact mistake the ``measurable`` flag exists to
prevent for the empty-table case.
"""
covered: list[int] = []
for event_index in event_indices:
window = range(max(0, event_index - horizon), event_index)
if any(
bool((rows.get(dates[index]) or {}).get("fundamental_usable"))
for index in window
):
covered.append(event_index)
return covered
def eligible_sessions(
rows: dict[date, dict], dates: list[date], start_index: int
) -> int:
"""Sessions a fundamental rule could have fired on, for the FA/year rate.
Annualising over the whole window instead would divide a rule's false alarms
by years in which it was structurally incapable of firing, reporting a
flattering rate that means nothing.
"""
return sum(
1
for session in dates[start_index:]
if bool((rows.get(session) or {}).get("fundamental_usable"))
)
def below_average_series(
series: rms.Series, window: int = BASELINE_SMA_WINDOW
) -> dict[date, float]:
"""100 while the close sits under its ``window``-session average, else 0."""
out: dict[date, float] = {}
closes = [value for _, value in series]
for index, (session, close) in enumerate(series):
if index + 1 < window:
continue
average = sum(closes[index + 1 - window: index + 1]) / window
out[session] = 100.0 if close < average else 0.0
return out
def _null_model(
alarm_count: int,
event_indices: list[int],
dates: list[date],
horizon: int,
start_index: int,
observed_warned: int,
draws: int = NULL_DRAWS,
seed: int = NULL_SEED,
) -> dict | None:
"""Recall from alarms scattered at random over the same evaluable sessions.
Drawn only from sessions a real rule could have fired on: over the whole
sample the null would be diluted by warm-up sessions and would understate
what chance achieves. That matters here -- with ~11 events and a 20-session
horizon, a sixth of the sample already sits inside a hit window.
Corrections cluster, and uniform placement does not, so this is the floor
rather than the bar: an alarm process that clusters would beat it for
reasons that have nothing to do with foresight.
"""
population = range(start_index, len(dates))
if alarm_count <= 0 or not event_indices or alarm_count > len(population):
return None
rng = random.Random(seed)
recalls: list[int] = []
for _ in range(draws):
picks = sorted(rng.sample(population, alarm_count))
recalls.append(evaluate_alarms(picks, event_indices, dates, horizon)["events_warned"])
mean = sum(recalls) / len(recalls)
variance = sum((value - mean) ** 2 for value in recalls) / len(recalls)
return {
"draws": draws,
"alarms_per_draw": alarm_count,
"events": len(event_indices),
"mean_warned": round(mean, 2),
"sd_warned": round(variance ** 0.5, 2),
"observed_warned": observed_warned,
"p_at_least_observed": round(
sum(1 for value in recalls if value >= observed_warned) / len(recalls), 3
),
}
def _reliability(
@@ -192,9 +569,9 @@ def _reliability(
events_detected: int,
events_in_holdout: int,
) -> dict:
"""How far the headline metrics can actually be trusted.
"""How far the *fitted* variant's headline metrics can be trusted.
Two things repeatedly invite over-reading this report:
Two things repeatedly invite over-reading it:
* The holdout carries only the corrections that fall in the last 30% of the
sample. A "2/4" is one event away from "3/4", and in practice the events
@@ -203,6 +580,9 @@ def _reliability(
* The score renormalises over available sensors, so a training window that
predates a sensor's history freezes a threshold on a different construct
than the holdout is measured against.
Neither applies to the shipped rule, whose thresholds are fixed constants --
but the second one does not vanish, it relocates: see ``_era_split``.
"""
expected = len(rms.WARNING_WEIGHTS)
train = [backing[d] for d in dates[:split] if d in backing]
@@ -221,6 +601,126 @@ def _reliability(
}
def _era_split(
alarms: list[int],
event_indices: list[int],
dates: list[date],
horizon: int,
start_index: int,
credit_from: date | None,
) -> dict | None:
"""Shipped-rule metrics either side of the credit sensor's first session.
Dropping the fitted threshold makes the whole sample evaluable, which is the
point -- but most of the extra events sit before 2023-08, where W3 does not
exist and Warning renormalises to ``(W1*45 + W2*30)/75``. The fixed 40
divider is then applied to a different construct than it was reasoned about,
so the coverage caveat does not disappear with the split; it relocates from
the threshold to the score. Reporting the two eras separately is what keeps
the fuller sample from being a differently misleading headline.
The pre-credit era is close to a "Warning without W3" ablation on real
sessions -- and a clean one, because the fundamental channel is not a term in
Warning at all, so the two eras differ by W3 and nothing else. That stays
true however much fundamental history accumulates.
Alarms and events are assigned to eras by index, so an alarm days before the
boundary that matched an event days after it lands in the earlier era. With
the eras years long and the events sparse, that costs nothing.
"""
if credit_from is None:
return None
boundary = next(
(index for index, session in enumerate(dates) if session >= credit_from), None
)
if boundary is None or boundary <= start_index or boundary >= len(dates):
return None
def slice_metrics(low: int, high: int) -> dict:
sessions = max(0, high - low)
metrics = _score_rule(
[a for a in alarms if low <= a < high],
[e for e in event_indices if low <= e < high],
dates, horizon, sessions,
)
metrics.pop("per_event", None)
metrics["sessions"] = sessions
return metrics
return {
"credit_from": credit_from.isoformat(),
"pre_credit": {
"label": "W1+W2 only",
"start": dates[start_index].isoformat(),
"end": dates[boundary - 1].isoformat(),
**slice_metrics(start_index, boundary),
},
"full_coverage": {
"label": "all three sensors",
"start": dates[boundary].isoformat(),
"end": dates[-1].isoformat(),
**slice_metrics(boundary, len(dates)),
},
}
def _warning_from_rows(
rows: dict[date, dict], dates: list[date]
) -> tuple[dict[date, float], dict[date, int]]:
"""Published Warning per session plus how many sensors backed it.
Read off ``_axis_rows`` rather than recomputed. v2 re-derived Warning by hand
from ``WARNING_WEIGHTS`` and would have kept measuring the old construct
after a scoring change; a second derivation here would have done the same to
any later change to how Warning is assembled -- silently, in the fitted
variant and the ``warning_bare`` ablation, while the shipped replay moved on
without it.
"""
out: dict[date, float] = {}
backing: dict[date, int] = {}
for session in dates:
row = rows.get(session)
if row is None or row["warning"] is None:
continue
out[session] = float(row["warning"])
backing[session] = int(row["warning_sensors"])
return out, backing
def _rule_row(
rule_id: str,
label: str,
kind: str,
note: str,
alarms: list[int],
event_indices: list[int],
dates: list[date],
horizon: int,
sessions: int,
measurable: bool = True,
) -> dict:
"""One comparison row.
``measurable=False`` marks a rule whose *input* is too thin to have been
tested, not one that failed. A fundamental rule scores 0/N whether it is
wrong or merely absent, and a 0/N sitting in this table would read as
tested-and-failed -- the same false precision the whole restructure exists to
remove. It stays false until the channel has covered
``MIN_EVENTS_FOR_CONFIDENCE`` corrections, because a 1/1 or 0/2 over a
two-week exposure is not a result either.
"""
metrics = _score_rule(alarms, event_indices, dates, horizon, sessions)
metrics.pop("per_event", None)
return {
"id": rule_id,
"label": label,
"kind": kind,
"note": note,
"measurable": measurable,
**metrics,
}
async def run_event_study(
db: AsyncSession,
threshold_pct: float = EVENT_THRESHOLD_PCT,
@@ -242,55 +742,202 @@ async def run_event_study(
)
divergence = breadth_service.compute_divergence_series(breadth, benchmark)
oas_series = await rms._fetch_fred_series("BAMLH0A0HYM2", start, end)
warning, backing = _warning_series(prices, divergence, dates, config, oas_series)
# State needs volatility, which the Warning-only study never fetched.
vix_series = await rms._fetch_fred_series("VIXCLS", start, end)
# The point-in-time fundamental series. It is not in either score; it drives
# the categorical channel the confluence rule below is measured on.
observations = await rms.get_fundamental_observations(db)
# The credit sensor cannot reach back as far as the price history does (the
# upstream series is capped at ~3 years), so the earlier part of the sample
# scores on W1+W2 alone via renormalisation. Report where W3 starts rather
# than letting the threshold quietly straddle two sensor sets.
credit_from = oas_series[0][0].isoformat() if oas_series else None
credit_from = oas_series[0][0] if oas_series else None
all_events = detect_events(closes, dates, threshold_pct)
all_event_indices = [event["index"] for event in all_events]
# --- one pass; every rule below reads its Warning from these rows ----
rows = _axis_rows(
prices,
vix_series,
oas_series,
rms._mapping_series(breadth),
rms._mapping_series(divergence),
dates,
config,
observations,
)
warning, backing = _warning_from_rows(rows, dates)
fires = replay_quadrant_changes(rows, dates)
# Nothing can alarm before the baseline seeds, so every rule is measured from
# the same session and the comparison stays like-for-like.
seeded = next(
(
index
for index, session in enumerate(dates)
if _publishable(rows.get(session)) and rows[session]["inputs_fresh"]
),
None,
)
if seeded is None:
return {"available": False, "reason": "no session with publishable coverage"}
evaluable_start = seeded + 1
evaluable_sessions = max(1, len(dates) - evaluable_start)
evaluable_events = [index for index in all_event_indices if index >= evaluable_start]
warning_alarms = entry_alarms(fires, WARNING_QUADRANTS)
shipped_metrics = _score_rule(
warning_alarms, evaluable_events, dates, horizon, evaluable_sessions
)
shipped_events = shipped_metrics.pop("per_event")
# --- the fitted variant, kept for continuity -------------------------
split = max(1, min(len(dates) - 1, int(len(dates) * TRAIN_FRACTION)))
train_values = [warning[d] for d in dates[:split] if d in warning]
warn_threshold = _percentile(train_values, WARN_PERCENTILE)
if warn_threshold is None:
return {"available": False, "reason": "insufficient warning history"}
all_events = detect_events(closes, dates, threshold_pct)
holdout_events = [event["index"] for event in all_events if event["index"] >= split]
alarms = alarm_episodes(warning, dates, warn_threshold, start_index=split)
metrics = evaluate_alarms(alarms, holdout_events, dates, horizon)
holdout_events = [index for index in all_event_indices if index >= split]
fitted_alarms = alarm_episodes(warning, dates, warn_threshold, start_index=split)
holdout_sessions = max(1, len(dates) - split)
metrics["false_alarms_per_year"] = round(
metrics["false_alarms"] / (holdout_sessions / 252.0), 2
fitted_metrics = _score_rule(
fitted_alarms, holdout_events, dates, horizon, holdout_sessions
)
fitted_events = fitted_metrics.pop("per_event")
reliability = _reliability(dates, split, backing, len(all_events), len(holdout_events))
# --- ablations and baselines, all on fixed thresholds ----------------
# Fitted thresholds are deliberately excluded here: a threshold fitted on the
# full sample would have lookahead the shipped rule does not, and one fitted
# on a training split could only be scored on the four holdout events. Fixed
# constants keep every row on the same events over the same sessions.
state_series = {
session: row["state"] for session, row in rows.items() if row["state"] is not None
}
vix_indicator = {
session: value
for session in dates
if (value := rms._value_asof(vix_series, session)) is not None
}
# The fundamental channel is categorical and never enters a score, so it is
# compared as its own rule and as a confluence gate rather than tuned as a
# weight. With an empty observation series both are unmeasurable, and say so.
fundamental_alarms = adverse_episodes(rows, dates, evaluable_start)
confluence_alarms = confluence_episodes(warning_alarms, rows, dates)
# Coverage-matched denominators. These rules only existed on the sessions the
# channel had usable context, so scoring them over the whole window would
# report an exposure they never had -- and one day of coverage would render
# as 0/10.
fundamental_events = covered_events(evaluable_events, rows, dates, horizon)
fundamental_sessions = eligible_sessions(rows, dates, evaluable_start)
fundamental_measurable = len(fundamental_events) >= MIN_EVENTS_FOR_CONFIDENCE
comparison = [
_rule_row(
"fundamental_adverse", "Fundamental context turns adverse", "fundamental",
"The third channel on its own: transitions into an adverse capex / "
"earnings-reaction state, with no market input at all.",
fundamental_alarms, fundamental_events, dates, horizon, fundamental_sessions,
measurable=fundamental_measurable,
),
_rule_row(
"confluence", "Confluence: Warning crossing while adverse", "fundamental",
"The shipped market crossing, kept only when the fundamental channel "
"agrees. Answers whether requiring agreement buys precision, at what "
"cost in recall.",
confluence_alarms, fundamental_events, dates, horizon, fundamental_sessions,
measurable=fundamental_measurable,
),
_rule_row(
"market_over_covered", "Quadrant alert, covered window only", "fundamental",
"The shipped market rule scored on exactly the events, sessions and "
"alarms the two rows above were scored on. Without it, any difference "
"between them and the headline could be the window rather than the "
"channel.",
# Alarms are restricted to the covered window too: counting crossings
# that fired when the channel had no context would compare the market
# rule's full exposure against the channel's partial one.
[
index
for index in warning_alarms
if index >= evaluable_start
and bool((rows.get(dates[index]) or {}).get("fundamental_usable"))
],
fundamental_events, dates, horizon, fundamental_sessions,
measurable=fundamental_measurable,
),
_rule_row(
"quadrant_stress_entry", "Quadrant alert, both axes high", "ablation",
"The same replay, recording only entries into the both-high quadrant. "
"State is coincident by construction, so requiring it should convert "
"leads into confirmations.",
entry_alarms(fires, STRESS_QUADRANT),
evaluable_events, dates, horizon, evaluable_sessions,
),
_rule_row(
"warning_bare", f"Warning >= {QUAD_Y_DIV:.0f} (bare crossing)", "ablation",
"The shipped divider with none of the quadrant machinery: no State "
"condition, no hysteresis, no confirmation, no cooldown.",
alarm_episodes(warning, dates, QUAD_Y_DIV, start_index=evaluable_start),
evaluable_events, dates, horizon, evaluable_sessions,
),
_rule_row(
"state_bare", f"State >= {QUAD_X_DIV:.0f} (bare crossing)", "ablation",
"The coincident axis alone. State measures stress that has already "
"arrived, so a competitive lead here would be surprising.",
alarm_episodes(state_series, dates, QUAD_X_DIV, start_index=evaluable_start),
evaluable_events, dates, horizon, evaluable_sessions,
),
_rule_row(
"smh_below_50dma", f"{leader} below its {BASELINE_SMA_WINDOW}-DMA", "baseline",
"The crudest possible trend rule, and free.",
alarm_episodes(
below_average_series(benchmark, BASELINE_SMA_WINDOW), dates,
50.0, start_index=evaluable_start,
),
evaluable_events, dates, horizon, evaluable_sessions,
),
_rule_row(
"vix_level", f"VIX >= {BASELINE_VIX_LEVEL:.0f}", "baseline",
"The market's own risk gauge, unweighted and unmodelled.",
alarm_episodes(
vix_indicator, dates, BASELINE_VIX_LEVEL, start_index=evaluable_start
),
evaluable_events, dates, horizon, evaluable_sessions,
),
]
null_model = _null_model(
len(warning_alarms), evaluable_events, dates, horizon,
evaluable_start, shipped_metrics["events_warned"],
# Passed rather than defaulted: a default argument binds the constant at
# import, so overriding it (in tests) would silently do nothing.
draws=NULL_DRAWS, seed=NULL_SEED,
)
eras = _era_split(
warning_alarms, evaluable_events, dates, horizon, evaluable_start, credit_from
)
basket_asof = date.fromisoformat(config["basket_asof"])
retrospective = dates[split] < basket_asof
retrospective = dates[evaluable_start] < basket_asof
evaluation = "exploratory" if retrospective else "holdout"
lead_text = (
f"median lead {metrics['median_lead_days']:.0f} sessions"
if metrics["median_lead_days"] is not None
f"median lead {shipped_metrics['median_lead_days']:.0f} sessions"
if shipped_metrics["median_lead_days"] is not None
else "no successful warning lead"
)
summary = (
f"{evaluation.capitalize()} chronological test: warning episodes preceded "
f"{metrics['events_warned']}/{metrics['events']} 10% corrections; "
f"{metrics['events_missed']} missed, {metrics['false_alarms_per_year']:.1f} "
f"false alarms/year, {lead_text}. "
f"{metrics['events']} of {reliability['events_detected']} detected corrections "
f"fall in the test period"
+ (
"; too few to read recall as a property of the score."
if reliability["underpowered"]
else "."
f"{evaluation.capitalize()} replay of the shipped quadrant alert over "
f"{evaluable_sessions} sessions: it entered Warning-high territory ahead of "
f"{shipped_metrics['events_warned']} of {shipped_metrics['events']} 10% "
f"corrections, with {shipped_metrics['false_alarms_per_year']:.1f} false "
f"alarms/year and {lead_text}. Its dividers are fixed constants rather than "
f"fitted, so there is no training split and every detected correction is "
f"evaluable — compare it against the ablations and baselines below before "
f"reading the ratio as good or bad."
)
)
per_event = metrics.pop("per_event")
report = {
"available": True,
"schema": STUDY_SCHEMA,
"methodology": rms.METHODOLOGY,
"generated_at": datetime.now(timezone.utc).isoformat(),
"evaluation": evaluation,
@@ -301,24 +948,69 @@ async def run_event_study(
"event_threshold_pct": threshold_pct,
"event_cooldown_days": EVENT_COOLDOWN_DAYS,
"horizon_days": horizon,
"train_fraction": TRAIN_FRACTION,
"warn_percentile": WARN_PERCENTILE,
"warn_threshold": round(warn_threshold, 1),
"credit_sensor_from": credit_from,
"credit_sensor_from": credit_from.isoformat() if credit_from else None,
"basket_hash": rms._basket_hash(config["breadth_basket"]),
"basket_asof": config["basket_asof"],
},
# The channel's actual exposure, which is what its rows are scored on.
# The series starts empty -- the observation lived in a single
# overwritten settings slot until 2026-08-12 -- and it accumulates one
# observation at a time, so for a long while these rows are unmeasurable
# rather than unsuccessful. Stating the exposure is what stops the table
# inventing a failed result out of a thin one.
"fundamental_coverage": {
"observations": len(observations),
"sessions_eligible": fundamental_sessions,
"evaluable_sessions": evaluable_sessions,
"events_covered": len(fundamental_events),
"events_evaluable": len(evaluable_events),
"minimum_events": MIN_EVENTS_FOR_CONFIDENCE,
"measurable": fundamental_measurable,
},
"sample": {
"start": dates[0].isoformat(),
"end": dates[-1].isoformat(),
"sessions": len(dates),
# Not "test_start": the shipped rule fits nothing, so this is where
# the baseline seeds and every rule becomes measurable, not where a
# holdout begins. The fitted variant's split lives under "fitted".
"evaluable_from": dates[evaluable_start].isoformat(),
"evaluable_sessions": evaluable_sessions,
"events_detected": len(all_events),
"events_evaluable": len(evaluable_events),
},
"shipped": {
"rule": {
"state_divider": QUAD_X_DIV,
"warning_divider": QUAD_Y_DIV,
"margin": QUAD_MARGIN,
"confirm_sessions": 2,
"cooldown_days": QUAD_COOLDOWN_DAYS,
"entry": "Warning-high quadrant (early warning or active stress)",
},
"metrics": shipped_metrics,
"events": shipped_events,
"quadrant_changes": len(fires),
"fires": fires,
"by_era": eras,
},
"comparison": comparison,
"null_model": null_model,
"fitted": {
"params": {
"train_fraction": TRAIN_FRACTION,
"warn_percentile": WARN_PERCENTILE,
"warn_threshold": round(warn_threshold, 1),
},
"sample": {
"train_end": dates[split - 1].isoformat(),
"test_start": dates[split].isoformat(),
"sessions": len(dates),
"holdout_sessions": holdout_sessions,
},
"metrics": metrics,
"metrics": fitted_metrics,
"events": fitted_events,
},
"reliability": reliability,
"events": per_event,
"recent_breadth": [
{"date": d.isoformat(), "breadth": breadth[d], "warning": warning.get(d)}
for d in dates[-90:]
@@ -328,10 +1020,13 @@ async def run_event_study(
logger.info(json.dumps({
"event": "regime_event_study_complete",
"evaluation": evaluation,
"events": metrics["events"],
"events_detected": reliability["events_detected"],
"warned": metrics["events_warned"],
"false_alarms_per_year": metrics["false_alarms_per_year"],
"shipped_events": shipped_metrics["events"],
"shipped_warned": shipped_metrics["events_warned"],
"shipped_false_alarms_per_year": shipped_metrics["false_alarms_per_year"],
"quadrant_changes": len(fires),
"fitted_events": fitted_metrics["events"],
"fitted_warned": fitted_metrics["events_warned"],
"null_p_at_least_observed": (null_model or {}).get("p_at_least_observed"),
"underpowered": reliability["underpowered"],
"sensor_coverage_mismatch": reliability["sensor_coverage_mismatch"],
}))
@@ -352,4 +1047,8 @@ async def get_event_study_report(db: AsyncSession) -> dict | None:
report = json.loads(setting.value)
except (TypeError, ValueError):
return None
return report if report.get("methodology") == rms.METHODOLOGY else None
if report.get("methodology") != rms.METHODOLOGY:
return None
# A pre-replay report parses fine and carries the current methodology, so the
# shape has to be checked separately or the panel renders a headline-less v4.
return report if report.get("schema") == STUDY_SCHEMA else None
@@ -22,6 +22,7 @@ from app.models.fundamental_snapshot import FundamentalSnapshot
from app.models.ohlcv import OHLCVRecord
from app.models.ticker import Ticker
from app.services import fundamentals_derivation as deriv
from app.services import ticker_service
@dataclass(frozen=True)
@@ -46,7 +47,11 @@ async def build_candidates(
"""Derive current cache candidates using only already-stored data."""
today = today or datetime.now(ZoneInfo("America/New_York")).date()
tickers = list(
(await db.execute(select(Ticker).order_by(Ticker.symbol))).scalars()
(
await db.execute(
ticker_service.active_only(select(Ticker).order_by(Ticker.symbol))
)
).scalars()
)
if not tickers:
return []
+62 -3
View File
@@ -3,7 +3,9 @@
from __future__ import annotations
import json
from collections import defaultdict
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from sqlalchemy import exists, func, select
from sqlalchemy.ext.asyncio import AsyncSession
@@ -15,6 +17,23 @@ from app.models.ticker import Ticker
_SEC_FORMS = ("10-K", "10-Q", "10-K/A", "10-Q/A")
# How recent the issuer's own newest filing must be for an *escalated* gap to
# stop pausing setups. A gap pauses an issuer until it is either resolved or
# superseded by a later ingested filing — which assumes the gap is temporary.
# It is not always: SEC's per-company Company-Facts files can go stale
# indefinitely (2026-08, 43 large caps whose Q2 10-Qs the frames API carried but
# whose companyfacts files never received), and since the supersede rule needs a
# *successfully ingested* later filing, a stale file also swallows the next
# quarter. The pause is then open-ended rather than seasonal.
#
# So the pause hands off to the alert: once `filing_gap_aged` has escalated a gap
# to an operator (`escalated_at`), the issuer resumes on the fundamentals it does
# have — provided those are recent. An issuer with nothing this fresh has no
# usable fundamentals at all and stays paused, which is the case the gate was
# built for. The retry queue is untouched: `active_gaps` still returns these, so
# the importer keeps retrying and a recovered filing still resolves normally.
GAP_GATE_RECENT_FILING_DAYS = 180
@dataclass(frozen=True)
class SetupQuality:
@@ -51,6 +70,42 @@ async def active_gaps(
return list((await db.execute(stmt)).scalars().all())
async def gap_exempt_ciks(
db: AsyncSession, gaps: list[SecFilingGap]
) -> set[str]:
"""CIKs whose gaps have stopped pausing setups (see GAP_GATE_RECENT_FILING_DAYS).
Every one of a CIK's active gaps must be escalated: one fresh gap alongside an
old one still means a filing we might yet ingest, which is worth pausing for.
Public because the importer alerts on this exact transition (a CIK dropping
out of this set is a pause coming back on) and the rule must not exist twice.
"""
by_cik: dict[str, list[SecFilingGap]] = defaultdict(list)
for gap in gaps:
by_cik[gap.cik].append(gap)
escalated = {
cik
for cik, items in by_cik.items()
if all(gap.escalated_at is not None for gap in items)
}
if not escalated:
return set()
cutoff = (
datetime.now(timezone.utc) - timedelta(days=GAP_GATE_RECENT_FILING_DAYS)
).date()
rows = await db.execute(
select(FundamentalSnapshot.cik)
.where(
FundamentalSnapshot.cik.in_(escalated),
FundamentalSnapshot.form.in_(_SEC_FORMS),
FundamentalSnapshot.filed_date >= cutoff,
)
.distinct()
)
return set(rows.scalars())
async def _latest_validation(db: AsyncSession) -> dict:
payload = (
await db.execute(
@@ -80,8 +135,12 @@ async def blocked_reasons_by_cik(
if ciks is not None and not ciks:
return {}
gaps = await active_gaps(db, ciks)
# Escalated gaps on issuers that still have recent fundamentals no longer
# pause setups, on either path below — the summary mirrors the same filings.
exempt = await gap_exempt_ciks(db, gaps)
reasons = {
gap.cik: "sec_filing_gap" for gap in await active_gaps(db, ciks)
gap.cik: "sec_filing_gap" for gap in gaps if gap.cik not in exempt
}
summary = await _latest_validation(db)
@@ -92,11 +151,11 @@ async def blocked_reasons_by_cik(
# stay capped for audit readability. Detailed entries supply the reason.
for cik in summary.get("setup_blocked_ciks") or []:
normalized = str(cik) if cik else ""
if normalized and wanted(normalized):
if normalized and wanted(normalized) and normalized not in exempt:
reasons.setdefault(normalized, "sec_filing_gap")
for item in summary.get("missing_xbrl") or []:
normalized = str(item.get("cik") or "")
if normalized and wanted(normalized):
if normalized and wanted(normalized) and normalized not in exempt:
reasons.setdefault(normalized, "sec_filing_gap")
for cik in summary.get("no_xbrl_ciks") or []:
normalized = str(cik) if cik else ""
+91
View File
@@ -0,0 +1,91 @@
"""Single source for JobRunState reads/writes.
Mirrors ``settings_store``: ``record_finish`` never commits the caller owns
the transaction and reads are batched so the admin listing stays one query.
"""
from __future__ import annotations
import logging
from collections.abc import Iterable
from datetime import datetime, timezone
from sqlalchemy import select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.dialects.sqlite import insert as sqlite_insert
from sqlalchemy.ext.asyncio import AsyncSession
from app.models.job_run_state import JobRunState
logger = logging.getLogger(__name__)
def _as_datetime(value: object) -> datetime | None:
"""Runtime snapshots carry ISO strings; the column wants a datetime."""
if isinstance(value, datetime):
return value
if isinstance(value, str) and value:
try:
return datetime.fromisoformat(value)
except ValueError:
return None
return None
async def get_map(db: AsyncSession, job_names: Iterable[str]) -> dict[str, JobRunState]:
"""Return {job_name: row} for the given jobs that have ever finished.
``populate_existing`` because rows are written by core upserts, which leave
any previously-loaded ORM instance in the identity map stale.
"""
result = await db.execute(
select(JobRunState)
.where(JobRunState.job_name.in_(list(job_names)))
.execution_options(populate_existing=True)
)
return {row.job_name: row for row in result.scalars().all()}
def _insert_for(db: AsyncSession):
"""ON CONFLICT is dialect-specific; prod is Postgres, tests are SQLite."""
dialect = db.get_bind().dialect.name
return pg_insert if dialect == "postgresql" else sqlite_insert
async def record_finish(db: AsyncSession, job_name: str, runtime: dict) -> None:
"""Upsert the last-run row from a scheduler runtime snapshot.
Atomic, and newer-wins. Select-then-insert loses races that really happen
here: pipelines are separate scheduler jobs that can overlap, and they share
step ids -- data_collector belongs to all four. Two of them finishing that
step together would both see no row and both insert, and the loser's
IntegrityError is swallowed by the caller, so the run silently vanishes.
The ``where`` guard is the other half: without it a slower pipeline
finishing an *older* run last would rewind finished_at and the status with
it, so the panel would report a stale outcome as the latest one.
"""
finished_at = _as_datetime(runtime.get("finished_at")) or datetime.now(timezone.utc)
message = runtime.get("message")
now = datetime.now(timezone.utc)
values = {
"job_name": job_name,
"status": str(runtime.get("status") or "completed"),
"started_at": _as_datetime(runtime.get("started_at")),
"finished_at": finished_at,
"processed": runtime.get("processed"),
"total": runtime.get("total"),
"message": str(message)[:4000] if message else None,
# Set explicitly: the model's onupdate hook does not fire for a core
# INSERT ... ON CONFLICT DO UPDATE.
"updated_at": now,
}
statement = _insert_for(db)(JobRunState).values(**values)
await db.execute(
statement.on_conflict_do_update(
index_elements=[JobRunState.job_name],
set_={key: statement.excluded[key] for key in values if key != "job_name"},
where=JobRunState.finished_at < statement.excluded.finished_at,
)
)
+4 -1
View File
@@ -18,6 +18,7 @@ from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from app.models.ticker import Ticker
from app.services import ticker_service
from app.services.price_service import query_ohlcv
logger = logging.getLogger(__name__)
@@ -169,7 +170,9 @@ async def compute_activation_ranks(db: AsyncSession) -> dict[str, dict[str, floa
before scanning; the research backtest ranked each weekly setup-candidate
cross-section, so this is the deliberate production approximation.
"""
result = await db.execute(select(Ticker).order_by(Ticker.symbol))
result = await db.execute(
ticker_service.active_only(select(Ticker).order_by(Ticker.symbol))
)
tickers = list(result.scalars().all())
benchmark_closes = await _load_activation_benchmark(db)
+492 -43
View File
@@ -1,4 +1,4 @@
"""AI/Tech Risk Monitor v3.
"""AI/Tech Risk Monitor v4.
The monitor is a risk thermometer, not a probability or trading rule. It keeps
two deliberately separate outputs:
@@ -7,14 +7,20 @@ two deliberately separate outputs:
* Warning: deterioration/divergence that may precede State (breadth divergence,
relative strength, credit impulse).
Both scores are quantitative and daily. The sourced hyperscaler capex and
earnings-reaction observations are a qualitative *overlay* in v3 rather than
weighted sensors: at a combined 20 points they could not reach the event
study's alarm threshold even when both pegged, so refreshing them appeared to
do nothing. They are reported next to the scores instead of inside them.
* Fundamental context: a categorical channel (supportive / neutral / adverse /
unknown) with an evidence-quality grade, derived by fixed rules from the
sourced hyperscaler capex and earnings-reaction observations.
Both scores are quantitative and daily. The fundamental channel is deliberately
**not** a term in either: the three are read together by confluence, because
adding a slow categorical judgement to a fast continuous score manufactures
precision by summing unlike things, and any fusion weight would be a policy
preference presented as a measurement until there is enough point-in-time
history to fit one. A missing observation therefore stays ``unknown`` instead of
silently redistributing its weight onto the technical sensors.
Daily snapshots are the point-in-time record. The first run under a new
``METHODOLOGY`` rewrites the latest ``REBUILD_SESSIONS`` trading sessions once;
``METHODOLOGY`` rewrites every session inside ``REBUILD_LOOKBACK_DAYS`` once;
ordinary runs thereafter only upsert the latest trading date. The overlay is
still gated by its effective date so a rebuild cannot stamp today's observation
onto historical snapshots.
@@ -35,6 +41,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.config import settings
from app.exceptions import ProviderError, ValidationError
from app.models.regime_fundamental_observation import RegimeFundamentalObservation
from app.models.regime_snapshot import RegimeSnapshot
from app.providers.alpaca import AlpacaOHLCVProvider
from app.services import breadth_service, settings_store
@@ -48,10 +55,18 @@ _CA_BUNDLE = os.environ.get("SSL_CERT_FILE", "")
KEY_CONFIG = "regime_monitor_config"
KEY_FUNDAMENTALS = "regime_fundamental_overrides"
METHODOLOGY = "v3"
METHODOLOGY = "v4"
# Snapshots are reseeded on a methodology bump, but fundamental observations are
# collected by hand/LLM and carried across it when the format is compatible.
CATEGORICAL_FUNDAMENTAL_METHODOLOGIES = frozenset({"v2", "v3"})
# EVERY methodology sharing the categorical format must be listed: this is checked
# against the *stored* blob, so omitting the current one discards the observation
# on its first write, which leaves fetched_at null and locked false -- and then
# update_regime_monitor refreshes it via the LLM on every single run, forever.
# "v5" is listed although no v5 scoring exists: a v5 was briefly built (a weighted
# fundamental modifier on Warning) and reverted, so a development box can have
# that string sitting in its settings blob. Keeping it costs nothing; omitting it
# costs the failure above.
CATEGORICAL_FUNDAMENTAL_METHODOLOGIES = frozenset({"v2", "v3", "v4", "v5"})
# Bumped when a fix changes what historical rows *should* contain without
# changing the live formula, so stored history needs one reseed. Deliberately
@@ -59,6 +74,10 @@ CATEGORICAL_FUNDAMENTAL_METHODOLOGIES = frozenset({"v2", "v3"})
# study, neither of which is warranted here -- the study recomputes its Warning
# series from source rather than reading snapshots, so a reseed cannot stale it.
# Snapshots written before this marker existed carry no key and read as 1.
# Deliberately NOT bumped for v4: a METHODOLOGY change already forces a full
# reseed (every stored row fails _parse_snapshot, so _latest_snapshot_row returns
# None and rebuilding is True). Bumping both would imply the reseed was
# revision-driven.
SENSOR_REVISION = 2
MIN_COVERAGE = 75.0
SOURCE_MAX_LAG_DAYS = 7
@@ -68,9 +87,18 @@ SOURCE_MAX_LAG_DAYS = 7
# exceeded 64.9 in 408 sessions while State reached 91.2). Thresholds are round
# numbers chosen so each band covers a sane share of history, not percentile
# fits -- percentile-derived bands would drift on every rebuild and silently
# rewrite what past snapshots meant. Realized shares over the 408 sessions to
# 2026-07-24: State 73/15/8/3%, Warning 69/20/8/3%.
STATE_BANDS = (20.0, 50.0, 80.0)
# rewrite what past snapshots meant.
#
# v4 moved State's top band 80 -> 65, and only that one. With credit calm it
# scores 0.0 (not None) and still holds its full 20 points, so price + breadth +
# volatility at *literal maximum* summed to exactly 80.0 -- the old threshold, to
# the decimal, with nothing to spare. A 2022-style AI/tech drawdown with calm
# credit computes to 70.3-74.0 depending on whether a death cross has formed, so
# at 80 the case this monitor exists to measure could not print the top band.
# 65 clears it under either assumption. Realized shares over the 408 sessions to
# 2026-07-24, reported not fitted: State 78.9/13.0/4.7/3.4%, Warning 69/20/8/3%.
# The v4 breaking share (3.4%) matches v3's, which was arrived at independently.
STATE_BANDS = (20.0, 50.0, 65.0)
WARNING_BANDS = (20.0, 40.0, 60.0)
QUADRANT_STATE_DIVIDER = 50.0
@@ -118,6 +146,27 @@ P3_DRAWDOWN_ANCHORS = (
(0.0, 0.0), (4.0, 10.0), (8.0, 25.0), (16.0, 50.0), (28.0, 78.0), (40.0, 100.0),
)
# Trend-break depth (% below the 200-DMA, stress score). v4; see _under_200 for
# why the crossing gets a floor of 20 rather than starting at 0. Calibrated to
# sit alongside P3 rather than swamp it -- the 200-DMA lags, so a 20% drawdown
# typically coincides with ~10% below the average, where this reads ~61 against
# P3's ~59. On the population the P1_SCORE_CAP rule actually names -- sessions
# with State >= 40 -- P1 is the sole price argmax on 17 of 47 (36.2%), against
# P2's 16 and P3's 14, so it informs the pillar without owning it and no cap
# was needed.
P1_TREND_BREAK_ANCHORS = (
(0.0, 20.0), (3.0, 35.0), (8.0, 55.0), (15.0, 75.0), (25.0, 100.0),
)
# VIX level anchors (v4). Full scale at 55 rather than at 2020's ~82: anchoring
# the top at a once-in-a-generation print would make VIX 50 -- a genuine crisis
# -- read only ~70. A typical correction (25-35) now reads 38-67 where v3 read
# 66.7-100. The anchors encode the long-run distribution as constants, the same
# argument the credit level uses.
P5_VIX_ANCHORS = (
(15.0, 0.0), (20.0, 20.0), (25.0, 38.0), (30.0, 55.0), (40.0, 80.0), (55.0, 100.0),
)
STATE_WEIGHTS = {
"price": 40.0,
"breadth": 25.0,
@@ -135,6 +184,34 @@ WARNING_WEIGHTS = {
"credit_impulse": 25.0,
}
# The sourced fundamental read is a **separate channel**, never a term in either
# score. It is reported as a categorical state beside State and Warning, and the
# three are read together by confluence rather than added up.
#
# Two things had to be true at once and only this shape gets both.
#
# **v3's reason for removing it was wrong.** v3 argued that F1+F3, at 12+8 of 100
# Warning points, "could not change any published conclusion" because pegged they
# produced a Warning of exactly 20.0. That holds only when every technical sensor
# reads exactly zero. Weighted, those points added +10 to +20 across the
# realistic range and moved the technical score needed to reach the 40 quadrant
# divider from 40 to 25. So the observation was not inert, and demoting it to
# decoration was not justified by that argument.
#
# **But no weight is measurable either.** A weighted modifier was built (v5,
# reverted) and its size could not be derived from anything: with ~10 correction
# events and essentially no fundamental history, any fusion weight is a policy
# preference presented as a measurement. Adding a slow categorical judgement to a
# fast continuous score also manufactures precision by summing unlike things, and
# it forces a missing observation to silently redistribute its weight onto the
# technical sensors -- the opposite of leaving it unknown.
#
# So the read gets a channel, not a coefficient. Revisit only with enough
# point-in-time history to test whether the state improves prediction
# *conditional on* Warning; a fitted model then has something to fit.
FUNDAMENTAL_STATES = ("supportive", "neutral", "adverse", "unknown")
EVIDENCE_QUALITY = ("complete", "partial", "stale", "manual", "unavailable")
# Fixed at the v2 launch. These are liquid S&P 500/Nasdaq AI, semiconductor,
# infrastructure, cloud, and enterprise-software names that the platform's
# normal universe sync already stores.
@@ -158,7 +235,12 @@ DEFAULT_CONFIG: dict = {
}
CAPEX_STATES = ("raising", "holding", "cutting", "unknown")
GNSD_STATES = ("yes", "no", "mixed")
# "mixed" is a genuinely observed mixed reaction; "unknown" is nobody looked or
# the extraction failed. They were the same value until 2026-08-13, so a failed
# LLM parse silently became neutral *evidence* -- an observation of normality
# manufactured out of a parse error. Same distinction the capex map already made
# with its own "unknown", and the same one the whole channel is built on.
GNSD_STATES = ("yes", "no", "mixed", "unknown")
# v2 scored raising and holding identically at 0, so in a capex boom the reading
# was pinned at 0 and could not express the raising -> holding deceleration that
# is the actual early warning. Display-only in v3, but it should still describe.
@@ -219,10 +301,24 @@ def band_for(score: float, bands: tuple[float, float, float] = STATE_BANDS) -> s
def _under_200(closes: list[float]) -> float | None:
"""Trend break graded by depth below the 200-DMA, not a bare yes/no.
Through v3 this returned 0 or 100, so P1 printed 100 the moment SMH and QQQ
were both under their average -- and because the price pillar takes
``max(P1, P2, P3)``, that pinned the pillar and stopped P3's anchored ladder
resolving anything for the whole of a selloff. It pegged on 46 of the 408
sessions to 2026-07-24; under this table, none.
The step at the crossing (0 -> 20) is deliberate: the break itself is a
genuine binary event and deserves a floor. Only the depth past it is graded.
"""
sma200 = _sma(closes, 200)
if sma200 is None:
if sma200 is None or sma200 <= 0:
return None
return 100.0 if closes[-1] < sma200 else 0.0
pct_below = (sma200 - closes[-1]) / sma200 * 100.0
if pct_below <= 0:
return 0.0
return _clamp(_interpolate(pct_below, P1_TREND_BREAK_ANCHORS))
def p1_trend_break(smh: list[float], qqq: list[float], leader_weight: float = 2.0) -> float | None:
@@ -288,9 +384,17 @@ def p4_relative_strength(smh: list[float], spy: list[float], lookback: int = 60)
def p5_volatility(vix: float | None) -> float | None:
"""VIX level against named anchors, so it keeps resolving past a 30 print.
v3 used ``(vix - 15) / 15``, which reached 100 at VIX 30 -- the same
saturation v3 itself had just removed from P3. VIX 30 is a bad week, 50 is a
crisis and 82 was March 2020, and all three scored identically. In the 408
sessions to 2026-07-24 that flattened five distinct April-2025 prints
(52.33, 46.98, 45.31, 40.72, 38.57) into a single 100.
"""
if vix is None:
return None
return _clamp((vix - 15.0) / 15.0 * 100.0)
return _clamp(_interpolate(vix, P5_VIX_ANCHORS))
def breadth_level_score(pct_above_200: float | None) -> float | None:
@@ -375,6 +479,104 @@ def score_warning_sensors(sensors: dict[str, float | None]) -> float | None:
return sum(s * w for s, w in live) / sum(w for _, w in live)
def _capex_signal(capex: dict[str, str] | None, names: list[str]) -> str:
"""Categorical read of hyperscaler capex direction. Never an average.
Averaging is what this must not do: it would let two ``cutting`` reads and
two ``unknown`` ones land on "neutral", presenting missing evidence as
evidence of normality. Any cut is adverse on partial evidence; only a fully
known, uniformly rising basket is supportive.
"""
states = [str((capex or {}).get(name, "unknown")).strip().lower() for name in names]
known = [state for state in states if state in ("raising", "holding", "cutting")]
if not known:
return "unknown"
if "cutting" in known:
return "adverse"
if "holding" in known:
return "neutral"
return "supportive"
def _reaction_signal(good_news_stock_down: str | None) -> str:
"""Good earnings being sold is a late-cycle tell; not being sold is healthy.
Anything that is not one of the three observed categories -- including the
explicit ``"unknown"`` an extraction failure now writes -- falls through to
``unknown`` rather than to ``mixed``. A parse error is not a reading.
"""
return {
"yes": "adverse",
"no": "supportive",
"mixed": "neutral",
}.get(str(good_news_stock_down or "").strip().lower(), "unknown")
def combine_fundamental_signals(capex_signal: str, reaction_signal: str) -> str:
"""Confluence, not arithmetic: precedence over the two categorical reads.
``unknown`` is deliberately unreachable by combination -- it survives only
when *nothing* was observed. A single adverse read carries, because partial
evidence of deterioration is still evidence of deterioration; supportive
requires every observed signal to agree.
"""
signals = (capex_signal, reaction_signal)
if "adverse" in signals:
return "adverse"
observed = [signal for signal in signals if signal != "unknown"]
if not observed:
return "unknown"
return "supportive" if all(signal == "supportive" for signal in observed) else "neutral"
def _usable_context(observed: bool, pending: bool, stale: bool, state: str) -> bool:
"""Whether a fundamental reading may count as evidence.
One definition, called by both the point-in-time record and the live
reading, because they publish the same field name to the same consumers and
a second copy would drift. Distinct from `available`, which is about timing
alone: an observation whose extraction failed on everything is effective and
fresh, and still knows nothing.
"""
return observed and not pending and not stale and state != "unknown"
def _evidence_quality(
capex: dict[str, str] | None,
good_news_stock_down: str | None,
names: list[str],
*,
observed: bool,
stale: bool,
source: str | None,
) -> str:
"""How much to trust the state above, as one field the reader can act on.
Ordered by what an operator most needs to know: nothing collected beats
everything else, then a reading too old to be current, then a hand override,
then completeness.
"""
if not observed:
return "unavailable"
if stale:
return "stale"
if str(source or "").strip().lower() == "manual":
return "manual"
known = sum(
1
for name in names
if str((capex or {}).get(name, "unknown")).strip().lower() != "unknown"
)
# `bool(names)` matters: with an empty basket `known == len(names)` is
# vacuously true, so nothing observed would grade as complete.
complete = (
bool(names)
and known == len(names)
and _reaction_signal(good_news_stock_down) != "unknown"
)
return "complete" if complete else "partial"
def _sensor(sensor_id: str, label: str, score: float | None, **details: object) -> dict:
return {
"id": sensor_id,
@@ -512,26 +714,66 @@ def _overlay_timing(
return effective, pending, age, stale
def fundamental_overlay(overrides: dict, config: dict, as_of: date) -> dict:
"""Point-in-time qualitative overlay. Never feeds State or Warning in v3.
def fundamental_context(overrides: dict, config: dict, as_of: date) -> dict:
"""Point-in-time fundamental channel. Never a term in State or Warning.
Called an "overlay" until 2026-08-12, which undersold it: it is the third
channel of the model, read alongside the two scores by confluence rather than
decorating them. The categorical ``state`` is what a reader and the chart
consume; ``evidence_quality`` is how far to trust it.
Both are derived from the stored categorical facts by fixed rules, not from
an LLM's numeric judgement. The LLM's job is extraction and explanation --
find the capex guidance, classify it, cite it -- and the rules turn those
facts into a state, so the same observation always yields the same category.
The effective-date gate stays even though nothing is scored from this: the
400-session rebuild replays historical dates, and stamping today's LLM read
onto 2024 snapshots would be plain lookahead in the stored record.
rebuild replays historical dates, and stamping today's read onto 2024
snapshots would be plain lookahead in the stored record.
This is the *record*. For "what do we know right now", use
``current_observation`` -- do not add a bypass flag here, because this runs
for every replayed date during a rebuild.
"""
effective, pending, age, stale = _overlay_timing(overrides, config, as_of)
names = list(config["tickers"]["hyperscalers"])
capex = None if pending else overrides.get("capex")
reaction = None if pending else overrides.get("good_news_stock_down")
observed = not pending and bool(overrides.get("fetched_at"))
capex_signal = _capex_signal(capex, names) if observed else "unknown"
reaction_signal = _reaction_signal(reaction) if observed else "unknown"
state = combine_fundamental_signals(capex_signal, reaction_signal)
return {
"state": state,
"evidence_quality": _evidence_quality(
capex, reaction, names,
observed=observed, stale=stale, source=overrides.get("source"),
),
"capex_signal": capex_signal,
"reaction_signal": reaction_signal,
# Two different questions, and conflating them is a trap:
#
# `available` is about *timing* -- there is an effective, non-stale record
# to display. `usable` is about *content* -- it also actually says
# something. A collected observation whose extraction failed on every
# hyperscaler is available (show it, with its date) but not usable: it
# knows nothing, so it must never count as evidence.
#
# The distinction is load-bearing for the event study. Coverage is
# measured in sessions with usable context, and if repeated extraction
# failures counted, they would slowly accumulate "exposure" until the
# fundamental rows flipped to measurable 0/8 -- a failed result reported
# for a channel that never knew anything, which is the exact confusion
# coverage-matching exists to prevent.
"available": not pending and not stale,
"usable": _usable_context(observed, pending, stale, state),
"pending": pending,
"stale": stale,
"effective_date": effective.isoformat() if effective else None,
"age_days": age,
"capex": None if pending else overrides.get("capex"),
"good_news_stock_down": None if pending else overrides.get("good_news_stock_down"),
"capex": capex,
"good_news_stock_down": reaction,
"capex_stress": None if pending else overrides.get("f1_score"),
"earnings_stress": None if pending else overrides.get("f3_score"),
"reasoning": None if pending else overrides.get("reasoning"),
@@ -543,7 +785,7 @@ def fundamental_overlay(overrides: dict, config: dict, as_of: date) -> dict:
def current_observation(overrides: dict, config: dict, as_of: date) -> dict:
"""The observation as it stands now, for the live reading only.
Same shape as ``fundamental_overlay``, but the effective date is *reported*
Same shape as ``fundamental_context``, but the effective date is *reported*
rather than used to blank the content. A refresh stamps
``_next_weekday(today)``, so gating the live card hid a just-collected read
for one day -- three over a weekend -- and refreshing appeared to do
@@ -551,14 +793,36 @@ def current_observation(overrides: dict, config: dict, as_of: date) -> dict:
published number; the stored snapshot keeps the gate.
"""
effective, pending, age, stale = _overlay_timing(overrides, config, as_of)
# The default override carries "unknown"/"mixed" placeholders for every
# The default override carries "unknown" placeholders for every
# hyperscaler. Those are the absence of an observation, not an observation
# of absence, and must never be presented as collected. ``fetched_at`` is
# the collection timestamp and is the only field written on every path that
# produces real content (LLM refresh and manual save both stamp it).
observed = bool(overrides.get("fetched_at"))
names = list(config["tickers"]["hyperscalers"])
capex_signal = _capex_signal(overrides.get("capex"), names) if observed else "unknown"
reaction_signal = (
_reaction_signal(overrides.get("good_news_stock_down")) if observed else "unknown"
)
state = combine_fundamental_signals(capex_signal, reaction_signal)
return {
"observed": observed,
"state": state,
"evidence_quality": _evidence_quality(
overrides.get("capex"), overrides.get("good_news_stock_down"), names,
observed=observed, stale=stale, source=overrides.get("source"),
),
"capex_signal": capex_signal,
"reaction_signal": reaction_signal,
# Same shape as the record means the same *fields*, not just the same
# ones this function happens to need: the frontend types both payloads
# identically, so an omission here is an undefined at runtime that
# TypeScript cannot catch across a trusted server boundary.
#
# Note this is stricter than the `available` directly below: a pending
# observation is the freshest thing we have and worth showing, but it is
# not yet in force, so it is not yet evidence.
"usable": _usable_context(observed, pending, stale, state),
# Live availability is about usefulness, not effectiveness: a pending
# observation is the freshest thing we have -- but nothing collected is
# never available.
@@ -596,8 +860,16 @@ def _compute_index(
breadth_series: Series | None = None,
divergence_series: Series | None = None,
breadth_counts: dict[date, int] | None = None,
observations: list[dict] | None = None,
) -> dict:
"""Compute the complete v2 State/Warning snapshot as of one trading date."""
"""Compute the complete State/Warning snapshot as of one trading date.
``observations`` is the point-in-time fundamental series and is authoritative
when supplied; ``overrides`` is the single-slot fallback for callers that
predate the table (the calibration harness). Either way the reading is scored
into the same ``fundamental_context`` -- only where it is read from differs,
so the live monitor and the event study cannot report different states.
"""
tickers = config["tickers"]
smh = _closes_asof(prices.get(tickers["leaders"][0], []), as_of)
qqq = _closes_asof(prices.get(tickers["confirm"][0], []), as_of)
@@ -622,7 +894,10 @@ def _compute_index(
sensors = warning_sensor_scores(divergence, smh, spy, oas_window)
relative_strength = sensors["relative_strength"]
credit_impulse = sensors["credit_impulse"]
overlay = fundamental_overlay(overrides, config, as_of)
observation = (
observation_asof(observations, as_of) if observations is not None else overrides
) or {}
context = fundamental_context(observation, config, as_of)
state_pillars = [
{
@@ -708,7 +983,7 @@ def _compute_index(
"date": as_of.isoformat(),
"state": state,
"warning": warning,
"fundamental_overlay": overlay,
"fundamental_context": context,
"quadrant_config": {
"state_divider": QUADRANT_STATE_DIVIDER,
"warning_divider": QUADRANT_WARNING_DIVIDER,
@@ -730,8 +1005,8 @@ def _compute_index(
"breadth_pct_above_200": round(breadth_pct, 1) if breadth_pct is not None else None,
"breadth_date": breadth_item[0].isoformat() if breadth_item else None,
"fundamentals_fetched_at": overrides.get("fetched_at"),
"fundamentals_effective_date": overlay.get("effective_date"),
"fundamentals_age_days": overlay.get("age_days"),
"fundamentals_effective_date": context.get("effective_date"),
"fundamentals_age_days": context.get("age_days"),
},
"data_quality": {
"minimum_coverage": MIN_COVERAGE,
@@ -772,7 +1047,7 @@ async def get_regime_config(db: AsyncSession) -> dict:
if stored.get("fundamental_staleness_days") is not None:
cfg["fundamental_staleness_days"] = int(stored["fundamental_staleness_days"])
except (TypeError, ValueError, ValidationError):
logger.warning("Corrupt %s; using v2 defaults", KEY_CONFIG)
logger.warning("Corrupt %s; using defaults", KEY_CONFIG)
return cfg
@@ -799,7 +1074,7 @@ async def get_fundamental_overrides(db: AsyncSession) -> dict:
"f1_score": None,
"f3_score": None,
"capex": {name: "unknown" for name in names},
"good_news_stock_down": "mixed",
"good_news_stock_down": "unknown",
"locked": False,
"reasoning": None,
"fetched_at": None,
@@ -820,9 +1095,9 @@ async def get_fundamental_overrides(db: AsyncSession) -> dict:
if stored.get("methodology") not in CATEGORICAL_FUNDAMENTAL_METHODOLOGIES:
return default
capex = _normalise_capex_states(stored.get("capex"), names)
reaction = str(stored.get("good_news_stock_down", "mixed")).strip().lower()
reaction = str(stored.get("good_news_stock_down", "unknown")).strip().lower()
if reaction not in GNSD_STATES:
reaction = "mixed"
reaction = "unknown"
return {
**default,
**stored,
@@ -862,6 +1137,100 @@ def _score_capex_states(capex: dict[str, str], names: list[str]) -> float | None
return round(score, 1) if score is not None else None
async def record_fundamental_observation(db: AsyncSession, observation: dict) -> None:
"""Append the observation to the point-in-time series, keyed on effective date.
Upsert rather than insert: re-saving on the same effective date is a
correction to that day's reading, not a second observation of it.
Silently does nothing without an effective date or a ``fetched_at``. Those
are the default placeholder blob -- the absence of an observation, which must
never enter the series as though someone had looked.
Deliberately does **not** commit. ``update_regime_monitor`` calls this inside
a run that owns its transaction and commits once after the snapshot loop;
committing here would take that boundary away from it. The two override
writers commit for themselves.
"""
effective = _parse_date(observation.get("effective_date"))
fetched_raw = observation.get("fetched_at")
if effective is None or not fetched_raw:
return
try:
fetched = datetime.fromisoformat(str(fetched_raw))
except ValueError:
fetched = datetime.now(timezone.utc)
if fetched.tzinfo is None:
fetched = fetched.replace(tzinfo=timezone.utc)
existing = await db.execute(
select(RegimeFundamentalObservation).where(
RegimeFundamentalObservation.effective_date == effective
)
)
row = existing.scalar_one_or_none()
payload = {
"f1_score": observation.get("f1_score"),
"f3_score": observation.get("f3_score"),
"capex_json": json.dumps(observation.get("capex") or {}),
"good_news_stock_down": str(observation.get("good_news_stock_down") or "unknown")[:10],
"reasoning": observation.get("reasoning"),
"source": str(observation.get("source") or "unknown")[:30],
"fetched_at": fetched,
}
if row is None:
db.add(RegimeFundamentalObservation(
effective_date=effective,
created_at=datetime.now(timezone.utc),
**payload,
))
else:
for key, value in payload.items():
setattr(row, key, value)
async def get_fundamental_observations(db: AsyncSession) -> list[dict]:
"""The whole observation series, oldest first, for point-in-time scoring."""
result = await db.execute(
select(RegimeFundamentalObservation).order_by(
RegimeFundamentalObservation.effective_date.asc()
)
)
out: list[dict] = []
for row in result.scalars().all():
try:
capex = json.loads(row.capex_json)
except (TypeError, ValueError):
capex = {}
out.append({
"effective_date": row.effective_date,
"f1_score": row.f1_score,
"f3_score": row.f3_score,
"capex": capex,
"good_news_stock_down": row.good_news_stock_down,
"reasoning": row.reasoning,
"source": row.source,
"fetched_at": row.fetched_at.isoformat() if row.fetched_at else None,
})
return out
def observation_asof(observations: list[dict] | None, as_of: date) -> dict | None:
"""Latest observation effective on or before ``as_of``.
This *is* the effective-date gate now. The settings-blob version had to
recompute it per call because there was only ever one observation to gate;
with a series, "which reading was live that day" is just a lookup.
"""
chosen: dict | None = None
for observation in observations or []:
if observation["effective_date"] <= as_of:
chosen = observation
else:
break
return chosen
async def set_fundamental_overrides(
db: AsyncSession,
capex: dict[str, str] | None = None,
@@ -894,7 +1263,15 @@ async def set_fundamental_overrides(
"fetched_at": now.isoformat(),
"effective_date": _next_weekday(now.date()).isoformat(),
})
await update_setting(db, KEY_FUNDAMENTALS, json.dumps(current))
# The blob (what the live card reads) and the series row (what the
# point-in-time replay reads) are the same observation. Committed together:
# `update_setting` commits internally, so using it here would leave a window
# where a failure publishes the reading to the card but not to the record,
# and the two would disagree permanently with nothing to detect it.
await settings_store.upsert_setting(db, KEY_FUNDAMENTALS, json.dumps(current))
if observation_changed:
await record_fundamental_observation(db, current)
await db.commit()
return current
@@ -998,12 +1375,59 @@ def _snapshot_revision(snapshot: dict) -> int:
return 1
def _context_from_legacy_overlay(overlay: dict) -> dict:
"""Rebuild the categorical channel from a pre-rename snapshot's overlay.
The channel was called ``fundamental_overlay`` until 2026-08-12 and stored
the same underlying facts -- the capex map, the earnings reaction, the
effective date. The rename shipped without a methodology bump (no score
changed), so those rows are still served and were never reseeded: reading
only the new key would turn every one of them into ``unknown`` and silently
discard real recorded evidence -- historical Path colours, and any exposure
the event study could legitimately count.
Derived, not guessed. The hyperscaler list comes from the overlay's own
capex keys, which is exactly the basket that was observed at the time rather
than today's configured one.
"""
capex = overlay.get("capex") or {}
reaction = overlay.get("good_news_stock_down")
names = list(capex)
pending = bool(overlay.get("pending"))
stale = bool(overlay.get("stale"))
observed = not pending and bool(overlay.get("fetched_at"))
capex_signal = _capex_signal(capex, names) if observed else "unknown"
reaction_signal = _reaction_signal(reaction) if observed else "unknown"
state = combine_fundamental_signals(capex_signal, reaction_signal)
return {
**overlay,
"state": state,
"evidence_quality": _evidence_quality(
capex, reaction, names,
observed=observed, stale=stale, source=overlay.get("source"),
),
"capex_signal": capex_signal,
"reaction_signal": reaction_signal,
"usable": _usable_context(observed, pending, stale, state),
}
def _parse_snapshot(raw: str) -> dict | None:
try:
parsed = json.loads(raw)
except (TypeError, ValueError):
return None
return parsed if parsed.get("methodology") == METHODOLOGY else None
if parsed.get("methodology") != METHODOLOGY:
return None
# Normalise here rather than at each call site: every reader of a stored
# snapshot goes through this function, so a legacy row cannot reach one of
# them un-adapted.
if "fundamental_context" not in parsed and "fundamental_overlay" in parsed:
parsed["fundamental_context"] = _context_from_legacy_overlay(
parsed["fundamental_overlay"] or {}
)
return parsed
async def _latest_snapshot_row(db: AsyncSession) -> tuple[RegimeSnapshot, dict] | None:
@@ -1022,6 +1446,10 @@ async def update_regime_monitor(
) -> dict:
config = await get_regime_config(db)
overrides = await get_fundamental_overrides(db)
# Carries the pre-v5 single-slot observation into the series on first run, so
# a deployment does not lose the live reading. A no-op once recorded, and a
# no-op for the placeholder blob (no fetched_at).
await record_fundamental_observation(db, overrides)
if _fundamentals_stale(overrides, config) and not overrides.get("locked"):
try:
overrides = await refresh_fundamental_overrides(db, config=config)
@@ -1071,6 +1499,9 @@ async def update_regime_monitor(
breadth_series = _mapping_series(breadth)
divergence_series = _mapping_series(divergence)
# Loaded once, after any refresh, so a reseed scores each replayed date with
# the observation that was effective on it rather than with today's.
observations = await get_fundamental_observations(db)
latest_result: dict | None = None
snapshots_written = 0
for snapshot_date in dates:
@@ -1084,6 +1515,7 @@ async def update_regime_monitor(
breadth_series,
divergence_series,
breadth_counts,
observations=observations,
)
written, latest_result = await _upsert_snapshot(
db,
@@ -1161,16 +1593,22 @@ async def get_regime_monitor(db: AsyncSession) -> dict:
quality["is_fresh"] = bool(quality.get("inputs_fresh")) and snapshot_age <= 4
result["data_quality"] = quality
# The snapshot's overlay is the point-in-time record; the reader also wants
# the current observation even when it is not effective until the next
# session, because otherwise refreshing it looks like it did nothing.
# The snapshot's `fundamental_context` is the point-in-time record; the
# reader also wants the current observation even when it is not effective
# until the next session, or refreshing it looks like it did nothing.
config = await get_regime_config(db)
overrides = await get_fundamental_overrides(db)
live = current_observation(overrides, config, date.today())
# Deliberately reads the *snapshot's* overlay, not the live one: this is how
# Deliberately reads the *snapshot's* record, not the live one: this is how
# the reader tells "shown here" from "in the stored record".
live["observed_in_snapshot"] = bool((result.get("fundamental_overlay") or {}).get("available"))
result["fundamental_context"] = live
live["observed_in_snapshot"] = bool(
(result.get("fundamental_context") or {}).get("available")
)
# `fundamental_context` is the stored channel and stays the snapshot's;
# `fundamental_live` is what we know right now. Collapsing the two under one
# key is what made a just-collected observation look like it had been
# backdated into history.
result["fundamental_live"] = live
result["available"] = True
return result
@@ -1188,10 +1626,17 @@ async def get_regime_history(db: AsyncSession, days: int = 800) -> list[dict]:
if data is None:
continue
state, warning = data.get("state") or {}, data.get("warning") or {}
context = data.get("fundamental_context") or {}
out.append({
"date": row.date.isoformat(),
"state": state.get("score") if state.get("band") is not None else None,
"warning": warning.get("score") if warning.get("band") is not None else None,
# The third channel, carried per point so the Path view can colour a
# dot by the fundamental context that was on the record that day.
# Rows written before the channel existed carry nothing, which reads
# as "unknown" -- correct, since nothing was observed then either.
"fundamental_state": context.get("state") or "unknown",
"evidence_quality": context.get("evidence_quality") or "unavailable",
"state_coverage": state.get("coverage"),
"warning_coverage": warning.get("coverage"),
"basket_hash": (data.get("basket") or {}).get("hash"),
@@ -1323,7 +1768,7 @@ async def refresh_fundamental_overrides(
f1 = _score_capex_states(capex, names)
reaction = str(parsed.get("good_news_stock_down", "")).strip().lower()
if reaction not in GNSD_STATES:
reaction = "mixed"
reaction = "unknown"
f3 = _GNSD_SCORES.get(reaction)
now = datetime.now(timezone.utc)
result = {
@@ -1338,7 +1783,11 @@ async def refresh_fundamental_overrides(
"locked": False,
"source": llm.get("provider"),
}
await update_setting(db, KEY_FUNDAMENTALS, json.dumps(result))
# One transaction: see set_fundamental_overrides on why these two writes must
# not be able to land separately.
await settings_store.upsert_setting(db, KEY_FUNDAMENTALS, json.dumps(result))
await record_fundamental_observation(db, result)
await db.commit()
logger.info(json.dumps({
"event": "regime_fundamentals_refreshed",
"f1": result["f1_score"],
+6 -2
View File
@@ -31,7 +31,7 @@ from app.services import fundamentals_quality_service, system_event_service
from app.services.price_service import query_ohlcv
from app.services.qualification import setup_qualifies
from app.services.sr_service import detect_gate_target_ladder
from app.services import settings_store
from app.services import settings_store, ticker_service
from app.services.trade_policy import (
MANUAL_BOOK,
SHADOW_BOOK,
@@ -735,7 +735,11 @@ async def scan_all_tickers(
# Plain ids/strings, not Ticker instances: the rollbacks below expire any
# ORM objects held across them, and touching an expired attribute afterwards
# triggers sync lazy-loading, which raises on an AsyncSession.
result = await db.execute(select(Ticker.id, Ticker.symbol).order_by(Ticker.symbol))
result = await db.execute(
ticker_service.active_only(
select(Ticker.id, Ticker.symbol).order_by(Ticker.symbol)
)
)
ticker_rows = [(int(ticker_id), symbol) for ticker_id, symbol in result.all()]
total = len(ticker_rows)
+7 -3
View File
@@ -20,7 +20,7 @@ from app.database import insert_for_session
from app.exceptions import NotFoundError, ValidationError
from app.models.score import CompositeScore, DimensionScore
from app.models.ticker import Ticker
from app.services import settings_store
from app.services import settings_store, ticker_service
logger = logging.getLogger(__name__)
@@ -883,7 +883,11 @@ async def get_rankings(db: AsyncSession) -> dict:
Returns dict suitable for RankingResponse.
"""
weights = await _get_weights(db)
tickers = (await db.execute(select(Ticker).order_by(Ticker.symbol))).scalars().all()
tickers = (
await db.execute(
ticker_service.active_only(select(Ticker).order_by(Ticker.symbol))
)
).scalars().all()
async def _load_scores() -> tuple[dict[int, CompositeScore], dict[int, dict[str, DimensionScore]]]:
comps = {
@@ -947,7 +951,7 @@ async def update_weights(
await _save_weights(db, full_weights)
# Recompute all composite scores
result = await db.execute(select(Ticker))
result = await db.execute(ticker_service.active_only(select(Ticker)))
tickers = list(result.scalars().all())
for ticker in tickers:
+112
View File
@@ -45,6 +45,33 @@ _CA_VERIFY: str | bool = _CA if _CA and Path(_CA).exists() else True
_FORMS_10 = frozenset({"10-K", "10-Q", "10-K/A", "10-Q/A"})
# Notification of removal from listing. "25" is issuer-filed, "25-NSE" exchange-
# filed. The Form 15 family is deliberately absent: it ends a *reporting*
# obligation and does not mean the security stopped trading.
_DELISTING_FORMS = frozenset({"25", "25-NSE"})
# ``descriptionClassSecurity`` is free text ("Common Stock", "Class A Common
# Stock, $0.01 par value", "6.25% Notes due 2030", "Warrants", "Depositary
# Shares"). Only a common-equity class means the ticker itself stopped trading.
_NON_COMMON_CLASS = re.compile(
r"\b(note|bond|debenture|preferred|warrant|right|unit|depositary|"
r"subordinated|debt|trust)s?\b",
re.IGNORECASE,
)
def _is_common_stock(description: str) -> bool:
"""Does this Form 25 security class describe common equity?
Requires an explicit common-stock match AND no debt/preferred/warrant marker,
so "Depositary Shares each representing 1/1000th of Preferred" cannot pass on
the word "shares" alone. Unrecognised text is rejected a symbol is retired
on this answer, so ambiguity must not read as yes.
"""
if _NON_COMMON_CLASS.search(description):
return False
return re.search(r"\bcommon\s+(stock|share)", description, re.IGNORECASE) is not None
class SecError(ProviderError):
"""SEC request failed (403, exhausted 429/5xx, timeout, transport, parse)."""
@@ -253,6 +280,91 @@ class SecClient:
"filings": filings,
}
async def delisting_filing(
self, cik: int | str, *, not_before: date | None = None
) -> dict[str, Any] | None:
"""Newest Form 25 removing this issuer's COMMON stock from listing.
Deliberately narrow, because the caller retires a symbol on the answer:
- **Form 25 only.** The Form 15 family terminates a reporting obligation
(often just a class falling under the holder threshold) and is no
evidence that trading stopped.
- **Class-checked.** Form 25 is filed per security class an issuer
delisting its notes, preferred, warrants or an ADR class while the
common keeps trading files one too. The filing's own
``descriptionClassSecurity`` is what separates those, so the primary
document is fetched and read rather than trusting the form type.
- **``not_before``** rejects a historical filing for some long-gone
class. Without it a 2019 Form 25 would retire a symbol whose bars
stopped in 2026, and stamp 2019 as the date.
Anything unreadable no primary document (pre-2009 filings have none),
malformed XML, unrecognised class returns ``None``. Fail closed: the
caller keeps warning instead of retiring on a guess.
Reads ``filings.recent`` directly; ``submissions()`` keeps only the
10-K/10-Q family, so Form 25 never survives its parser.
"""
base = await self.get_json(f"{_DATA}/submissions/CIK{cik10(cik)}.json")
arrays = (base.get("filings") or {}).get("recent") or {}
forms = arrays.get("form") or []
dates = arrays.get("filingDate") or []
accessions = arrays.get("accessionNumber") or []
docs = arrays.get("primaryDocument") or []
candidates: list[tuple[date, str, str, str]] = []
for i, form in enumerate(forms):
if form not in _DELISTING_FORMS or i >= len(dates) or not dates[i]:
continue
try:
filed = date.fromisoformat(dates[i])
except ValueError:
continue
if not_before is not None and filed < not_before:
continue
if i >= len(accessions) or not accessions[i]:
continue
candidates.append((filed, form, accessions[i], docs[i] if i < len(docs) else ""))
for filed, form, accession, _doc in sorted(candidates, reverse=True):
security = await self._form25_security_class(cik, accession)
if security is None:
continue
if not _is_common_stock(security):
continue
return {
"form": form,
"filing_date": filed,
"security_class": security,
}
return None
async def _form25_security_class(
self, cik: int | str, accession: str
) -> str | None:
"""``descriptionClassSecurity`` from a Form 25's primary XML, or None.
The rendered ``primaryDocument`` is an XSL view of this file; the raw
``primary_doc.xml`` beside it is the structured original.
"""
folder = accession.replace("-", "")
url = (
f"{_WWW}/Archives/edgar/data/{int(cik)}/{folder}/primary_doc.xml"
)
try:
body = await self.get_text(url)
except SecNotFoundError:
return None
match = re.search(
r"<descriptionClassSecurity>(.*?)</descriptionClassSecurity>",
body,
re.IGNORECASE | re.DOTALL,
)
if match is None:
return None
return " ".join(match.group(1).split()) or None
async def companyfacts(self, cik: int | str) -> dict[str, Any]:
"""Raw companyfacts JSON ({cik, entityName, facts})."""
return await self.get_json(f"{_DATA}/api/xbrl/companyfacts/CIK{cik10(cik)}.json")
+63 -10
View File
@@ -113,8 +113,41 @@ _WEIGHTED_AVG_SHARE_CONCEPTS = [
# us-gaap instant (balance-sheet) concepts, at end == reportDate.
_CASH = ["CashAndCashEquivalentsAtCarryingValue"]
_ST_INVESTMENTS = ["ShortTermInvestments", "MarketableSecuritiesCurrent"] # pick one
# Debt is tagged in four mutually exclusive styles across large filers, and
# composing a total means knowing which span each concept covers (measured
# 2026-08 over a 20-issuer sample; the counts below are from it).
#
# ``LongTermDebt`` already spans current + noncurrent maturities — Apple tags all
# three and 71.34bn + 11.01bn = 82.30bn confirms it — so its complement is only
# genuinely short-term borrowing.
_LONG_TERM_DEBT_AGG = ["LongTermDebt"]
_LONG_TERM_DEBT_PARTS = ["LongTermDebtNoncurrent", "LongTermDebtCurrent"]
# Noncurrent-only balance-sheet lines, needing a current complement added.
# ``LongTermDebtAndCapitalLeaseObligations`` is what KO, HD, T, XOM and CVX tag
# and nothing read it before: AT&T reported no total_debt at all against 134bn
# tagged, and Coca-Cola reported 0.25bn of commercial paper against 39bn.
_LONG_TERM_DEBT_NONCURRENT = [
"LongTermDebtNoncurrent",
"LongTermDebtAndCapitalLeaseObligations",
]
_LONG_TERM_DEBT_CURRENT = ["LongTermDebtCurrent"]
# REITs that tag no aggregate at all, carrying a secured and an unsecured side
# instead. Both sides are required, because ``NotesPayable`` does not mean the
# same thing across issuers (measured 2026-08 over 14 REITs):
# - MAA tags NotesPayable 5.66bn = UnsecuredDebt 5.30bn + SecuredDebt 0.36bn
# exactly, so there it IS the total and adding SecuredDebt double-counts.
# - EQR/VMRK tags NotesPayable alongside a *larger* SecuredDebt (5.38bn vs
# 6.38bn in 2013), so there it is only the unsecured component.
# ``UnsecuredDebt`` is what separates them: where it is tagged it is the
# unambiguous unsecured side and NotesPayable is ignored; where it is absent,
# NotesPayable is that side. Requiring both sides is also what keeps this branch
# from inventing a total out of a fragment — Boston Properties tags SecuredDebt
# 4.28bn and nothing else against ~15bn of real debt, and Regency tags an
# UnsecuredDebt of 0.03bn that is a credit-line draw, not its 5bn of notes.
_SECURED_DEBT = ["SecuredDebt"]
_UNSECURED_DEBT = ["UnsecuredDebt", "NotesPayable"] # first present wins
# ``DebtCurrent`` spans short-term borrowing AND current maturities, so it is the
# whole current complement where present and must never be added alongside them.
_ALL_CURRENT_DEBT = ["DebtCurrent"]
_SHORT_TERM_DEBT = ["ShortTermBorrowings", "CommercialPaper"] # pick one
@@ -459,15 +492,35 @@ def _compose_cash(facts: list[Fact], report_date: date) -> float | None:
def _compose_debt(facts: list[Fact], report_date: date) -> float | None:
long_term = _select_instant(facts, _LONG_TERM_DEBT_AGG, report_date)
if long_term is None:
nc = _select_instant(facts, ["LongTermDebtNoncurrent"], report_date)
cur = _select_instant(facts, ["LongTermDebtCurrent"], report_date)
long_term = None if nc is None and cur is None else (nc or 0.0) + (cur or 0.0)
short_term = _select_instant(facts, _SHORT_TERM_DEBT, report_date)
if long_term is None and short_term is None:
return None
return (long_term or 0.0) + (short_term or 0.0)
"""Total debt at ``report_date``, or None when no long-term component is found.
**A short-term component alone is never a total.** Chevron tags its full debt
only in the 10-K, so its 10-Q carries ``ShortTermBorrowings`` of 0.40bn and
nothing else; returning that as total debt reads as a near-unlevered issuer
carrying 50bn. Since ``_net_debt`` needs both sides and yields nothing when
either is missing, None costs a leverage read while the partial value
produces a confidently wrong one.
"""
# An aggregate spanning current + noncurrent: only true short-term is missing.
total = _select_instant(facts, _LONG_TERM_DEBT_AGG, report_date)
if total is not None:
return total + (_select_instant(facts, _SHORT_TERM_DEBT, report_date) or 0.0)
noncurrent = _select_instant(facts, _LONG_TERM_DEBT_NONCURRENT, report_date)
if noncurrent is None:
secured = _select_instant(facts, _SECURED_DEBT, report_date)
unsecured = _select_instant(facts, _UNSECURED_DEBT, report_date)
if secured is None or unsecured is None:
return None # one side of a REIT's debt is not its total
noncurrent = secured + unsecured
current = _select_instant(facts, _ALL_CURRENT_DEBT, report_date)
if current is None:
current = (
(_select_instant(facts, _LONG_TERM_DEBT_CURRENT, report_date) or 0.0)
+ (_select_instant(facts, _SHORT_TERM_DEBT, report_date) or 0.0)
)
return noncurrent + current
def _select_shares(
+222 -10
View File
@@ -36,7 +36,11 @@ Guardrails (design + reviews):
excluded from actionable setups until its filing is recovered.
- ``promote`` inserts snapshots ``ON CONFLICT (accession) DO NOTHING`` (immutable),
reports differing existing accessions, and applies ticker updates in the same
transaction.
transaction. A difference in ``cik`` **alone** is reported separately as an
``accession_cik_collision``: every fact matched, so two tracked CIKs are
claiming one filing and the fix is the universe, not the parser. It never
self-heals on its own the losing CIK stores no row, so it is backfilled and
re-reported every run until its ticker is re-pointed or retired.
- ``reparse=True`` is the one exception to immutability, and it is deliberate:
it restages every accession with the current parser and **rewrites** the rows
that now reconstruct differently. Immutability protects SEC's record (one row
@@ -51,10 +55,10 @@ import json
import logging
from collections import Counter, defaultdict
from dataclasses import dataclass, field, replace
from datetime import date, datetime, timedelta, timezone
from datetime import date, datetime, time, timedelta, timezone
from typing import Any, Callable
from sqlalchemy import delete, select, update
from sqlalchemy import delete, exists, select, update
from app.database import insert_for_session
from app.models.data_import_run import DataImportRun
@@ -67,7 +71,7 @@ from app.services import sec_universe
from app.services.data_import import STATUS_PROMOTED, ValidationResult
from app.services.sec_client import SecClient, SecError, cik10
from app.services.sec_facts_parser import FilingMeta, SnapshotRow
from app.services.sec_universe import ResolvedUniverse
from app.services.sec_universe import CIK_OVERRIDES_KEY, ResolvedUniverse
logger = logging.getLogger(__name__)
@@ -81,6 +85,26 @@ MIN_BACKFILL_COVERAGE = 0.5
# three); past that it is misfiled, not late, and blocking forever costs more
# than the missing filing does — see the unresolved-filing guardrail below.
MISSING_XBRL_RETRY_DAYS = 3
# Aggregate ceiling on deferral. MISSING_XBRL_RETRY_DAYS bounds how long ONE
# filing blocks; it does not bound how long the import as a whole can stay
# deferred. Those differ because a blocking filing is only queued by promote(),
# which a deferred run never reaches — so during a rolling supply of
# unresolvable filings (earnings season, when SEC's Company-Facts aggregation is
# furthest behind) each new arrival restarts the clock before the previous one
# clears, and nothing is written at all: not the good rows, not the gap rows
# that would stop those filings blocking again.
#
# Once promotions have been stale this long, every unresolved filing is treated
# as past the window. promote() then queues them all (see the _past_retry_window
# call there), source_max_date advances, and _missing() forces queued rows
# aged-out on later runs so they never block again — the import self-heals
# through the paths that already exist.
#
# Well above MISSING_XBRL_RETRY_DAYS so ordinary overlapping blocks never trip
# it. Affected symbols stay barred from setups either way: setup_blocked_ciks is
# built from every missing filing regardless of window.
PROMOTION_CEILING_DAYS = 7
FILING_GAP_ESCALATE_DAYS = 14
# Share-count band a co-registrant-recovered row must land in, relative to the
# issuer's own last snapshot. Wide enough for buybacks/issuance, nowhere near
@@ -157,6 +181,9 @@ class SecFundamentalsImporter:
self._retry_rows: list[dict[str, Any]] = []
self._latest_index_date: date | None = None
self._backfill = False
# Set by validate() when the aggregate ceiling forced the block open;
# read by promote() to alert that it did.
self._ceiling_tripped: dict[str, Any] | None = None
# -- SourceImporter protocol -------------------------------------------
@@ -258,7 +285,18 @@ class SecFundamentalsImporter:
if old is not None:
fields = _diff_fields(row, old)
if fields:
staged.discrepancies.append({"accession": row.accession, "fields": fields})
# Carry both CIKs. promote() reads a bare ["cik"] as an
# attribution collision rather than a changed
# reconstruction, which holds only because _COMPARE_COLS
# spans every stored fact: a fact column added to the
# model but not to _SNAPSHOT_COLS would go uncompared and
# let a real difference through as a collision.
staged.discrepancies.append({
"accession": row.accession,
"fields": fields,
"cik": row.cik,
"stored_cik": old.cik,
})
return staged
async def _stage_issuer(
@@ -421,6 +459,24 @@ class SecFundamentalsImporter:
# reconstructible by re-walking the index.
blocking = _within_retry_window(staged.missing_xbrl)
aged_out = _past_retry_window(staged.missing_xbrl)
# ...unless promotions have been stale past the aggregate ceiling, in
# which case the deferral has cost more than the filings it withholds.
# Ageing them here (not just locally) is deliberate: promote() re-derives
# the queue from the same list, so this is what gets them queued.
self._ceiling_tripped = None
if blocking and db is not None and await self._promotions_stale(db):
for item in staged.missing_xbrl:
item["age_days"] = max(
item.get("age_days", 0), MISSING_XBRL_RETRY_DAYS + 1
)
self._ceiling_tripped = {
"forced": len(blocking),
"unresolved": len(staged.missing_xbrl),
}
blocking = _within_retry_window(staged.missing_xbrl)
aged_out = _past_retry_window(staged.missing_xbrl)
if blocking:
messages.append(
f"{len(blocking)} tracked XBRL filing(s) unresolved within the "
@@ -464,6 +520,9 @@ class SecFundamentalsImporter:
"missing_xbrl": staged.missing_xbrl[:50],
"missing_xbrl_count": len(staged.missing_xbrl),
"missing_xbrl_blocking": len(blocking),
# Present only when the aggregate ceiling forced this run through, so
# a promoted run that carries known-unresolved filings says so.
"promotion_ceiling_tripped": self._ceiling_tripped,
"recovered_from_coregistrant": staged.recovered[:50],
"recovered_count": len(staged.recovered),
# Complete compact gate input; detailed audit lists above stay capped.
@@ -511,8 +570,15 @@ class SecFundamentalsImporter:
inserted = 0
updated = 0
# Only accessions whose reconstruction actually changed are rewritten;
# an unchanged stored row is left completely alone.
changed = {d["accession"] for d in staged.discrepancies} if self.reparse else set()
# an unchanged stored row is left completely alone. A cik-only difference
# is excluded on purpose: the facts are identical there, so rewriting
# would re-stamp the filing onto the colliding co-registrant — taking it
# from the issuer that actually filed it, which no parser fix asks for.
changed = (
{d["accession"] for d in staged.discrepancies if d["fields"] != ["cik"]}
if self.reparse
else set()
)
for row in staged.rows:
if row.accession in staged.existing_accessions:
if row.accession in changed:
@@ -597,10 +663,51 @@ class SecFundamentalsImporter:
if gap["accession"] not in existing_gap_accessions
]
# Two tracked issuers claiming one filing is not a reconstruction change:
# every fact matched and only the CIK stamp differs, so re-parsing or
# reparsing fixes nothing — the universe resolution does. It is reported
# separately because it also does not self-heal: the loser of the
# collision never stores a row, so `_ciks_with_snapshots` never sees it,
# and it is full-history backfilled (and re-reported) on every run until
# a human re-points or retires the ticker. Observed 2026-08 for EQR,
# which SEC's own company_tickers.json maps to ERP Operating LP, the
# non-traded co-registrant of the issuer now trading as VMRK.
collisions = [d for d in staged.discrepancies if d["fields"] == ["cik"]]
if collisions:
named = ", ".join(
f"{d['accession']} (stored {d['stored_cik']}, parsed {d['cik']})"
for d in collisions[:10]
)
db.add(SystemEvent(
severity="warning",
source="sec_facts",
code="accession_cik_collision",
message=(
f"{len(collisions)} filing(s) are claimed by two tracked CIKs — "
"the reconstruction is identical, only the attribution differs, "
"so one of the two is a co-registrant the universe should not "
f"track. Re-point or retire the ticker (see {CIK_OVERRIDES_KEY}); "
f"this repeats every run until then: {named}"
)[:4000],
dedup_key=f"sec_facts:accession_cik_collision:{run_id}",
created_at=_now(),
))
# Warn (in-transaction, so it commits atomically with the promotion) when
# any existing accession reconstructed differently — kept immutable.
if staged.discrepancies:
accns = ", ".join(d["accession"] for d in staged.discrepancies[:10])
reconstruction_diffs = [
d for d in staged.discrepancies if d["fields"] != ["cik"]
]
if reconstruction_diffs:
# Name the columns, not just the accession: "differs in revenue"
# (our numbers moved) and "differs in period_start" (the filing was
# re-placed in the calendar) need different responses, and the alert
# is where that call gets made. The fields are already computed for
# validation_json — they were simply dropped from the message.
accns = ", ".join(
f"{d['accession']} ({', '.join(d['fields'])})"
for d in reconstruction_diffs[:10]
)
disposition = (
f"REWRITTEN by reparse run {run_id}" if self.reparse else "kept immutable"
)
@@ -609,13 +716,32 @@ class SecFundamentalsImporter:
source="sec_facts",
code="snapshot_reparse" if self.reparse else "snapshot_discrepancy",
message=(
f"{len(staged.discrepancies)} stored accession(s) reconstructed "
f"{len(reconstruction_diffs)} stored accession(s) reconstructed "
f"differently; {disposition}: {accns}"
)[:4000],
dedup_key=f"sec_facts:discrepancy:{run_id}",
created_at=_now(),
))
# A ceiling-forced promotion is the safety valve firing — it must be
# visible, or the import silently starts carrying known-unresolved
# filings. The affected symbols stay barred from setups regardless.
if self._ceiling_tripped:
db.add(SystemEvent(
severity="warning",
source="sec_facts",
code="promotion_ceiling_forced",
message=(
f"Promoted with {self._ceiling_tripped['unresolved']} unresolved "
f"filing(s) — {self._ceiling_tripped['forced']} still inside the "
f"{MISSING_XBRL_RETRY_DAYS}-day retry window — because nothing had "
f"promoted in {PROMOTION_CEILING_DAYS} days. They are queued for "
"retry and their symbols remain blocked from setups."
)[:4000],
dedup_key=f"sec_facts:promotion_ceiling_forced:{run_id}",
created_at=now,
))
# Persistent current gaps get one actionable escalation rather than a
# daily warning. The nullable marker makes this durable and noise-free.
escalation_cutoff = now - timedelta(days=FILING_GAP_ESCALATE_DAYS)
@@ -650,6 +776,60 @@ class SecFundamentalsImporter:
.values(escalated_at=now)
)
# The escalation above fires once per gap, so nothing would report the
# *end* of the reprieve it grants. An escalated gap stops pausing setups
# while the issuer's own fundamentals are still recent, and that lapses
# on its own — the stored filings age past the window, or a newer gap
# appears — putting the pause back on with no alert anywhere. Track the
# exemption as state and alert on the transition, once per lapse.
current_gaps = await fundamentals_quality_service.active_gaps(db)
escalated_gaps = [g for g in current_gaps if g.escalated_at is not None]
if escalated_gaps:
exempt_ciks = await fundamentals_quality_service.gap_exempt_ciks(
db, escalated_gaps
)
newly_exempt = [
g for g in escalated_gaps
if g.cik in exempt_ciks and g.exempted_at is None
]
lapsed = [
g for g in escalated_gaps
if g.cik not in exempt_ciks and g.exempted_at is not None
]
if newly_exempt:
# Silent on purpose: filing_gap_aged already announced this gap,
# and setups resuming is the behaviour that alert describes.
await db.execute(
update(SecFilingGap)
.where(SecFilingGap.id.in_([g.id for g in newly_exempt]))
.values(exempted_at=now)
)
if lapsed:
named = ", ".join(
f"{gap.cik}/{gap.accession}" for gap in lapsed[:10]
)
db.add(SystemEvent(
severity="warning",
source="sec_facts",
code="filing_gap_repaused",
message=(
f"{len(lapsed)} SEC filing gap(s) pause setups again: the "
"issuer's own fundamentals have aged out of the "
f"{fundamentals_quality_service.GAP_GATE_RECENT_FILING_DAYS}"
"-day window, or a newer gap arrived, so there is nothing "
f"recent left to score on: {named}"
)[:4000],
dedup_key=f"sec_facts:filing_gap_repaused:{run_id}",
created_at=now,
))
# Cleared, not stamped: the issuer can recover and age out again,
# and each lapse is worth its own alert.
await db.execute(
update(SecFilingGap)
.where(SecFilingGap.id.in_([g.id for g in lapsed]))
.values(exempted_at=None)
)
# Recovered rows are real data from an unexpected place — record where they
# came from, so a wrong recovery is auditable rather than invisible.
if staged.recovered:
@@ -757,6 +937,38 @@ class SecFundamentalsImporter:
if accession not in resolved
]
async def _promotions_stale(self, db) -> bool:
"""Has nothing promoted within ``PROMOTION_CEILING_DAYS``?
Only true for a source that HAS promoted before. A never-promoted import
is initial setup, not a wedge: forcing its first promotion through would
mask a misconfiguration rather than recover from a transient SEC gap.
Measured from ``self.today`` rather than the wall clock, so the ceiling
honors the same injected date that ages the filings it releases.
"""
cutoff = datetime.combine(
self.today - timedelta(days=PROMOTION_CEILING_DAYS),
time.min,
tzinfo=timezone.utc,
)
ever, recent = (
await db.execute(
select(
exists().where(
DataImportRun.source == SOURCE,
DataImportRun.status == STATUS_PROMOTED,
),
exists().where(
DataImportRun.source == SOURCE,
DataImportRun.status == STATUS_PROMOTED,
DataImportRun.started_at >= cutoff,
),
)
)
).one()
return bool(ever) and not bool(recent)
async def _last_processed_index_date(self, db) -> date | None:
return (
await db.execute(
+6 -2
View File
@@ -24,7 +24,7 @@ from typing import Iterable
from sqlalchemy import select, update
from app.models.ticker import Ticker
from app.services import settings_store
from app.services import settings_store, ticker_service
from app.services.earnings_alignment import normalise_symbol
from app.services.sec_client import SecClient
@@ -55,7 +55,11 @@ async def resolve_ciks(db, client: SecClient) -> ResolvedUniverse:
returns the mapping + proposed `tickers.cik` writes; mutates nothing."""
ticker_to_cik = await client.company_tickers()
overrides = await cik_overrides(db)
rows = (await db.execute(select(Ticker.id, Ticker.symbol, Ticker.cik))).all()
rows = (
await db.execute(
ticker_service.active_only(select(Ticker.id, Ticker.symbol, Ticker.cik))
)
).all()
result = ResolvedUniverse()
for tid, symbol, current_cik in rows:
+199 -3
View File
@@ -1,13 +1,65 @@
"""Ticker Registry service: add, delete, and list tracked tickers."""
"""Ticker Registry service: add, delete, list, and retire tracked tickers."""
import logging
import re
from datetime import date, timedelta
from sqlalchemy import select
from sqlalchemy import func, or_, select, update
from sqlalchemy.ext.asyncio import AsyncSession
from app.exceptions import DuplicateError, NotFoundError, ValidationError
from app.models.ticker import Ticker
logger = logging.getLogger(__name__)
# Reasons a symbol may be marked delisted, narrowest first.
REASON_FORM_25 = "form_25" # SEC Form 25/25-NSE/15 confirmed the exchange exit
REASON_MANUAL = "manual" # an operator decided
# How long a symbol must be without bars before we spend an SEC request asking
# whether it delisted. Guards against a market-data outage probing the whole
# universe at once; a real delisting is still stale days later.
MIN_STALE_DAYS_BEFORE_PROBE = 3
# Rule 12d2-2: a Form 25 removal takes effect ten days after filing, so the
# filing date is not the date the security stopped trading.
FORM_25_EFFECTIVE_DAYS = 10
# How far before the last bar a Form 25 may be filed and still explain this gap.
# An exchange can file shortly before trading actually stops; anything older
# concerns a class that was already gone while the symbol kept printing bars.
FILING_LOOKBACK_DAYS = 30
def _sec_client_factory():
"""Build the SEC client for a delisting probe (patched in tests).
Imported lazily so the SEC/httpx stack stays off the import path of every
module that only wants ``active_only``.
"""
from app.services.sec_client import SecClient
return SecClient()
def active_only(stmt, *, as_of: date | None = None):
"""Restrict a Ticker query to symbols that still trade.
Opt-in on purpose rather than folded into a shared getter: list and admin
views deliberately keep delisted rows so the delisting is *visible*, which a
silent default would undo. Apply this on the live signal path scanning,
ranking, scoring, breadth, ingestion and nowhere else.
``delisted_on`` is an *effective* date, and a Form 25 is known ten days
before it takes effect, so a future date must not drop the symbol yet it
is still trading and still worth scanning and ingesting. Compared in SQL
against the database's own date; ``as_of`` overrides it for tests.
"""
cutoff = func.current_date() if as_of is None else as_of
return stmt.where(
or_(Ticker.delisted_on.is_(None), Ticker.delisted_on > cutoff)
)
async def add_ticker(db: AsyncSession, symbol: str) -> Ticker:
"""Add a new ticker after validation.
@@ -52,6 +104,150 @@ async def delete_ticker(db: AsyncSession, symbol: str) -> None:
async def list_tickers(db: AsyncSession) -> list[Ticker]:
"""Return all tracked tickers sorted alphabetically by symbol."""
"""Return all tracked tickers sorted alphabetically by symbol.
Delisted symbols are included and carry ``delisted_on`` the registry is
where an operator needs to *see* that a symbol retired, not where it should
quietly disappear.
"""
result = await db.execute(select(Ticker).order_by(Ticker.symbol.asc()))
return list(result.scalars().all())
async def mark_delisted(
db: AsyncSession,
symbol: str,
*,
delisted_on: date,
reason: str = REASON_MANUAL,
) -> bool:
"""Record that a symbol stopped trading. True if this changed anything.
Idempotent, so the staleness path can call it every run without churning the
row: re-marking is a no-op. The one exception is an SEC confirmation landing
on a row an operator marked by hand Form 25 carries the real effective
date, so it replaces the operator's estimate. Nothing downgrades a confirmed
row back to a manual one.
"""
normalised = symbol.strip().upper()
result = await db.execute(select(Ticker).where(Ticker.symbol == normalised))
ticker = result.scalar_one_or_none()
if ticker is None:
raise NotFoundError(f"Ticker not found: {normalised}")
if ticker.delisted_on is not None:
upgrading = (
reason == REASON_FORM_25 and ticker.delisted_reason != REASON_FORM_25
)
if not upgrading:
return False
await db.execute(
update(Ticker)
.where(Ticker.id == ticker.id)
.values(delisted_on=delisted_on, delisted_reason=reason)
)
await db.commit()
logger.info(
"ticker %s marked delisted on %s (%s)", normalised, delisted_on, reason
)
return True
async def confirm_delisting(
db: AsyncSession,
symbol: str,
*,
last_bar: date | None,
today: date | None = None,
) -> date | None:
"""Ask SEC whether ``symbol`` actually delisted; mark it if so.
Called when OHLCV goes stale, because "no new bars" alone cannot tell a
delisting from a halt or a rename. Returns the effective date whenever the
symbol is known to have delisted whether this call established that or an
earlier one did and ``None`` while it remains unproven, so the caller warns
only about gaps that still have no explanation.
Returning the already-known date matters between filing and effect: trading
usually stops before the ten-day Rule 12d2-2 delay expires, so the symbol is
correctly still active (see ``active_only``) while producing no bars. Without
this the staleness warning would fire daily across that window the exact
noise the delisting flow exists to remove.
Deliberately driven by staleness rather than by the SEC fundamentals import:
that importer stalls for days at a time on unrelated Company-Facts gaps, and
detection wired into it would stall with it.
The probe waits for ``MIN_STALE_DAYS_BEFORE_PROBE``. A delisted symbol stays
stale forever, so the delay costs nothing, and it keeps a broad market-data
outage where every tracked symbol reports stale at once from turning into
one SEC request per symbol per run.
"""
from app.services.sec_client import SecError
normalised = symbol.strip().upper()
result = await db.execute(select(Ticker).where(Ticker.symbol == normalised))
ticker = result.scalar_one_or_none()
if ticker is None:
return None
known = ticker.delisted_on
# Already confirmed by SEC — nothing left to learn, but the caller still
# needs the date to know this gap is explained. A row an operator marked by
# hand is worth probing: Form 25 upgrades the estimated date.
if ticker.delisted_reason == REASON_FORM_25:
return known
if not ticker.cik:
return known
# No bars at all is an ingestion problem, not evidence of a delisting.
if last_bar is None:
return known
if ((today or date.today()) - last_bar).days < MIN_STALE_DAYS_BEFORE_PROBE:
return known
try:
async with _sec_client_factory() as client:
# Only a Form 25 filed around or after the last bar can explain THIS
# gap. An older one belongs to a class that stopped trading before
# the symbol was still printing bars, and must not retire it.
filing = await client.delisting_filing(
ticker.cik, not_before=last_bar - timedelta(days=FILING_LOOKBACK_DAYS)
)
except SecError:
# Never let a probe failure escalate a routine staleness warning.
logger.warning("delisting probe failed for %s", normalised, exc_info=True)
return known
if filing is None:
return known
# Removal takes effect ten days after filing, so the filing date is not the
# date the symbol stopped trading.
effective = filing["filing_date"] + timedelta(days=FORM_25_EFFECTIVE_DAYS)
if await mark_delisted(
db, normalised, delisted_on=effective, reason=REASON_FORM_25
):
return effective
return known
async def clear_delisted(db: AsyncSession, symbol: str) -> bool:
"""Un-retire a symbol. True if it had been marked.
The counterpart that makes automatic marking acceptable: a false positive
costs one row update, where a delete would have cost the price history.
"""
normalised = symbol.strip().upper()
result = await db.execute(select(Ticker).where(Ticker.symbol == normalised))
ticker = result.scalar_one_or_none()
if ticker is None:
raise NotFoundError(f"Ticker not found: {normalised}")
if ticker.delisted_on is None:
return False
await db.execute(
update(Ticker)
.where(Ticker.id == ticker.id)
.values(delisted_on=None, delisted_reason=None)
)
await db.commit()
logger.info("ticker %s un-marked as delisted", normalised)
return True
+22 -1
View File
@@ -357,8 +357,25 @@ async def bootstrap_universe(
db.add(Ticker(symbol=symbol))
deleted_count = 0
skipped_delisted: list[str] = []
if symbols_to_delete:
result = await db.execute(delete(Ticker).where(Ticker.symbol.in_(symbols_to_delete)))
# A delisted row was retained on purpose — its price history is exactly
# what a survivorship-honest backtest needs, and the delete cascades it
# away. Pruning must not undo that. (Pruning a symbol that is merely no
# longer an index constituent still destroys history; that needs a
# tracked/membership state separate from delisting.)
protected = (
await db.execute(
select(Ticker.symbol).where(
Ticker.symbol.in_(symbols_to_delete),
Ticker.delisted_on.is_not(None),
)
)
).scalars().all()
skipped_delisted = sorted(protected)
deletable = [s for s in symbols_to_delete if s not in set(protected)]
if deletable:
result = await db.execute(delete(Ticker).where(Ticker.symbol.in_(deletable)))
deleted_count = int(result.rowcount or 0)
await db.commit()
@@ -378,4 +395,8 @@ async def bootstrap_universe(
"already_tracked": len(target_symbols & existing_symbols),
"deleted": deleted_count,
"added_symbols": symbols_to_add[:50],
# Delisted rows a prune declined to destroy, so the caller can see the
# count did not match what they asked to remove.
"kept_delisted": skipped_delisted[:50],
"kept_delisted_count": len(skipped_delisted),
}
+11
View File
@@ -24,6 +24,17 @@ entries: the application scheduler owns both jobs.
tickers are excluded from actionable setups until a snapshot is recovered or
a later valid 10-K/10-Q supersedes the gap. Migration `028` materializes older
promoted gaps into this queue once, so setup reads never scan import history.
- A gap that survives 14 days raises `filing_gap_aged` and, from that point,
stops pausing setups **if** the issuer's own newest stored 10-K/10-Q is less
than `GAP_GATE_RECENT_FILING_DAYS` (180) old. This is the hand-off from pause
to alert, and it exists because the pause would otherwise be open-ended:
SEC's per-company Company-Facts files can go stale indefinitely (2026-08: 43
large caps whose Q2 10-Qs the `frames` API carried but whose
`companyfacts/CIK*.json` never received), and the supersede rule needs a
*successfully ingested* later filing, so a stale file swallows the next
quarter too. Retrying is unaffected — the gap stays queued and a recovered
filing still resolves it normally. An issuer with no filing that recent has no
usable fundamentals at all and stays paused.
The systemd service uses one application worker. The import framework also holds
a PostgreSQL advisory lock per source, so an overlapping manual/scheduled run is
+12
View File
@@ -214,3 +214,15 @@ it. See the [frozen specification](portfolio-capacity-bracket.md) and the
[capacity findings](portfolio-capacity-bracket-findings.md#correction-2026-08-05-ev-per-trade-was-the-wrong-lens).
The next real evidence is **forward**, not backward: the live paper-trade record.
## AI/Tech Risk Monitor
An observational risk thermometer (State + Warning) shown on the Risk page. It
gates nothing — no entries, exits, sizing or ranking — so it is not a strategy
document, but its calibration follows the same rules as one.
- [Methodology, v4](regime-monitor-v4.md) — sensors, weights, bands, and the
reasoning behind each cut from v2 onward.
- Reproduce any number in it with `scripts/run_regime_monitor_calibration.py`,
which replays the series offline and refuses to report unless it first
reproduces the published v2 and v3 figures.
+4 -320
View File
@@ -1,322 +1,6 @@
# AI/Tech Risk Monitor v3 methodology
# Moved
Named "Regime Monitor" until 2026-08-07; the filename, the `regime_monitor` job
id, the `/regime` route and the `METHODOLOGY`/snapshot fields keep the old word,
because those are persisted or externally linked. Only the wording changed.
The methodology doc now lives at [regime-monitor-v4.md](regime-monitor-v4.md).
The AI/Tech Risk Monitor is an observational risk thermometer. It does not
gate entries, exits, position size, ranking, or alerts about individual setups.
v3 supersedes v2. Every parameter below was calibrated against the 408 v2
sessions ending 2026-07-24, reproduced offline from the same Alpaca and FRED
inputs the live job uses; the reproduction matched the stored prod distribution
exactly (State avg 22.6/22.7, p80 35.1, max 91.2, P3 pegged 39, W1 live 108).
## What changed and why
**Fundamentals left the score.** F1 (capex) and F3 (good-news-stock-down)
carried 12 + 8 of 100 Warning points. Pegged at maximum stress they produced a
Warning of exactly 20.0 — below the event study's 25.3 alarm threshold, and
still inside the "stable" band. The sourced observation could not change any
published conclusion, so refreshing it looked like it did nothing. They are now
a qualitative overlay reported beside the scores. Capex also stopped scoring
`raising` and `holding` identically at 0: `holding` is the deceleration case and
now scores 50, so a boom no longer reads the same as a stall.
**The drawdown sensor stopped saturating.** v2 used `dd_pct * 5`, reaching 100 at
a 20% drawdown — the 90th percentile of the observed distribution. 39 of 408
sessions sat at exactly 100 with no resolution left, and the price pillar showed
the top band on 13.5% of sessions. v3 uses named anchors with headroom past the
observed 36% maximum, and blends leader/confirm 2:1 as P1 and P2 already did
instead of taking `max()`. P3's realized share of State falls from 65% to 40%,
matching its nominal weight.
**Warning gained a sensor with range.** The HY OAS *level* is pinned at zero
below the 3.5 mild anchor (2.77 at the cutover), so credit contributed nothing
in a calm tape. Its 20-session rate of change still does, and spread widening is
a classic lead.
**The credit percentile leg was removed.** Its reference window silently shrank
from 10 years to 3 when ICE restricted the upstream series in April 2026, after
which it scored 20 points of stress at a spread the same sensor's anchors call
"mild". See Calibration below.
**Breadth loss counts during declines.** v2's divergence gate was
`price_ret >= 0`, so the sensor zeroed during every selloff. On 2026-07-24 the
basket shed 10 points of participation in 20 sessions while SMH fell 11.9% and
Warning printed exactly 0. v3 tapers to a floor instead: deterioration counts
fully when price masks it (true divergence, the dangerous pre-top case) and at
35% when price confirms it. Breadth *level* lives in State, but breadth
*velocity* appears nowhere else, so this is not double counting.
**Bands are per axis.** v2 Warning never exceeded 64.9 in 408 sessions while
State reached 91.2, yet both used 30/60/80 with quadrant dividers at 60. The
upper half of the Warning axis was unreachable.
## Outputs
**State** — current structural stress:
- Price structure, 40%: `max(P1, P2, P3)`, one capped vote for correlated reads.
- Fixed-basket breadth level, 25%.
- HY option-adjusted credit spread level, 20%.
- VIX level, 15%.
**Warning** — deterioration and divergence:
- Fixed-basket breadth divergence, 45%.
- 60-session SMH/SPY relative-strength deterioration, 30%.
- HY OAS 20-session widening, 25%.
Combined, RSP/SPY (former F4), and the NVDA canary (former P6) do not enter v3.
## Calibration
P3 drawdown anchors, as (drawdown %, score): 0→0, 4→10, 8→25, 16→50, 28→78,
40→100, flat outside. Credit impulse is relative (+35% over 20 sessions = 100)
rather than absolute, because +0.5pp means something very different at an OAS of
2.7 than at 8.0.
Bands are round, meaning-anchored numbers, not percentile fits — percentile
thresholds would drift on every rebuild and silently rewrite what past snapshots
meant. Realized shares over the calibration window:
| Axis | stable | watch | elevated | breaking | thresholds |
|------|--------|-------|----------|----------|------------|
| State | 73.3% | 15.0% | 8.3% | 3.4% | 20 / 50 / 80 |
| Warning | 69.4% | 19.6% | 7.6% | 3.4% | 20 / 40 / 60 |
Quadrant dividers sit at each axis's watch/elevated boundary: State 50,
Warning 40.
Scores renormalize over available fixed weights, but a band is published only at
75% or greater coverage. Trend deltas are suppressed when the participating
pillar set changes. Zero means ordinary/healthy; only stress contributes.
Credit level is the named HY OAS anchors alone: 3.5 mild, 5.0 elevated, 7.0
stressed, linear between, and nothing else. v2 blended those anchors at 70% with
a 30% upper-tail percentile over a nominally 10-year window.
That leg was removed rather than repaired. ICE restricted FRED to a rolling
3-year window for `BAMLH0A0HYM2` in April 2026 — the series metadata states it
outright ("Starting in April 2026, this series will only include 3 years of
observations"), and an unbounded request returns the same 795 observations as a
30-year one. The v2 percentile therefore ranked the current spread against three
uniformly tight years (range 2.594.61 over the calibration window), which made
it fire early and saturate absurdly: at an OAS of 3.50 — the level the anchors
call *mild*, scoring zero stress — the blended sensor read 20.1, and the
percentile leg pegged at 100 by an OAS of 4.5. Across the 408 sessions it
roughly tripled the credit sensor's average (2.70 vs 1.00) and more than doubled
its nonzero days (60 vs 27).
The anchors already encode the long-run distribution as constants, so the
percentile was a second, noisier estimate of the same thing. What it was
genuinely reaching for — "unusual versus recent history" — is now W3 on the
Warning axis, computed as a rate of change, which is where deterioration
belongs. Removing it moved State's average by 0.4 and its maximum by 3.8, left
Warning bit-identical, and did not shift any band threshold.
A long-history alternative (`BAA10Y`, Fed-published, 7,712 observations back to
1997) was considered and rejected: ranking an HY spread against investment-grade
history is not a coherent statistic, and it would rescue a leg that is redundant
anyway.
Every snapshot now records `data_quality.credit_history_days` and
`vix_history_days`. This defect was invisible for roughly three months because
nothing asserted the window the code claimed; the spans make a future upstream
truncation show up in the record instead of quietly reshaping a sensor.
**Survivorship caveat.** The basket was frozen 2026-07-15 but the calibration
window reaches back to 2024, so names were partly selected for having done well.
Every distribution above inherits that bias. It is the same bias v2 carried, so
the v2/v3 comparison is like-for-like, but the absolute band shares are
optimistic.
## Point-in-time record
The first run under a new `METHODOLOGY` rebuilds the latest 400 trading sessions
with sufficient sensor warm-up; routine runs thereafter insert/update only the
latest trading date. The history API and main chart show only snapshots matching
the current methodology, so a bump reseeds the series rather than splicing two
formulas into one line.
The fundamental overlay keeps its effective date (normally the next session after
collection) and is never replayed backward, so a rebuild cannot stamp today's
observation onto historical snapshots. Because the observation is stored in a
single slot, a refresh replaces the previously effective record: the snapshot
therefore reports the overlay as `pending` until the new effective date.
Two functions, deliberately: `fundamental_overlay` is the **record** and keeps
the gate — it runs for every replayed date during a rebuild, so it must never
grow a bypass flag. `current_observation` is the **live reading** behind
`fundamental_context`, and *reports* the effective date instead of blanking the
content.
Until 2026-08-07 the live reading called the gated function, so a just-collected
observation stayed hidden until the next weekday — three days over a weekend —
and refreshing appeared to do nothing. That was the opposite of what this section
already claimed. Showing it early cannot leak into a published number, because
nothing in the overlay is scored (see "Fundamentals left the score").
`current_observation` gates on `observed` (a non-null `fetched_at`, the one field
every path writing real content stamps). Without it, the default override —
`unknown` for every hyperscaler and `mixed` for the reaction — was reported as a
live observation with `available: true`, so the card presented placeholders as a
collected reading. Those are the absence of an observation, not an observation of
absence. `fundamental_overlay` never had this problem: no observation means no
effective date, which means `pending`, which already blanks the content.
Each snapshot stores the fixed basket symbols, hash, and freeze date.
Reconstructed history before that freeze date is retrospective/exploratory.
## Presentation
The page is deliberately thin: two gauges, one chart card, one pillar table, the
overlay, and a provenance strip. Time and Path are two projections of the same
snapshot series and share one card and one query key — they were previously two
panels, which read as two datasets. Methodology rationale lives in this document,
not on the page; page text is limited to what changes how the reader interprets
today's number. The quadrant dividers rendered in Path view come from
`quadrant_config` and are the same constants the alert path consumes
(`alert_service`), so the chart cannot drift from what actually fires.
## Warning study
The study calls the outcome a **10% correction**, not a regime break. The first
70% of sessions freezes the 80th-percentile warning threshold; alarm episodes are
measured on the final 30%. Because v3 dropped fundamentals from the score, the
study now measures exactly the live Warning score rather than a technical-only
approximation of it, and both are computed from one shared sensor definition
(`warning_sensor_scores`) so they cannot drift apart.
A cached report is discarded when its methodology no longer matches, so the panel
reverts to "not run yet" after a bump rather than showing stale numbers. **Re-run
the Event Study job after cutting over to v3.**
### Reading the result
The report carries a `reliability` block and the UI renders its warnings, because
the headline numbers invite over-reading in two specific ways.
**The holdout is thin.** The study detects 11 corrections across 5 years but the
70/30 split leaves only 4 in the test period. Recall is therefore one event away
from a materially different headline, and in practice the event that flips is
decided by where the frozen threshold happens to land rather than by whether the
score saw anything. The v3 cutover run illustrates it: v3 scored 2/4 against v2's
3/4, but "v3 without the credit sensor" scores 3/4 at a *higher* threshold
(35.5) than shipped v3 misses it at (32.3) — because the alarm rule needs a
rising edge, and a lower threshold can mean the alarm already fired outside the
20-session horizon and never reset below. Below `MIN_EVENTS_FOR_CONFIDENCE`
holdout events the report says so explicitly.
Some events carry no information at all for comparison: in that run every
variant caught 2026-03-06, every variant missed 2026-06-05, and every variant
"caught" 2025-11-20 with a 1-session lead, which is coincident rather than a
warning.
**Sensor coverage can straddle the split.** The score renormalises over available
sensors, so a training window predating a sensor's history freezes the threshold
on a different construct than the holdout is measured against. At the v3 cutover
only 39% of training sessions had all three Warning sensors versus 100% of the
test period, because credit history begins 2023-07-25.
Restricting the threshold to sensor-matched training sessions was tried and is
*not* the fix: those sessions are a calm recent stretch, so the threshold drops
from 32.3 to 22.5 and false alarms rise from 3.3 to 8.6 per year. It trades a
coverage bias for a regime-selection bias. The honest position is that the
threshold is hypersensitive to window choice at this sample size; the report
states its limits rather than pretending to a precision it does not have.
## Open calibration questions
Raised 2026-08-07 during the page refactor. **None are implemented.** Each one
changes a published score, so acting on any of them means cutting `METHODOLOGY`
to v4 — which reseeds 400 sessions and discards the cached event study. They are
recorded here rather than hand-patched into v3.
**1. State's top band is a credit-event band.** `f2_credit_spreads` returns
`0.0` — not `None` — for any OAS below the 3.5 mild anchor, so credit stays
*available* at weight 20 and is not renormalized out. It is simply pinned at
zero. Verified: with price, breadth and volatility all pegged at 100 and OAS at
the cutover's 2.77, State computes to exactly **80.0** at 100% coverage — the
"breaking" threshold to the decimal. So the top State band requires either a
credit event or all three remaining pillars simultaneously at maximum. A pure
AI/Tech drawdown with calm credit — the scenario this monitor exists to
measure — cannot print it with anything to spare. Anchors-only credit was
nonzero on 27 of 408 calibration sessions, so that 20-point weight sits at zero
roughly 93% of the time. This is structurally the same defect v3 corrected on
the Warning axis ("the upper half of the Warning axis was unreachable"), and it
means the State bands were fit against a v2 credit distribution that v3 no
longer produces.
**2. V1 saturates at VIX 30.** `(vix - 15) / 15 * 100` reaches 100 at VIX 30 and
has no resolution above it: VIX 30, 50 and 82 all score identically. That is the
same failure mode, at a similar percentile, as the `dd_pct * 5` formula this
version replaced for pegging at a 20% drawdown. If addressed, it should get an
anchor table in the P3 style rather than a rescaled slope.
**3. `max(P1, P2, P3)` defeats P3's anchoring.** The `max` is deliberate ("one
capped vote for correlated reads"), but `_under_200` is binary, so P1 prints 100
whenever SMH and QQQ are both below their 200-DMA. P3's anchor ladder therefore
only resolves anything while price is *above* the 200-DMA — that is, before the
drawdown it measures is underway. Note also that "P3's realized share of State
falls from 65% to 40%" is argmax-share accounting, which is a slippery statistic
under `max()`.
## Fixed 2026-08-07: the OAS fetch window did not cover a rebuild
`HY_OAS_WINDOW_DAYS` was 400 **calendar** days, but a rebuild replays
`leader_series[-REBUILD_SESSIONS:]` — 400 **trading** sessions, about 579
calendar days. The oldest ~180 calendar days of any rebuild therefore got no OAS
data at all, so `f2_credit_spreads` and `w3_credit_impulse` both returned `None`.
Verified: State then lands at 80% coverage and Warning at exactly 75.0% —
`MIN_COVERAGE` — so **both still publish bands**. The rebuilt series would look
homogeneous while its oldest rows had been scored without credit, the tell being
a null `data_quality.credit_history_days` on exactly those rows.
The window is now 700 days: it must cover the oldest replayed date (~579) plus
W3's lookback and slack, while staying under ICE's ~3-year cap so FRED still
honours the request. This required **no methodology bump** — C1 reads
`oas_values[-1]` and W3 reads `oas_values[-21]`, both indexed from the end, so
widening only prepends older observations and every live score is bit-identical.
Confirmed by evaluating both windows against a varying synthetic series: today's
C1/W3 match exactly, while the oldest rebuild row goes from `None`/`None` to real
values.
Expect `credit_history_days` on new snapshots to rise from ~400 to ~700. That is
the widened request, not new upstream history — and it makes the chip a better
truncation canary, since a 700-day request returning ~1095 days' worth is now
the visible ceiling.
**Widening the window alone does not repair stored history.** Routine runs
recompute only the latest trading date, and `rebuilding` was keyed on "no v3
snapshot exists at all" — which is false once the cutover has run — so every row
already written would have kept its credit gap indefinitely. `SENSOR_REVISION`
fixes that: it is stamped into each snapshot, snapshots predating it read as 1,
and a stored revision below the current one triggers exactly one reseed.
It is deliberately not `METHODOLOGY`. That constant partitions the history API
and discards the cached event study; neither is warranted here, because the study
recomputes its Warning series from source (`_warning_series` calls
`warning_sensor_scores` against freshly fetched prices and OAS) rather than
reading snapshots, so a reseed cannot stale it.
The reseed is bounded by `REBUILD_LOOKBACK_DAYS` in calendar days rather than a
session count, because the binding constraint is the OAS fetch: each replayed row
needs W3's 20-business-day lookback inside `HY_OAS_WINDOW_DAYS`. At 672 days the
replay reaches ~464 sessions, W3's oldest requirement lands exactly on the first
fetched OAS day, and the ~400-session series the v3 cutover wrote is fully
covered. A test asserts that relationship so the two constants cannot drift into
recreating the gap.
The fix was sequenced deliberately: acting on items 13 above bumps
`METHODOLOGY`, which fires `rebuilding`, which would have baked the credit-less
rows into the fresh series. Fixing the window afterwards would mean reseeding
twice.
## Operator rule
Quadrant alerts default off for new/reset configurations. When enabled they
require fresh inputs, at least 75% coverage on both axes, two consecutive daily
confirmations, hysteresis, and cooldown. Every alert states: **Risk thermometer —
not a trade signal.**
v3's text is in git history (`git log --follow docs/research/regime-monitor-v4.md`).
This stub exists because commit messages up to 2026-08-08 cite the old path.
+837
View File
@@ -0,0 +1,837 @@
# AI/Tech Risk Monitor v4 methodology
Named "Regime Monitor" until 2026-08-07; the filename's `regime` stem, the
`regime_monitor` job id, the `/regime` route and the `METHODOLOGY`/snapshot
fields keep the old word, because those are persisted or externally linked.
The AI/Tech Risk Monitor is an observational risk thermometer. It does not
gate entries, exits, position size, ranking, or alerts about individual setups.
**v4 supersedes v3** (2026-08-08). Unlike v3, whose calibration was ad-hoc and
never landed, every number below is reproducible:
```
.venv/Scripts/python.exe scripts/run_regime_monitor_calibration.py --methodology v2_reconstruction,v2_reconstruction_oas400,v3,v4,v4-vix-only,v4-p1-only --cache-dir .calib-cache
```
`v3` and `v4` are mandatory — the row-wise `state_v4 <= state_v3` invariant is
a hard gate and needs both — and the replayed **start** date is asserted
against the published window. The session *count* alone proves nothing, since
the harness slices the tail of the price series to whatever was asked for.
The harness replays the 408 sessions ending 2026-07-24 from the live inputs
(Alpaca for all 33 symbols, FRED for VIX and HY OAS) with no database, and
reproduces the published v2 and v3 figures before it will emit anything:
| figure | published | replayed |
|---|---|---|
| v2 State avg | 22.6 | 22.68 |
| v2 State p80 | 35.1 | **35.1** |
| v2 State max | 91.2 | **91.2** |
| v2 P3 pegged | 39 | **39** |
| v2 W1 live | 108 | **108** |
| v3 State max | 87.4 | **87.4** |
| v3 band shares | 73.3 / 15.0 / 8.3 / 3.4 | 73.0 / 15.4 / 8.1 / 3.4 |
It refuses to emit a band recommendation, and exits non-zero, unless every hard
gate passes — 33 symbols fetched with full warm-up, the whole basket on every
session, the calendar anchors, 100% coverage on every row, and a row-wise
`state_v4 <= state_v3` invariant. Reading a calibration result out of a run whose
pipeline did not validate is meant to be structurally impossible.
## The fundamental channel (2026-08-12)
The monitor has **three channels**, not two scores with a decoration:
- **State** — current observable technical stress (price, breadth, credit, volatility).
- **Warning** — observable deterioration that may precede stress (breadth
divergence, relative strength, credit impulse).
- **Fundamental context** — a categorical state (`supportive` / `neutral` /
`adverse` / `unknown`) with an `evidence_quality` grade.
The third is **never a term in the other two**. They are read together by
confluence:
| Warning | Fundamentals | Reading |
|---|---|---|
| Calm | Supportive/neutral | Normal |
| Elevated | Supportive/neutral | Technical warning, not fundamentally confirmed |
| Calm | Adverse | Fundamental concern; tape has not confirmed |
| Elevated | Adverse | Confluence — highest attention |
`METHODOLOGY` stays **v4**: no score changed, so partitioning the history API and
discarding the event study cache would be churn. `STUDY_SCHEMA` moved to 3
instead, and is now the only thing that discards a stale report.
### Why the read is a channel and not a weight
Two things are true at once, and only this shape honours both.
**v3's reason for removing fundamentals from the score was wrong.** Not stale —
wrong. v3 argued that F1 (capex) and F3 (good-news-stock-down), carrying 12 + 8
of 100 Warning points, "could not change any published conclusion" because pegged
they produced a Warning of exactly 20.0, below the alarm threshold. That
arithmetic holds only when *every* technical sensor reads exactly zero, which is
the one case that never matters. Warning is a weighted average, so the sensors
add:
| technical Warning | without fundamentals | with them pegged | delta |
|---|---|---|---|
| 0 | 0.0 | 20.0 | +20.0 |
| 20 | 20.0 | 36.0 | +16.0 |
| 25 | 25.0 | **40.0** | +15.0 |
| 35 | 35.0 | **48.0** | +13.0 |
| 50 | 50.0 | 60.0 | +10.0 |
| 80 | 80.0 | 84.0 | +4.0 |
Pegged fundamentals lowered the technical Warning needed to reach the 40 quadrant
divider from 40 to 25. That is a 15-point shift in where the alert fires, which
is emphatically a changed conclusion. The v3 section below is kept as written,
with this correction attached, because its reasoning is cited elsewhere in this
file and a silent overwrite would hide that the error was ever made.
**But no weight is measurable either.** A weighted modifier was built and
reverted: 025 points added onto the technical Warning, sized so a maxed-out read
carried a calm tape over the 40 divider on its own. Nothing could justify the 25.
With ~10 correction events and essentially no fundamental history, any fusion
weight is a policy preference presented as a measurement — and the debate it
invites ("does the read deserve 10%, 20%, 30%?") has no evidence that can settle
it. Adding a slow categorical judgement to a fast continuous score also
manufactures precision by summing unlike things, and it forces a missing
observation to silently redistribute its weight onto the technical sensors, which
is the opposite of leaving it unknown.
So: the read gets a channel, not a coefficient. Both facts survive — the v3
removal was badly argued *and* no weight is defensible — because "report it
separately" is the only design that neither buries the observation nor invents a
number for it.
### Derivation
Deterministic, from the stored categorical facts. The LLM is an **extraction and
explanation layer**: it finds the capex guidance, classifies it, and cites it.
Fixed rules turn those facts into a state, so the same observation always yields
the same category.
`capex_signal`: any `cutting` → adverse; else any `holding` → neutral; else all
known `raising` → supportive; nothing known → unknown.
`reaction_signal`: `yes` → adverse, `mixed` → neutral, `no` → supportive,
`unknown` → unknown.
`mixed` and `unknown` are different reaction states and were merged until
2026-08-13. A failed LLM parse fell back to `mixed`, so an extraction error
became *neutral evidence* — an observation of normality manufactured out of a
bug. `mixed` now means an observed mixed reaction; anything unreadable, missing
or unattempted is `unknown` and contributes nothing.
Combined by precedence, never by averaging: **any adverse read carries**; both
unknown → unknown; every observed signal supportive → supportive; otherwise
neutral.
`unknown` is deliberately unreachable by combination. Averaging would let two
`cutting` reads and two `unknown` ones land on "neutral", presenting missing
evidence as evidence of normality — the same conflation `current_observation`
already refuses between "no observation" and "an observation of zero". Two cuts
and two unknowns read **adverse with `evidence_quality: partial`**.
`evidence_quality` is ordered by what an operator needs first: `unavailable`
(nothing collected) → `stale` (past `fundamental_staleness_days`) → `manual`
(hand override) → `complete` / `partial`.
### Presentation and alerts
The Path view colours each dot by the fundamental state recorded that day; the
axes are untouched, because context is confluence information rather than a
position on either axis. The card leads with the state and evidence grade.
Alerts stay **separate**, off one toggle:
- quadrant change — the market axes moved (existing);
- `regime_fundamental` — the context changed, e.g. neutral → adverse;
- `regime_confluence` — Warning elevated *and* fundamentals adverse.
`unknown` never alerts: an absence of evidence is not a change in the evidence,
and alerting on it would train the reader to ignore the channel. Both new
triggers seed silently on first run, as the quadrant alert does.
### The observation is now a real time series
`regime_fundamental_observations` (migration 033), one row per `effective_date`,
upserted. Before this it lived in a single `SystemSetting` slot that every
refresh overwrote, so no history existed at all — which made the read impossible
to replay, impossible to backtest, and meant a rebuild recorded every historical
session as if nothing had been observed. `update_regime_monitor` carries the
pre-existing single-slot observation into the series on its next run.
### What this does not establish
The table starts empty and fills one observation at a time, so the fundamental
rows are **untested, not failed**. Two things enforce that rather than one:
- they are **coverage-matched** — scored only on sessions where the channel had
usable context and on corrections whose warning horizon fell inside it, with a
market-only comparator over the identical window so any difference between them
is the channel and not the window;
- `measurable` stays false until `MIN_EVENTS_FOR_CONFIDENCE` corrections are
covered, and the panel prints "insufficient exposure" rather than a ratio.
Without the first, one day of coverage would render as 0/10 — recreating, one
observation later, exactly the tested-versus-unavailable confusion the flag was
added to prevent. The market rows are unchanged, and the 1/10 shipped-rule figure
remains a verdict on the technical sensors and the alert machinery alone.
The rationale for expecting the read to matter is the operator's: hyperscaler
capex is the demand side of the entire AI trade, and good earnings being sold is
a classic late-cycle tell. Both are plausible. Neither is measured here, and this
file's convention is that published numbers are reproducible.
**The path forward is accumulation, then a test — in that order.** Once enough
point-in-time observations exist, test whether the state improves prediction
*conditional on* Warning. If it does, a fitted and calibrated model has something
to fit; until then there is nothing to calibrate against. Backfilling would get
there faster: capex direction is derivable from the 10-Q/10-K capex line, which
the SEC fundamentals import already carries, and "good news, stock down" from
earnings dates plus next-day returns, which the Dolt earnings import already
carries. That last one is worth computing deterministically rather than asking
the LLM to judge, for the same reason the state derivation is rule-based.
## What changed in v4
**V1 stopped saturating at VIX 30.** `(vix - 15) / 15` reached 100 at VIX 30 —
the same defect v3 had *just* removed from P3, left in place one sensor over. VIX
30 is a bad week, 50 is a crisis and 82 was March 2020, and all three scored
identically. In the calibration window this flattened five distinct April-2025
prints (52.33, 46.98, 45.31, 40.72, 38.57) into a single 100. It pegged on 14 of
408 sessions; under the anchors below, none.
**The trend break is graded by depth, not a yes/no.** `_under_200` returned a
bare 0/100, so P1 printed 100 the moment SMH and QQQ were both under their
average — and because the price pillar takes `max(P1, P2, P3)`, that pinned the
pillar and stopped P3's anchored ladder resolving anything for the whole of a
selloff. It pegged on 46 of 408 sessions; now none. A 2% break reads ~30 where it
used to read 100.
`max()` was **kept**. The defect was the step function feeding it, not the vote
itself, and v3's "one capped vote for correlated reads" rationale still holds.
The `P1_SCORE_CAP` fallback drafted during design was to fire if P1 became the
sole price argmax on **more than 80% of sessions with State ≥ 40** — i.e. if it
had quietly become a second drawdown sensor. Measured on that population: 47
qualifying sessions, P1 sole argmax on **17 of them (36.2%)**, against P2's 16
and P3's 14. Well under the threshold, so the cap is not shipped.
**The top State band moved 80 → 65.** See Calibration; this is the one change
that is about the band rather than a sensor.
**Scope.** All three are State-side. `WARNING_BANDS`, `WARNING_WEIGHTS`,
`QUADRANT_WARNING_DIVIDER` and the event study's frozen threshold are untouched.
`QUADRANT_STATE_DIVIDER` stays 50 because only `breaking` moved.
## What changed in v3
**Fundamentals left the score.** F1 (capex) and F3 (good-news-stock-down)
carried 12 + 8 of 100 Warning points. Pegged at maximum stress they produced a
Warning of exactly 20.0 — below the event study's 25.3 alarm threshold, and
still inside the "stable" band. The sourced observation could not change any
published conclusion, so refreshing it looked like it did nothing. They are now
a qualitative overlay reported beside the scores. Capex also stopped scoring
`raising` and `holding` identically at 0: `holding` is the deceleration case and
now scores 50, so a boom no longer reads the same as a stall.
> **Corrected 2026-08-12.** The claim in this paragraph is false. "Pegged
> they produced a Warning of exactly 20.0" describes only the case where every
> technical sensor reads zero; Warning is a weighted average, so in the general
> case those 20 points added +10 to +20 and moved the technical score needed to
> reach the 40 quadrant divider from 40 to 25. The observation was removed for
> being *underweighted*, on reasoning that mistook a corner case for the whole
> range. See "The fundamental channel" above for what replaced it — a separate
> categorical channel, not a restored weight. The capex `holding` rescale in the second half
> of this paragraph stands and is still live.
**The drawdown sensor stopped saturating.** v2 used `dd_pct * 5`, reaching 100 at
a 20% drawdown — the 90th percentile of the observed distribution. 39 of 408
sessions sat at exactly 100 with no resolution left, and the price pillar showed
the top band on 13.5% of sessions. v3 uses named anchors with headroom past the
observed 36% maximum, and blends leader/confirm 2:1 as P1 and P2 already did
instead of taking `max()`. P3's realized share of State falls from 65% to 40%,
matching its nominal weight.
**Warning gained a sensor with range.** The HY OAS *level* is pinned at zero
below the 3.5 mild anchor (2.77 at the cutover), so credit contributed nothing
in a calm tape. Its 20-session rate of change still does, and spread widening is
a classic lead.
**The credit percentile leg was removed.** Its reference window silently shrank
from 10 years to 3 when ICE restricted the upstream series in April 2026, after
which it scored 20 points of stress at a spread the same sensor's anchors call
"mild". See Calibration below.
**Breadth loss counts during declines.** v2's divergence gate was
`price_ret >= 0`, so the sensor zeroed during every selloff. On 2026-07-24 the
basket shed 10 points of participation in 20 sessions while SMH fell 11.9% and
Warning printed exactly 0. v3 tapers to a floor instead: deterioration counts
fully when price masks it (true divergence, the dangerous pre-top case) and at
35% when price confirms it. Breadth *level* lives in State, but breadth
*velocity* appears nowhere else, so this is not double counting.
**Bands are per axis.** v2 Warning never exceeded 64.9 in 408 sessions while
State reached 91.2, yet both used 30/60/80 with quadrant dividers at 60. The
upper half of the Warning axis was unreachable.
## Outputs
**State** — current structural stress:
- Price structure, 40%: `max(P1, P2, P3)`, one capped vote for correlated reads.
- Fixed-basket breadth level, 25%.
- HY option-adjusted credit spread level, 20%.
- VIX level, 15%.
**Warning** — deterioration and divergence:
- Fixed-basket breadth divergence, 45%.
- 60-session SMH/SPY relative-strength deterioration, 30%.
- HY OAS 20-session widening, 25%.
**Fundamental context** — a categorical third channel, not a term in either
score. See "The fundamental channel" above.
Combined, RSP/SPY (former F4), and the NVDA canary (former P6) do not enter v3
or v4.
## Calibration
### Interpolated sensor tables
All three are `(x, stress score)` pairs read by `_interpolate`, flat outside the
first and last anchor.
| sensor | anchors |
|---|---|
| P3 drawdown (% below the 52w high) | 0→0, 4→10, 8→25, 16→50, 28→78, 40→100 |
| **P1 trend break** (% below the 200-DMA) | 0→**20**, 3→35, 8→55, 15→75, 25→100 |
| **V1 volatility** (VIX level) | 15→0, 20→20, 25→38, 30→55, 40→80, 55→100 |
P1's floor of 20 at the crossing is deliberate: the break itself is a genuine
binary event and deserves a floor; only the depth past it is graded. P1 is
calibrated to sit alongside P3 rather than swamp it — the 200-DMA lags, so a 20%
drawdown typically coincides with ~10% below the average, where P1 reads ~61
against P3's ~59.
V1 reaches full scale at 55 rather than at 2020's ~82: anchoring the top at a
once-in-a-generation print would make VIX 50 — a genuine crisis — read only ~70.
The anchors encode the long-run distribution as constants, the same argument the
credit level uses. Unlike P1 and V1, whose slopes ease off monotonically, P3's do
not (2.5, 3.75, 3.125, 2.33, 1.83) — its gentle onset is intentional and the
monotone-slope test excludes it.
Credit impulse is relative (+35% over 20 sessions = 100) rather than absolute,
because +0.5pp means something very different at an OAS of 2.7 than at 8.0.
### Bands
Round, meaning-anchored numbers, **not** percentile fits — those would drift on
every rebuild and silently rewrite what past snapshots meant.
**Why `breaking` moved 80 → 65.** With credit calm, `f2_credit_spreads` returns
`0.0` (not `None`), so it keeps its full 20 points pinned at zero. Price, breadth
and volatility at *literal maximum* therefore sum to:
(100×40 + 100×25 + 0×20 + 100×15) / 100 = 80.0 exactly
`band_for` uses `>=`, so v3's top band was reachable only by touching its floor
to the decimal, with nothing above it. The band was fit on v2, when credit's
since-removed percentile leg still contributed regularly; the sensor is not
wrong — a calm-credit selloff genuinely *is* less stressed than one with credit
contagion — the threshold was stale.
Chosen by scenario arithmetic on unchanged weights (`_scenarios` in the harness
computes these, so they are machine-checked, not prose):
| scenario | price | breadth | C1 | V1 | State |
|---|---|---|---|---|---|
| Ordinary tape (3% dd, breadth 65%, VIX 16, OAS 2.8) | 7.5 | 0 | 0 | 4.0 | **3.6** |
| 10% correction, calm credit (2% below, breadth 35%, VIX 24) | 31.2 | 62.5 | 0 | 34.4 | **33.3** |
| **2022-style drawdown, calm credit, no death cross** | 90.8 | 100 | 0 | 60.0 | **70.3** |
| **same, with death cross** (P2 pegged) | 100 | 100 | 0 | 60.0 | **74.0** |
| Credit event on top (OAS 6.0, VIX 45) | 100 | 100 | 75.0 | 86.7 | **93.0** |
| March 2020 (everything pegged) | 100 | 100 | 100 | 100 | **100** |
Rows 3 and 4 are the case this monitor exists to measure, and they must print
`breaking`. At 80 they do not. **65** clears them under either P2 assumption,
which matters because P2 is set by the 50/200-DMA gap and no drawdown figure
implies it; 70 would have left 0.33 points of headroom in row 3, reproducing the
defect being fixed.
Realized shares, **reported not fitted**, over the 408 sessions to 2026-07-24:
| Axis | stable | watch | elevated | breaking | thresholds |
|------|--------|-------|----------|----------|------------|
| State (v4) | 78.9% | 13.0% | 4.7% | **3.4%** | 20 / 50 / **65** |
| Warning | 69.4% | 19.6% | 7.6% | 3.4% | 20 / 40 / 60 |
The v4 `breaking` share lands on 3.4% — the same as v3's — having been chosen by
scenario reasoning rather than aimed at that number. Sensitivity: 60 gives 5.1%,
70 gives 1.2%.
Quadrant dividers sit at each axis's watch/elevated boundary: State 50,
Warning 40. Only `breaking` moved in v4, so the dividers and every alert
threshold are unchanged. `test_quadrant_dividers_match_the_band_boundaries` now
enforces that relationship, which nothing did before.
Scores renormalize over available fixed weights, but a band is published only at
75% or greater coverage. Trend deltas are suppressed when the participating
pillar set changes. Zero means ordinary/healthy; only stress contributes.
Credit level is the named HY OAS anchors alone: 3.5 mild, 5.0 elevated, 7.0
stressed, linear between, and nothing else. v2 blended those anchors at 70% with
a 30% upper-tail percentile over a nominally 10-year window.
That leg was removed rather than repaired. ICE restricted FRED to a rolling
3-year window for `BAMLH0A0HYM2` in April 2026 — the series metadata states it
outright ("Starting in April 2026, this series will only include 3 years of
observations"), and an unbounded request returns the same 795 observations as a
30-year one. The v2 percentile therefore ranked the current spread against three
uniformly tight years (range 2.594.61 over the calibration window), which made
it fire early and saturate absurdly: at an OAS of 3.50 — the level the anchors
call *mild*, scoring zero stress — the blended sensor read 20.1, and the
percentile leg pegged at 100 by an OAS of 4.5. Across the 408 sessions it
roughly tripled the credit sensor's average (2.70 vs 1.00) and more than doubled
its nonzero days (60 vs 27).
The anchors already encode the long-run distribution as constants, so the
percentile was a second, noisier estimate of the same thing. What it was
genuinely reaching for — "unusual versus recent history" — is now W3 on the
Warning axis, computed as a rate of change, which is where deterioration
belongs. Removing it moved State's average by 0.4 and its maximum by 3.8, left
Warning bit-identical, and did not shift any band threshold.
A long-history alternative (`BAA10Y`, Fed-published, 7,712 observations back to
1997) was considered and rejected: ranking an HY spread against investment-grade
history is not a coherent statistic, and it would rescue a leg that is redundant
anyway.
Every snapshot now records `data_quality.credit_history_days` and
`vix_history_days`. This defect was invisible for roughly three months because
nothing asserted the window the code claimed; the spans make a future upstream
truncation show up in the record instead of quietly reshaping a sensor.
**Survivorship caveat.** The basket was frozen 2026-07-15 but the calibration
window reaches back to 2024, so names were partly selected for having done well.
Every distribution above inherits that bias. It is the same bias v2 carried, so
the v2/v3 comparison is like-for-like, but the absolute band shares are
optimistic.
**Which OAS window the published v2 figures used.** v2 requested 13 years of HY
OAS and sliced `HY_OAS_REFERENCE_YEARS = 10.0` per session; ICE serves only ~3
years (778 observations from 2023-08-08), so the effective window was that. But
production v2 also fetched only 400 *calendar* days at one point — the bug fixed
2026-08-07 — and whether the published numbers predate that was not recoverable
from the text. Settled by replay rather than assumed: the
`v2_reconstruction_oas400` variant truncates the OAS **source series** to 400
days (patching the per-session window cannot simulate data that was simply
absent) and yields avg 26.54, p80 42.52, max **100.00**, against published
22.6 / 35.1 / 91.2. Full coverage reproduces all three. So the published figures
correspond to the untruncated fetch.
**The top VIX anchors are exercised, not just asserted.** The window contains a
52.33 close (2025-04-08), so the 40 → 80 → 55 → 100 segment is fed by real data
rather than justified from long-run history alone.
## Point-in-time record
The first run under a new `METHODOLOGY` rebuilds every session inside
`REBUILD_LOOKBACK_DAYS` — 672 calendar days, roughly 464 trading sessions;
routine runs thereafter insert/update only the latest trading date. The bound is
in calendar days rather than a session count because the binding constraint is
the OAS fetch: each replayed row needs W3's lookback inside
`HY_OAS_WINDOW_DAYS`, so replaying further back would recreate the credit gap a
reseed exists to close. The history API and main chart show only snapshots matching
the current methodology, so a bump reseeds the series rather than splicing two
formulas into one line.
The fundamental channel keeps its effective date (normally the next session after
collection) and is never replayed backward, so a rebuild cannot stamp today's
observation onto historical snapshots. Since the observations became a real
series (`regime_fundamental_observations`, migration 033), the effective-date
lookup *is* the gate: a replayed session gets whichever observation was live on
it, and sessions before the first one read `unknown`.
Two functions, deliberately: `fundamental_context` is the **record** and keeps
the gate — it runs for every replayed date during a rebuild, so it must never
grow a bypass flag. `current_observation` is the **live reading** behind
`fundamental_live`, and *reports* the effective date instead of blanking the
content.
Until 2026-08-07 the live reading called the gated function, so a just-collected
observation stayed hidden until the next weekday — three days over a weekend —
and refreshing appeared to do nothing. That was the opposite of what this section
already claimed. Showing it early cannot leak into a published score, because
nothing in the channel is scored.
`current_observation` gates on `observed` (a non-null `fetched_at`, the one field
every path writing real content stamps). Without it, the default override —
`unknown` for every hyperscaler and, since 2026-08-13, `unknown` for the reaction
— was reported as a live observation with `available: true`, so the card
presented placeholders as a collected reading. Those are the absence of an
observation, not an observation of absence. `fundamental_context` never had this
problem: no observation means no effective date, which means `pending`, which
already blanks the content.
**`usable` is what may confirm; `available` is only what to display.** Three
distinct things, and collapsing any two of them is a bug:
- `state` — the last thing observed. Survives going stale, so the card can show it.
- `available`*timing*: there is an effective, non-stale record to display.
- `usable`*content*: available **and** the observation actually determined
something (`state != "unknown"`).
The confluence alert and all three coverage-matched study rules gate on `usable`.
Gating on `available` instead has two failure modes, and both were live at some
point in this design:
1. a reading past `fundamental_staleness_days` would corroborate every Warning
crossing indefinitely — the strongest claim this channel makes, from the data
with the least right to make it;
2. an LLM run that failed to extract anything produces a perfectly fresh
observation that knows nothing. Counting it as exposure means repeated
extraction failures slowly accumulate coverage until the fundamental rows flip
to a *measurable* 0/8 — a failed result published for a channel that never saw
a thing, which is precisely what coverage-matching exists to prevent.
**Pre-rename snapshots are adapted, not discarded.** The channel was stored as
`fundamental_overlay` until 2026-08-12. The rename shipped without a methodology
bump — no score changed — so those rows are still served and were never reseeded.
Reading only the new key would have turned every one of them into `unknown`,
silently dropping real recorded evidence: historical Path colours, and exposure
the event study can legitimately count. `_parse_snapshot` derives the channel
from a legacy overlay's own stored facts (its capex map supplies the basket, so
the derivation uses the names observed at the time rather than today's config).
Normalising there rather than at each call site means no reader can receive an
un-adapted row. Delete only after a reseed has rewritten the whole window.
**The blob and the series row are one transaction.** They are the same
observation seen by the live card and by the point-in-time replay; committing
them separately leaves a window where a failure publishes one and not the other,
and the two then disagree permanently with nothing to detect it. Both writers use
`settings_store.upsert_setting` (which does not commit) plus a single commit;
`record_fundamental_observation` deliberately takes no commit of its own so
`update_regime_monitor` keeps its own transaction boundary.
Each snapshot stores the fixed basket symbols, hash, and freeze date.
Reconstructed history before that freeze date is retrospective/exploratory.
## Presentation
The page is deliberately thin: two gauges, one chart card, one pillar table, the
overlay, and a provenance strip. Time and Path are two projections of the same
snapshot series and share one card and one query key — they were previously two
panels, which read as two datasets. Methodology rationale lives in this document,
not on the page; page text is limited to what changes how the reader interprets
today's number. The quadrant dividers rendered in Path view come from
`quadrant_config` and are the same constants the alert path consumes
(`alert_service`), so the chart cannot drift from what actually fires.
## Warning study
The study calls the outcome a **10% correction**, not a regime break. It measures
two rules against that outcome, plus enough context to tell whether either number
is any good.
A cached report is discarded when its methodology no longer matches *or* when
`STUDY_SCHEMA` moves, so the panel reverts to "not run yet" rather than showing
stale numbers or a report missing half its blocks. **Re-run the Event Study job
after a methodology cutover or a schema bump.**
### The headline is the rule that actually fires
Until 2026-08-12 the study measured a bare rising-edge crossing of an
80th-percentile threshold fitted on the first 70% of sessions. **Nothing consumes
that rule.** What reaches Telegram is `_collect_regime_quadrant`: a quadrant
change with State ≥ 50 and Warning ≥ 40 as fixed dividers, a ±5 hysteresis
deadband, a two-session confirmation, a 3-day cooldown, and a 75% coverage gate
on both axes. The two differ on every one of those axes, including the threshold
itself (a fitted ~32 against a shipped 40).
`replay_quadrant_changes` replays the shipped state machine over the whole
sample. Three details are reproduced rather than cleaned up, because a state
machine written from first principles gets each of them wrong:
- the prior session is classified against the **current baseline**, not against
its own predecessor, so confirmation asks "did yesterday already look like this
change" rather than "did yesterday change too";
- the baseline advances only when an alert actually fires, so a change blocked by
confirmation or cooldown is re-evaluated against the old quadrant next session;
- one cooldown is shared by every quadrant change, so a 3→4 alert can swallow a
4→2 alert three days later.
Two consequences worth stating. The alarm is dated at the **confirmation**, not
at the first crossing, which costs one session of lead by construction. And the
rule alerts on changes in both directions, so the replay's exits are recorded but
filtered out by `entry_alarms` — only entering a Warning-high quadrant is a
warning about anything.
The replay reuses `_compute_index` rather than re-deriving the axes. That is the
same anti-drift argument that produced `warning_sensor_scores`: the v2 study
re-derived Warning by hand and would have kept measuring the old construct
through a scoring change. State has no equivalent shared helper, so the snapshot
builder itself is the shared definition.
**Nothing is fitted, so nothing needs protecting from a training set.** There is
no split, and every detected correction is evaluable instead of the four that
happen to land in the last 30%. The `underpowered` and "threshold frozen on a
different construct" caveats do not apply to this variant.
### Reading the result
A bare "2 of 4" is unreadable in either direction, so the report scores four more
rules through the same `evaluate_alarms` harness over the same events and
sessions, and adds a null. All use fixed thresholds — a threshold fitted on the
full sample would have lookahead the shipped rule does not, and one fitted on a
split could only be scored on the holdout events.
| kind | rules | the question |
|---|---|---|
| ablation | Warning ≥ 40 bare, State ≥ 50 bare | does the quadrant machinery earn its place? |
| baseline | leader below its 50-DMA, VIX ≥ 20 | does the score earn its complexity? |
| null | K random alarms at the observed firing rate | is any of this better than chance? |
The two kinds must not be read as one list. If a baseline matches the score, the
composite is not earning its complexity and that is the finding — it does not
mean the monitor is worthless, since State and Warning exist to be *read*, but it
caps how much further calibration is justified. If the bare Warning crossing
beats the shipped rule, the machinery (not the sensor) is what is costing recall.
The null draws only from sessions a rule could actually have fired on. Over the
whole sample it would be diluted by warm-up sessions and would understate what
chance achieves — which matters, because with ~11 events and a 20-session horizon
roughly a sixth of the sample already sits inside a hit window. It is seeded, so
a re-run cannot move the report. Corrections cluster and uniform placement does
not, so it is the **floor, not the bar**: an alarm process that clustered would
beat it for reasons unrelated to foresight.
### First result (2026-08-12): the shipped rule is not distinguishable from chance
Replayed over 2021-07-14 → 2026-08-12. The 200-DMA warm-up means the baseline
only seeds on 2022-05-26, so 1056 of 1276 sessions are evaluable and 10 of the 11
detected corrections fall inside them.
| rule | kind | warned | FA/yr | median lead |
|---|---|---|---|---|
| **Quadrant alert (shipped)** | | **1/10** | **0.9** | 19d |
| Quadrant alert, both axes high | ablation | 0/10 | 0.9 | — |
| Warning ≥ 40, bare crossing | ablation | 3/10 | 4.8 | 20d |
| State ≥ 50, bare crossing | ablation | 0/10 | 0.7 | — |
| SMH below its 50-DMA | baseline | 7/10 | 6.7 | 8d |
| VIX ≥ 20 | baseline | 4/10 | 7.2 | 9.5d |
| Random alarms, same firing rate | null | 0.9 ± 0.8 | — | — |
**P(chance ≥ 1/10) = 0.65.** Alarms scattered at random over the same sessions at
the rule's own firing rate match or beat it two times in three. Whatever the
score knows, this rule is not transmitting it.
Three readings, in order of how much they should change:
**The machinery costs more than it protects.** The bare Warning crossing catches
3 with a 20-session lead; wrapping it in the quadrant rule drops that to 1. The
State condition is the largest single cost — requiring both axes high catches
nothing at all, which is what a coincident axis gating a leading one predicts.
Hysteresis, the two-session confirmation and the shared cooldown between them
take the rest, and the cooldown is shared across *every* quadrant change, so
exits consume the budget that entries need. Only 5 of the 15 replayed changes are
Warning-high entries.
**The crude baselines beat everything on recall, at a price.** SMH below its
50-DMA catches 7 of 10 — but at 6.7 false alarms a year against the shipped
rule's 0.9. That is a 7× recall improvement for 7× the noise, so it is not a
clean dominance and this table cannot settle it; the missing axis is what a false
alarm actually costs, which nothing here measures. What it does settle is that
the composite is not buying recall the 50-DMA does not already have.
**The 0.9 false alarms/year is not the achievement it looks like.** A rule that
almost never fires has few false alarms by construction. Read the two columns
together or not at all.
Recorded from an offline replay (live Alpaca + FRED, no database, breadth
computed from the same Alpaca closes rather than the stored universe). The job in
Admin → Jobs is the canonical path and reads breadth from the DB, so re-run it to
confirm these figures before treating them as the record.
**This is a verdict on the market channels only.** The fundamental and confluence
rows in the same table are marked `measurable: false` and print "not measurable"
rather than a ratio: with an empty observation series they never fire, and a 0/10
sitting in a comparison column would read as tested-and-failed. `false` here means
the input does not exist yet, not that the rule lost.
(The figures above were also produced under a briefly-built weighted modifier and
came back bit-identical, which is what confirmed the modifier was inert over the
whole window — the numbers depend on the technical sensors alone either way.)
**Not acted on.** Nothing in the alert path was changed on the strength of this.
The obvious candidates — dropping the State condition from the entry test,
separating the entry and exit cooldowns, or lowering the Warning divider — are
threshold changes to a live alerting rule and want their own decision.
### The coverage gap relocates, it does not close
Dropping the fitted threshold makes the whole sample evaluable, but most of the
extra events predate 2023-08. W3 does not exist there, so Warning renormalises to
`(W1×45 + W2×30)/75` and the fixed 40 divider is applied to a different construct
than it was reasoned about. The report therefore splits shipped-rule metrics at
the credit sensor's first session and the panel states both, because replacing
one misleading headline with a differently misleading one would be no gain.
Convenient side effect: the pre-credit era *is* the "Warning without W3"
ablation, measured on real sessions rather than simulated ones, so that ablation
is not run separately.
Alarms and events are assigned to eras by index, so an alarm days before the
boundary matching an event days after it lands in the earlier era. With the eras
years long and the events sparse, that costs nothing.
### The fitted variant, kept for continuity
The 70/30 percentile study is still computed and still reported, collapsed, with
its `reliability` block intact — it is a genuinely different question, and it is
what earlier revisions of this document report. Its caveats stand:
**The holdout is thin.** The study detects 11 corrections across 5 years but the
70/30 split leaves only 4 in the test period. Recall is one event away from a
materially different headline, and in practice the event that flips is decided by
where the frozen threshold happens to land rather than by whether the score saw
anything. The v3 cutover run illustrates it: v3 scored 2/4 against v2's 3/4, but
"v3 without the credit sensor" scores 3/4 at a *higher* threshold (35.5) than
shipped v3 misses it at (32.3) — because the alarm rule needs a rising edge, and a
lower threshold can mean the alarm already fired outside the 20-session horizon
and never reset below. Below `MIN_EVENTS_FOR_CONFIDENCE` holdout events the
report says so explicitly.
Some events carry no information at all for comparison: in that run every
variant caught 2026-03-06, every variant missed 2026-06-05, and every variant
"caught" 2025-11-20 with a 1-session lead, which is coincident rather than a
warning. The headline recall does not currently discount those; a minimum-lead
rule is the obvious next change and has not been made.
**Sensor coverage straddles the split.** The score renormalises over available
sensors, so a training window predating a sensor's history freezes the threshold
on a different construct than the holdout is measured against. At the v3 cutover
only 39% of training sessions had all three Warning sensors versus 100% of the
test period, because credit history begins 2023-07-25.
Restricting the threshold to sensor-matched training sessions was tried and is
*not* the fix: those sessions are a calm recent stretch, so the threshold drops
from 32.3 to 22.5 and false alarms rise from 3.3 to 8.6 per year. It trades a
coverage bias for a regime-selection bias. The honest position is that a fitted
threshold is hypersensitive to window choice at this sample size — which is the
strongest argument for making the unfitted shipped rule the headline.
### Considered and not done
**An ETF credit proxy (HYG/IEF) to extend W3 back over the whole sample.** It
would trade "two sensors versus three" for "proxy sensor versus real sensor" —
still a construct straddle, but no longer flagged by the coverage split. This is
the same objection that rejected `BAA10Y` as a percentile reference. If ever
revisited, check the impulse correlation on the three years of real-OAS overlap
first and report it as a sensitivity, never as the headline.
**A depth sweep (5%/7%/15% corrections) for more events.** `EVENT_COOLDOWN_DAYS`
is 40, so at shallower thresholds re-triggers inside a single decline merge or
drop and the denominator moves for cooldown reasons rather than market ones.
## Resolved in v4 (raised 2026-08-07, shipped 2026-08-08)
The three questions this section used to hold are now answered. Kept here
because the reasoning that resolved them is not obvious from the code.
**1. `breaking` had zero headroom — resolved by moving the band, not the sensor.**
`f2_credit_spreads` returns `0.0`, not `None`, below the 3.5 mild anchor, so
credit stays *available* at weight 20 and is pinned at zero on roughly 93% of
sessions rather than being renormalized out. Price + breadth + volatility at
literal maximum therefore summed to exactly 80.0 — v3's threshold, to the
decimal.
The sensor is **deliberately unchanged**. A calm-credit selloff genuinely is less
stressed than one with credit contagion, so scoring it lower is correct; what was
stale was `STATE_BANDS`, fit on v2 while credit's since-removed percentile leg
still contributed. Making credit `None` when calm was considered and rejected: it
would leave State on 80% coverage, which still publishes, but consumes the whole
buffer — any *second* missing pillar would then suppress the band, and the 7d/30d
trend deltas would null out every time OAS crossed 3.5, because `_delta`
suppresses on a change of participating pillars. See Calibration for the
scenario arithmetic behind 65.
**2. V1 saturated at VIX 30 — resolved with an anchor table.** See "What changed
in v4".
**3. `max(P1, P2, P3)` defeated P3's anchoring — resolved by grading `_under_200`,
keeping `max()`.** The `max` was deliberate ("one capped vote for correlated
reads") and survives; the binary step feeding it was the defect.
**Its limit, stated precisely.** `_death_cross` is `clamp(-gap_pct * 20)`, so P2
pegs at a 5% 50/200-DMA gap — routine in a real downtrend. In a *deep* selloff
the price pillar therefore still reaches 100 via P2 even with P1 graded. What v4
repairs is the shallow-to-moderate break, which is where resolution was most
obviously missing: a 10% correction 2% below the average now scores 31 where v3
scored 100. It would be wrong to claim "the price pillar no longer pegs".
P2 did not peg once in the 408-session calibration window, so this is a property
of the sensor rather than an observed problem. Grading P2 the same way is the
natural next item if it starts binding; the replay reports a P2-pegged census
alongside P3 and V1 so the evidence accumulates.
## Fixed 2026-08-07: the OAS fetch window did not cover a rebuild
`HY_OAS_WINDOW_DAYS` was 400 **calendar** days, but a rebuild replays
`leader_series[-REBUILD_SESSIONS:]` — 400 **trading** sessions, about 579
calendar days. The oldest ~180 calendar days of any rebuild therefore got no OAS
data at all, so `f2_credit_spreads` and `w3_credit_impulse` both returned `None`.
Verified: State then lands at 80% coverage and Warning at exactly 75.0% —
`MIN_COVERAGE` — so **both still publish bands**. The rebuilt series would look
homogeneous while its oldest rows had been scored without credit, the tell being
a null `data_quality.credit_history_days` on exactly those rows.
The window is now 700 days: it must cover the oldest replayed date (~579) plus
W3's lookback and slack, while staying under ICE's ~3-year cap so FRED still
honours the request. This required **no methodology bump** — C1 reads
`oas_values[-1]` and W3 reads `oas_values[-21]`, both indexed from the end, so
widening only prepends older observations and every live score is bit-identical.
Confirmed by evaluating both windows against a varying synthetic series: today's
C1/W3 match exactly, while the oldest rebuild row goes from `None`/`None` to real
values.
Expect `credit_history_days` on new snapshots to rise from ~400 to ~700. That is
the widened request, not new upstream history — and it makes the chip a better
truncation canary, since a 700-day request returning ~1095 days' worth is now
the visible ceiling.
**Widening the window alone does not repair stored history.** Routine runs
recompute only the latest trading date, and `rebuilding` was keyed on "no v3
snapshot exists at all" — which is false once the cutover has run — so every row
already written would have kept its credit gap indefinitely. `SENSOR_REVISION`
fixes that: it is stamped into each snapshot, snapshots predating it read as 1,
and a stored revision below the current one triggers exactly one reseed.
It is deliberately not `METHODOLOGY`. That constant partitions the history API
and discards the cached event study; neither is warranted here, because the study
recomputes its Warning series from source (`_warning_series` calls
`warning_sensor_scores` against freshly fetched prices and OAS) rather than
reading snapshots, so a reseed cannot stale it.
The reseed is bounded by `REBUILD_LOOKBACK_DAYS` in calendar days rather than a
session count, because the binding constraint is the OAS fetch: each replayed row
needs W3's 20-business-day lookback inside `HY_OAS_WINDOW_DAYS`. At 672 days the
replay reaches ~464 sessions, W3's oldest requirement lands exactly on the first
fetched OAS day, and the ~400-session series the v3 cutover wrote is fully
covered. A test asserts that relationship so the two constants cannot drift into
recreating the gap.
The fix was sequenced deliberately: acting on items 13 above bumped
`METHODOLOGY`, which fires `rebuilding`, which would have baked the credit-less
rows into the fresh series. Fixing the window first meant the v4 reseed replayed
a clean window; doing it the other way round would have meant reseeding twice.
## Operator rule
Quadrant alerts default off for new/reset configurations. When enabled they
require fresh inputs, at least 75% coverage on both axes, two consecutive daily
confirmations, hysteresis, and cooldown. Every alert states: **Risk thermometer —
not a trade signal.**
-18
View File
@@ -1,18 +0,0 @@
<!doctype html>
<html lang="en" class="dark">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>FundamentalsPanel harness</title>
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link
href="https://fonts.googleapis.com/css2?family=Space+Grotesk:wght@400;500;600;700&family=Instrument+Sans:wght@400;500;600;700&family=IBM+Plex+Mono:wght@400;500;600&display=swap"
rel="stylesheet"
/>
</head>
<body class="bg-[#0a0b11] text-gray-100 font-sans">
<div id="root"></div>
<script type="module" src="/src/dev/harness.tsx"></script>
</body>
</html>
+23 -2
View File
@@ -207,14 +207,28 @@ export function backfillTickerNames() {
}
// Jobs
export type JobCategory = 'pipeline' | 'pipeline_step' | 'scheduled' | 'manual';
export type NextRunSource = 'own_schedule' | 'via_pipeline' | 'manual_only';
export interface JobStatus {
name: string;
label: string;
enabled: boolean;
next_run_at: string | null;
via_pipeline?: boolean;
registered: boolean;
category?: JobCategory;
/** Server-assigned ordering; the payload already arrives grouped by it. */
sort_order?: [number, number];
/** Parent pipelines for a step. Many-to-many: data_collector runs in all four. */
pipelines?: string[];
/** Step names, for a pipeline row. */
steps?: string[];
next_run_at: string | null;
next_run_source?: NextRunSource;
/** For a step: the soonest enabled parent's next run, and which parent. */
via_next_run_at?: string | null;
via_next_run_job?: string | null;
running?: boolean;
/** runtime_* is live, in-memory state only — it resets when the app restarts. */
runtime_status?: string | null;
runtime_processed?: number | null;
runtime_total?: number | null;
@@ -223,6 +237,13 @@ export interface JobStatus {
runtime_started_at?: string | null;
runtime_finished_at?: string | null;
runtime_message?: string | null;
/** last_run_* is persisted and survives restarts. Kept separate from
* runtime_* so a stale error cannot pin the status chip or the banner. */
last_run_at?: string | null;
last_run_status?: string | null;
last_run_message?: string | null;
last_run_processed?: number | null;
last_run_total?: number | null;
}
export interface TriggerJobResponse {
+1 -1
View File
@@ -6,7 +6,7 @@ import { useAuthStore } from '../stores/authStore';
* Typed error class for API errors, providing structured error handling
* across the application.
*/
export class ApiError extends Error {
class ApiError extends Error {
constructor(message: string) {
super(message);
this.name = 'ApiError';
-4
View File
@@ -34,10 +34,6 @@ export interface EquityPoint {
benchmark_pnl: number;
}
export function getEquityCurve() {
return apiClient.get<EquityPoint[]>('paper-trades/equity-curve').then((r) => r.data);
}
export interface PerfPoint {
date: string;
manual_pnl: number;
+228 -96
View File
@@ -1,4 +1,5 @@
import { useJobs, useToggleJob, useTriggerJob } from '../../hooks/useAdmin';
import type { JobCategory, JobStatus } from '../../api/admin';
import { SkeletonTable } from '../ui/Skeleton';
function formatNextRun(iso: string | null): string {
@@ -10,7 +11,8 @@ function formatNextRun(iso: string | null): string {
const mins = Math.round(diffMs / 60_000);
if (mins < 60) return `in ${mins}m`;
const hrs = Math.round(mins / 60);
return `in ${hrs}h`;
if (hrs < 48) return `in ${hrs}h`;
return `in ${Math.round(hrs / 24)}d`;
}
function formatAgo(iso: string | null | undefined): string {
@@ -29,85 +31,95 @@ function lastRunColor(status: string | null | undefined): string {
return 'text-gray-500';
}
export function JobControls() {
const { data: jobs, isLoading } = useJobs();
const toggleJob = useToggleJob();
const triggerJob = useTriggerJob();
const anyJobRunning = (jobs ?? []).some((job) => job.running);
const runningJob = jobs?.find((job) => job.running);
const pausedJob = jobs?.find((job) => !job.running && job.runtime_status === 'rate_limited');
const runningJobLabel = runningJob?.label;
if (isLoading) return <SkeletonTable rows={4} cols={3} />;
/** The four kinds of job, in the order the API already sorts them. A job whose
* category the client does not recognise still renders, under "Other" better
* a stray section than a job that silently vanishes from the admin page. */
const SECTIONS: { key: JobCategory; title: string; hint: string }[] = [
{
key: 'pipeline',
title: 'Pipelines',
hint: 'own schedule · run their steps in order',
},
{
key: 'pipeline_step',
title: 'Pipeline steps',
hint: 'no timer of their own · still triggerable individually',
},
{
key: 'scheduled',
title: 'Standalone scheduled',
hint: 'own schedule · independent of any pipeline',
},
{ key: 'manual', title: 'Manual only', hint: 'never fires on its own' },
];
/** One consistent answer per job: its own timer, its parent's, or "manual only".
* A step has no schedule of its own, so reporting one was the original bug. */
function NextRun({ job, labels }: { job: JobStatus; labels: Record<string, string> }) {
const muted = 'text-[11px] text-gray-500';
if (job.next_run_source === 'manual_only') {
return <span className={muted}>manual only</span>;
}
if (job.next_run_source === 'via_pipeline') {
if (!job.via_next_run_at || !job.via_next_run_job) {
return <span className={muted}>runs via pipeline</span>;
}
return (
<div className="space-y-3">
{runningJob && (
<div className="rounded-xl border border-blue-400/30 bg-blue-500/10 px-4 py-3">
<div className="flex flex-wrap items-center justify-between gap-3">
<div>
<div className="text-xs font-semibold text-blue-300">
Active job: {runningJob.label}
</div>
<div className="mt-0.5 text-[11px] text-blue-100/80">
Manual triggers are blocked until this run finishes.
</div>
</div>
<div className="text-[11px] text-blue-200">
{runningJob.runtime_processed ?? 0}
{typeof runningJob.runtime_total === 'number'
? ` / ${runningJob.runtime_total}`
: ''}
</div>
</div>
<div className="mt-2 h-1.5 w-full rounded-full bg-slate-700/80 overflow-hidden">
<div
className="h-full bg-blue-400 transition-all duration-500"
style={{
width: `${
typeof runningJob.runtime_progress_pct === 'number'
? Math.max(5, Math.min(100, runningJob.runtime_progress_pct))
: 30
}%`,
}}
/>
</div>
{runningJob.runtime_current_ticker && (
<div className="mt-1 text-[11px] text-blue-100/80">
Current: {runningJob.runtime_current_ticker}
</div>
)}
{runningJob.runtime_message && (
<div className="mt-1 text-[11px] text-blue-100/80">
{runningJob.runtime_message}
</div>
)}
</div>
)}
<span className={muted}>
Next via {labels[job.via_next_run_job] ?? job.via_next_run_job}{' '}
{formatNextRun(job.via_next_run_at)}
</span>
);
}
if (!job.next_run_at) return null;
return <span className={muted}>Next run {formatNextRun(job.next_run_at)}</span>;
}
{!runningJob && pausedJob && (
<div className="rounded-xl border border-amber-400/30 bg-amber-500/10 px-4 py-3">
<div className="flex flex-wrap items-center justify-between gap-3">
<div>
<div className="text-xs font-semibold text-amber-300">
Last run paused: {pausedJob.label}
/** Membership, shown rather than nested: a step can belong to several pipelines
* (data_collector is in all four), so duplicating rows under each parent would
* render Trigger buttons that are not distinct actions. */
function Membership({ job, labels }: { job: JobStatus; labels: Record<string, string> }) {
const name = (id: string) => labels[id] ?? id;
if (job.category === 'pipeline' && job.steps?.length) {
return (
<div className="mt-1 text-[11px] leading-relaxed text-gray-600">
{job.steps.map(name).join(' → ')}
</div>
<div className="mt-0.5 text-[11px] text-amber-100/90">
{pausedJob.runtime_message || 'Rate limit hit. The collector stopped early and will resume from last progress on the next run.'}
);
}
if (job.category === 'pipeline_step' && job.pipelines?.length) {
return (
<div className="mt-1 text-[11px] leading-relaxed text-gray-600">
runs in: {job.pipelines.map(name).join(', ')}
</div>
</div>
<div className="text-[11px] text-amber-200">
{pausedJob.runtime_processed ?? 0}
{typeof pausedJob.runtime_total === 'number'
? ` / ${pausedJob.runtime_total}`
: ''}
</div>
</div>
</div>
)}
);
}
return null;
}
{jobs?.map((job) => (
<div key={job.name} className="glass p-4 glass-hover">
interface JobCardProps {
job: JobStatus;
labels: Record<string, string>;
anyJobRunning: boolean;
runningJobLabel?: string;
onToggle: (job: JobStatus) => void;
onTrigger: (job: JobStatus) => void;
togglePending: boolean;
triggerPending: boolean;
}
function JobCard({
job,
labels,
anyJobRunning,
runningJobLabel,
onToggle,
onTrigger,
togglePending,
triggerPending,
}: JobCardProps) {
return (
<div className="glass p-4 glass-hover">
<div className="flex flex-wrap items-center justify-between gap-4">
<div className="flex items-center gap-3">
{/* Status dot */}
@@ -122,7 +134,9 @@ export function JobControls() {
/>
<div>
<span className="text-sm font-medium text-gray-200">{job.label}</span>
<div className="flex items-center gap-3 mt-0.5">
<div className="mt-0.5 flex flex-wrap items-center gap-3">
{/* Live state only a persisted error must not read as the
current status forever, so this never consults last_run_*. */}
<span
className={`text-[11px] font-medium ${
job.running
@@ -148,26 +162,23 @@ export function JobControls() {
? 'Active'
: 'Inactive'}
</span>
{job.via_pipeline ? (
<span className="text-[11px] text-gray-500">runs via pipeline</span>
) : (
job.enabled && job.next_run_at && (
<span className="text-[11px] text-gray-500">
Next run {formatNextRun(job.next_run_at)}
</span>
)
)}
{job.enabled && <NextRun job={job} labels={labels} />}
{!job.registered && (
<span className="text-[11px] text-red-400">Not registered</span>
)}
</div>
{!job.running && job.runtime_finished_at && (
<div className={`mt-1 text-[11px] ${lastRunColor(job.runtime_status)}`}>
Last run {formatAgo(job.runtime_finished_at)}
{job.runtime_status ? ` · ${job.runtime_status}` : ''}
{job.runtime_message ? `${job.runtime_message}` : ''}
<Membership job={job} labels={labels} />
{/* Persisted, so this survives a deploy — unlike runtime_* above. */}
{!job.running && job.last_run_at && (
<div className={`mt-1 text-[11px] ${lastRunColor(job.last_run_status)}`}>
Last run {formatAgo(job.last_run_at)}
{job.last_run_status ? ` · ${job.last_run_status}` : ''}
{job.last_run_message ? `${job.last_run_message}` : ''}
</div>
)}
{!job.running && !job.last_run_at && (
<div className="mt-1 text-[11px] text-gray-600">No run recorded yet</div>
)}
{job.running && (
<div className="mt-2 space-y-1.5">
<div className="flex items-center justify-between text-[11px] text-gray-400">
@@ -180,7 +191,7 @@ export function JobControls() {
<span>{Math.max(0, Math.min(100, job.runtime_progress_pct)).toFixed(0)}%</span>
)}
</div>
<div className="h-1.5 w-56 rounded-full bg-slate-700/80 overflow-hidden">
<div className="h-1.5 w-56 overflow-hidden rounded-full bg-slate-700/80">
<div
className="h-full bg-blue-400 transition-all duration-500"
style={{
@@ -203,8 +214,8 @@ export function JobControls() {
<div className="flex items-center gap-2">
<button
type="button"
onClick={() => toggleJob.mutate({ jobName: job.name, enabled: !job.enabled })}
disabled={toggleJob.isPending}
onClick={() => onToggle(job)}
disabled={togglePending}
className={`rounded-lg border px-3 py-1.5 text-xs transition-all duration-200 disabled:opacity-50 ${
job.enabled
? 'border-red-500/20 bg-red-500/10 text-red-400 hover:bg-red-500/20'
@@ -215,14 +226,14 @@ export function JobControls() {
</button>
<button
type="button"
onClick={() => triggerJob.mutate(job.name)}
disabled={triggerJob.isPending || !job.enabled || anyJobRunning}
className="btn-primary px-3 py-1.5 text-xs disabled:opacity-50 disabled:cursor-not-allowed"
onClick={() => onTrigger(job)}
disabled={triggerPending || !job.enabled || anyJobRunning}
className="btn-primary px-3 py-1.5 text-xs disabled:cursor-not-allowed disabled:opacity-50"
>
<span>
{job.running
? 'Running…'
: triggerJob.isPending
: triggerPending
? 'Triggering…'
: anyJobRunning
? 'Blocked'
@@ -237,7 +248,128 @@ export function JobControls() {
</div>
)}
</div>
);
}
export function JobControls() {
const { data: jobs, isLoading } = useJobs();
const toggleJob = useToggleJob();
const triggerJob = useTriggerJob();
const all = jobs ?? [];
// Job id -> display label, so a step can name its parent pipeline.
const labels = Object.fromEntries(all.map((job) => [job.name, job.label]));
const anyJobRunning = all.some((job) => job.running);
const runningJob = all.find((job) => job.running);
const pausedJob = all.find((job) => !job.running && job.runtime_status === 'rate_limited');
if (isLoading) return <SkeletonTable rows={4} cols={3} />;
const known = new Set<string>(SECTIONS.map((s) => s.key));
const groups: { key: string; title: string; hint: string; jobs: JobStatus[] }[] = [
...SECTIONS.map((section) => ({
...section,
jobs: all.filter((job) => job.category === section.key),
})),
{
key: 'other',
title: 'Other',
hint: 'uncategorised',
jobs: all.filter((job) => !job.category || !known.has(job.category)),
},
];
const cardProps = {
labels,
anyJobRunning,
runningJobLabel: runningJob?.label,
onToggle: (job: JobStatus) =>
toggleJob.mutate({ jobName: job.name, enabled: !job.enabled }),
onTrigger: (job: JobStatus) => triggerJob.mutate(job.name),
togglePending: toggleJob.isPending,
triggerPending: triggerJob.isPending,
};
return (
<div className="space-y-6">
{runningJob && (
<div className="rounded-xl border border-blue-400/30 bg-blue-500/10 px-4 py-3">
<div className="flex flex-wrap items-center justify-between gap-3">
<div>
<div className="text-xs font-semibold text-blue-300">
Active job: {runningJob.label}
</div>
<div className="mt-0.5 text-[11px] text-blue-100/80">
Manual triggers are blocked until this run finishes.
</div>
</div>
<div className="text-[11px] text-blue-200">
{runningJob.runtime_processed ?? 0}
{typeof runningJob.runtime_total === 'number'
? ` / ${runningJob.runtime_total}`
: ''}
</div>
</div>
<div className="mt-2 h-1.5 w-full overflow-hidden rounded-full bg-slate-700/80">
<div
className="h-full bg-blue-400 transition-all duration-500"
style={{
width: `${
typeof runningJob.runtime_progress_pct === 'number'
? Math.max(5, Math.min(100, runningJob.runtime_progress_pct))
: 30
}%`,
}}
/>
</div>
{runningJob.runtime_current_ticker && (
<div className="mt-1 text-[11px] text-blue-100/80">
Current: {runningJob.runtime_current_ticker}
</div>
)}
{runningJob.runtime_message && (
<div className="mt-1 text-[11px] text-blue-100/80">{runningJob.runtime_message}</div>
)}
</div>
)}
{!runningJob && pausedJob && (
<div className="rounded-xl border border-amber-400/30 bg-amber-500/10 px-4 py-3">
<div className="flex flex-wrap items-center justify-between gap-3">
<div>
<div className="text-xs font-semibold text-amber-300">
Last run paused: {pausedJob.label}
</div>
<div className="mt-0.5 text-[11px] text-amber-100/90">
{pausedJob.runtime_message || 'Rate limit hit. The collector stopped early and will resume from last progress on the next run.'}
</div>
</div>
<div className="text-[11px] text-amber-200">
{pausedJob.runtime_processed ?? 0}
{typeof pausedJob.runtime_total === 'number'
? ` / ${pausedJob.runtime_total}`
: ''}
</div>
</div>
</div>
)}
{groups.map(
(group) =>
group.jobs.length > 0 && (
<section key={group.key} className="space-y-3">
<h3 className="text-xs font-medium uppercase tracking-widest text-gray-500">
{group.title}
<span className="ml-2 num text-gray-600">{group.jobs.length}</span>
<span className="ml-2 normal-case tracking-normal text-gray-600">
{group.hint}
</span>
</h3>
{group.jobs.map((job) => (
<JobCard key={job.name} job={job} {...cardProps} />
))}
</section>
),
)}
</div>
);
}
@@ -11,6 +11,8 @@ const DEFAULTS: ScheduleConfig = {
schedule_near_close_pipeline_cron: '30 15 * * mon-fri',
schedule_after_close_pipeline_cron: '45 16 * * mon-fri',
schedule_intraday_pipeline_cron: '0 10-15 * * mon-fri',
schedule_backtest_cron: '0 3 * * sun',
schedule_ticker_universe_cron: '0 1 * * *',
};
const FIELDS: { key: keyof ScheduleConfig; label: string; hint: string; mono?: boolean }[] = [
@@ -55,6 +57,18 @@ const FIELDS: { key: keyof ScheduleConfig; label: string; hint: string; mono?: b
hint: 'Refresh prices + resolve outcomes mid-session. Default hourly 10:0015:00 ET weekdays.',
mono: true,
},
{
key: 'schedule_backtest_cron',
label: 'Backtest',
hint: 'Replay history and refresh the Track Record report. Default Sunday 03:00 ET. Was a 168h interval, which restarted on every deploy and so could defer indefinitely.',
mono: true,
},
{
key: 'schedule_ticker_universe_cron',
label: 'Ticker universe sync',
hint: 'Refresh the tracked-symbol universe. Default 01:00 ET daily, before the morning pipeline.',
mono: true,
},
];
export function ScheduleSettings() {
+147 -46
View File
@@ -2,7 +2,6 @@ import { useMemo, useState } from 'react';
import { useQuery } from '@tanstack/react-query';
import {
CartesianGrid,
Cell,
Line,
LineChart,
ReferenceArea,
@@ -19,6 +18,8 @@ import { getRegimeHistory, getRegimeMonitor } from '../../api/regime';
import { Callout } from '../ui/Callout';
import { SkeletonCard } from '../ui/Skeleton';
import { formatDate } from '../../lib/format';
import { FUNDAMENTAL_VISUAL, QUADRANT_WASH, REGIME_VISUAL } from '../../lib/regime';
import type { EvidenceQuality, FundamentalState } from '../../lib/types';
// Lazy-loaded (see RegimePage) so recharts stays in the regime-tab chunk.
// Time and Path are two projections of one series, so they share a card and a
@@ -38,10 +39,16 @@ type RangeKey = (typeof RANGES)[number]['key'];
/** Sessions drawn in Path view. The full series is unreadable as a path. */
const PATH_TRAIL = 60;
const STATE_COLOR = '#60a5fa';
const WARNING_COLOR = '#fb923c';
const STATE_COLOR = REGIME_VISUAL.state;
const WARNING_COLOR = REGIME_VISUAL.warning;
const FUNDAMENTAL_SYMBOL: Record<FundamentalState, string> = {
supportive: '▲',
neutral: '●',
adverse: '◆',
unknown: '○',
};
// Fall back to the v3 constants, not v2's shared 60/60, so a missing
// Fall back to the shipped constants, not v2's shared 60/60, so a missing
// quadrant_config cannot draw dividers that disagree with the alert path.
const DEFAULT_STATE_DIVIDER = 50;
const DEFAULT_WARNING_DIVIDER = 40;
@@ -50,13 +57,19 @@ interface PathPoint {
x: number;
y: number;
date: string;
/** The third channel as recorded that day. Colours the dot; never moves it. */
fundamental: FundamentalState;
evidence: EvidenceQuality;
/** Raw dated observations are interactive dots; the smoothed copy is line-only. */
raw: boolean;
recency: number;
}
/** Centered moving average to de-noise the path; today (last) kept exact. */
function smoothTrail(points: PathPoint[], half = 2): PathPoint[] {
const n = points.length;
return points.map((p, i) => {
if (i === n - 1) return { ...p };
if (i === n - 1) return { ...p, raw: false };
let sx = 0;
let sy = 0;
let c = 0;
@@ -65,14 +78,58 @@ function smoothTrail(points: PathPoint[], half = 2): PathPoint[] {
sy += points[j].y;
c += 1;
}
return { x: sx / c, y: sy / c, date: p.date };
return { ...p, x: sx / c, y: sy / c, raw: false };
});
}
/** Recency gradient: 0 = oldest (muted slate), 1 = newest (bright blue). */
function recencyColor(t: number): string {
const lerp = (a: number, b: number) => Math.round(a + (b - a) * t);
return `rgba(${lerp(71, 96)}, ${lerp(85, 165)}, ${lerp(105, 250)}, ${(0.3 + 0.7 * t).toFixed(2)})`;
function FundamentalGlyph({
cx,
cy,
state,
size,
opacity = 1,
}: {
cx: number;
cy: number;
state: FundamentalState;
size: number;
opacity?: number;
}) {
const visual = FUNDAMENTAL_VISUAL[state] ?? FUNDAMENTAL_VISUAL.unknown;
const common = { fill: visual.color, opacity, stroke: '#11131c', strokeWidth: 1 };
if (visual.glyph === 'up') {
return <polygon points={`${cx},${cy - size} ${cx - size},${cy + size} ${cx + size},${cy + size}`} {...common} />;
}
if (visual.glyph === 'diamond') {
return <polygon points={`${cx},${cy - size} ${cx - size},${cy} ${cx},${cy + size} ${cx + size},${cy}`} {...common} />;
}
if (visual.glyph === 'ring') {
return <circle cx={cx} cy={cy} r={size - 0.5} fill="transparent" opacity={opacity} stroke={visual.color} strokeWidth={1.5} />;
}
return <circle cx={cx} cy={cy} r={size - 0.5} {...common} />;
}
function PathPointShape({ cx = 0, cy = 0, payload }: { cx?: number; cy?: number; payload?: PathPoint }) {
if (!payload) return <g />;
return (
<FundamentalGlyph
cx={cx}
cy={cy}
state={payload.fundamental}
size={3.25 + payload.recency * 1.25}
opacity={0.58 + payload.recency * 0.42}
/>
);
}
function LatestPointShape({ cx = 0, cy = 0, payload }: { cx?: number; cy?: number; payload?: PathPoint }) {
if (!payload) return <g />;
return (
<g>
<circle cx={cx} cy={cy} r={7} fill="transparent" stroke="#ffffff" strokeWidth={1.75} />
<FundamentalGlyph cx={cx} cy={cy} state={payload.fundamental} size={4.5} />
</g>
);
}
function SegmentedControl<T extends string>({
@@ -94,8 +151,8 @@ function SegmentedControl<T extends string>({
type="button"
aria-pressed={value === option}
onClick={() => onChange(option)}
className={`rounded px-2 py-1 text-[11px] font-medium tabular-nums transition-colors ${
value === option ? 'bg-white/10 text-blue-300' : 'text-gray-500 hover:text-gray-300'
className={`min-h-9 rounded px-3 py-2 text-xs font-medium tabular-nums transition-colors ${
value === option ? 'bg-white/10 text-blue-300' : 'text-gray-400 hover:text-gray-200'
}`}
>
{option}
@@ -107,14 +164,19 @@ function SegmentedControl<T extends string>({
function PathTip({ active, payload }: { active?: boolean; payload?: { payload: PathPoint }[] }) {
if (!active || !payload?.length) return null;
const p = payload[0].payload;
const p = payload.find((item) => item.payload.raw)?.payload ?? payload[0].payload;
const visual = FUNDAMENTAL_VISUAL[p.fundamental] ?? FUNDAMENTAL_VISUAL.unknown;
const evidence = p.evidence === 'unavailable' ? 'Unavailable' : `${p.evidence.replace(/_/g, ' ')} evidence`;
return (
<div className="glass px-2.5 py-1.5 text-[11px]">
<div className="glass px-3 py-2 text-xs">
<div className="text-gray-300">{formatDate(p.date)}</div>
<div className="text-gray-400">
State <span style={{ color: STATE_COLOR }}>{Math.round(p.x)}</span> · Warning{' '}
<span style={{ color: WARNING_COLOR }}>{Math.round(p.y)}</span>
</div>
<div className="mt-0.5 text-gray-400">
Fundamentals <span style={{ color: visual.color }}>{visual.label}</span> · {evidence}
</div>
</div>
);
}
@@ -144,7 +206,15 @@ export default function RegimeChart() {
}, [history.data, view, range]);
const pathPoints = useMemo<PathPoint[]>(
() => series.map((p) => ({ x: p.state as number, y: p.warning as number, date: p.date })),
() => series.map((p, index, points) => ({
x: p.state as number,
y: p.warning as number,
date: p.date,
fundamental: p.fundamental_state ?? 'unknown',
evidence: p.evidence_quality ?? 'unavailable',
raw: true,
recency: points.length <= 1 ? 1 : index / (points.length - 1),
})),
[series],
);
const trail = useMemo(() => (view === 'Path' ? smoothTrail(pathPoints) : []), [pathPoints, view]);
@@ -159,7 +229,7 @@ export default function RegimeChart() {
<div className="glass p-5">
<div className="flex flex-wrap items-center justify-between gap-3">
<div className="flex items-center gap-3">
<span className="text-[11px] uppercase tracking-wider text-gray-500">
<span className="text-xs uppercase tracking-wider text-gray-400">
{view === 'Time' ? 'State & Warning over time' : `State × Warning path · last ${PATH_TRAIL} sessions`}
</span>
<SegmentedControl options={VIEWS} value={view} onChange={setView} label="Chart view" />
@@ -168,7 +238,7 @@ export default function RegimeChart() {
<SegmentedControl options={RANGES.map((r) => r.key)} value={range} onChange={setRange} label="Time range" />
) : (
latest && (
<span className="text-[11px] text-gray-500">
<span className="text-xs text-gray-400">
now: State <span style={{ color: STATE_COLOR }}>{Math.round(latest.x)}</span> · Warning{' '}
<span style={{ color: WARNING_COLOR }}>{Math.round(latest.y)}</span>
</span>
@@ -182,14 +252,18 @@ export default function RegimeChart() {
<Callout variant="empty">Not enough coverage-qualified history yet it accumulates as the daily job runs.</Callout>
) : (
<>
<div className="mt-3 h-72">
<div
className="mt-3 h-72"
role="img"
aria-label={view === 'Time' ? 'State and Warning scores over time' : 'State by Warning path with fundamental context symbols'}
>
<ResponsiveContainer width="100%" height="100%">
{view === 'Time' ? (
<LineChart data={series} margin={{ top: 6, right: 8, left: 0, bottom: 0 }}>
<CartesianGrid stroke="rgba(255,255,255,0.05)" vertical={false} />
<XAxis
dataKey="date"
tick={{ fill: '#6b7280', fontSize: 10 }}
tick={{ fill: '#9aa0b0', fontSize: 10 }}
tickFormatter={(d) => formatDate(String(d))}
minTickGap={28}
tickLine={false}
@@ -200,7 +274,7 @@ export default function RegimeChart() {
<YAxis
domain={[0, 100]}
ticks={[0, 25, 50, 75, 100]}
tick={{ fill: '#6b7280', fontSize: 10 }}
tick={{ fill: '#9aa0b0', fontSize: 10 }}
width={34}
tickLine={false}
axisLine={false}
@@ -225,10 +299,13 @@ export default function RegimeChart() {
</LineChart>
) : (
<ScatterChart margin={{ top: 10, right: 16, bottom: 22, left: 0 }}>
<ReferenceArea x1={0} x2={xDiv} y1={yDiv} y2={100} fill="#f59e0b" fillOpacity={0.07} stroke="none" />
<ReferenceArea x1={xDiv} x2={100} y1={yDiv} y2={100} fill="#f97316" fillOpacity={0.07} stroke="none" />
<ReferenceArea x1={0} x2={xDiv} y1={0} y2={yDiv} fill="#10b981" fillOpacity={0.07} stroke="none" />
<ReferenceArea x1={xDiv} x2={100} y1={0} y2={yDiv} fill="#ef4444" fillOpacity={0.08} stroke="none" />
{/* One neutral at four opacities: denser = more axes elevated.
Hue here would collide with the fundamental glyphs drawn
on top of it see QUADRANT_WASH. */}
<ReferenceArea x1={0} x2={xDiv} y1={yDiv} y2={100} fill="#ffffff" fillOpacity={QUADRANT_WASH.early_warning} stroke="none" />
<ReferenceArea x1={xDiv} x2={100} y1={yDiv} y2={100} fill="#ffffff" fillOpacity={QUADRANT_WASH.active_stress} stroke="none" />
<ReferenceArea x1={0} x2={xDiv} y1={0} y2={yDiv} fill="#ffffff" fillOpacity={QUADRANT_WASH.healthy} stroke="none" />
<ReferenceArea x1={xDiv} x2={100} y1={0} y2={yDiv} fill="#ffffff" fillOpacity={QUADRANT_WASH.stabilizing} stroke="none" />
<CartesianGrid stroke="rgba(255,255,255,0.04)" />
<ReferenceLine x={xDiv} stroke="rgba(255,255,255,0.12)" />
<ReferenceLine y={yDiv} stroke="rgba(255,255,255,0.12)" />
@@ -237,36 +314,37 @@ export default function RegimeChart() {
dataKey="x"
domain={[0, 100]}
ticks={[0, 20, 40, 60, 80, 100]}
tick={{ fill: '#6b7280', fontSize: 10 }}
tick={{ fill: '#9aa0b0', fontSize: 10 }}
tickLine={false}
axisLine={{ stroke: 'rgba(255,255,255,0.08)' }}
label={{ value: 'State →', position: 'insideBottom', offset: -12, fill: '#6b7280', fontSize: 10 }}
label={{ value: 'State →', position: 'insideBottom', offset: -12, fill: '#9aa0b0', fontSize: 10 }}
/>
<YAxis
type="number"
dataKey="y"
domain={[0, 100]}
ticks={[0, 20, 40, 60, 80, 100]}
tick={{ fill: '#6b7280', fontSize: 10 }}
tick={{ fill: '#9aa0b0', fontSize: 10 }}
width={30}
tickLine={false}
axisLine={false}
label={{ value: 'Warning', angle: -90, position: 'insideLeft', fill: '#6b7280', fontSize: 10 }}
label={{ value: 'Warning', angle: -90, position: 'insideLeft', fill: '#9aa0b0', fontSize: 10 }}
/>
<ZAxis range={[13, 13]} />
<ZAxis range={[18, 18]} />
<Tooltip cursor={{ strokeDasharray: '3 3', stroke: 'rgba(255,255,255,0.2)' }} content={<PathTip />} />
<Scatter data={trail} line={{ stroke: 'rgba(96,165,250,0.18)', strokeWidth: 1.5 }} isAnimationActive={false}>
{trail.map((_, i) => (
<Cell key={i} fill={recencyColor(trail.length <= 1 ? 1 : i / (trail.length - 1))} />
))}
</Scatter>
<Scatter
data={trail}
line={{ stroke: 'rgba(255,255,255,0.18)', strokeWidth: 1.5 }}
shape={(props: { cx?: number; cy?: number }) => <circle cx={props.cx} cy={props.cy} r={0} />}
tooltipType="none"
isAnimationActive={false}
/>
<Scatter data={pathPoints} shape={<PathPointShape />} isAnimationActive={false} />
{latest && (
<Scatter
data={[latest]}
isAnimationActive={false}
shape={(props: { cx?: number; cy?: number }) => (
<circle cx={props.cx} cy={props.cy} r={6} fill="#ffffff" stroke={STATE_COLOR} strokeWidth={2} />
)}
shape={<LatestPointShape />}
/>
)}
</ScatterChart>
@@ -275,7 +353,7 @@ export default function RegimeChart() {
</div>
{view === 'Time' ? (
<div className="mt-2 flex flex-wrap items-center gap-4 text-[11px] text-gray-400">
<div className="mt-2 flex flex-wrap items-center gap-4 text-xs text-gray-400">
<span className="flex items-center gap-1.5">
<span className="inline-block h-2 w-3 rounded-sm" style={{ background: STATE_COLOR }} />
State
@@ -284,20 +362,43 @@ export default function RegimeChart() {
<span className="inline-block h-2 w-3 rounded-sm" style={{ background: WARNING_COLOR }} />
Warning
</span>
<span className="text-gray-600">dashed = each axis's elevated threshold ({xDiv} / {yDiv})</span>
<span className="text-gray-400">dashed = each axis's elevated threshold ({xDiv} / {yDiv})</span>
</div>
) : (
<div className="mt-2 grid grid-cols-1 gap-x-4 gap-y-1 text-[11px] text-gray-500 sm:grid-cols-2">
<span><span className="text-amber-400">Early warning</span> calm, fragility rising</span>
<span><span className="text-orange-400">Active stress</span> damaged and deteriorating</span>
<span><span className="text-emerald-400">Healthy</span> calm, broadly supported</span>
<span><span className="text-red-400">Stabilizing</span> damage remains, warning lower</span>
<span className="text-gray-600 sm:col-span-2">White dot = today; trail brightens toward the present, smoothed.</span>
<div className="mt-2 grid grid-cols-1 gap-x-4 gap-y-1 text-xs text-gray-400 sm:grid-cols-2">
{/* Swatches, not coloured words: the quadrant names used the
fundamental channel's colours, so "Stabilizing" was rendered in
the adverse hue while meaning damage receding. */}
{([
['active_stress', 'Active stress', 'damaged and deteriorating'],
['early_warning', 'Early warning', 'calm, fragility rising'],
['stabilizing', 'Stabilizing', 'damage remains, warning lower'],
['healthy', 'Healthy', 'calm, broadly supported'],
] as const).map(([key, name, gloss]) => (
<span key={key} className="flex items-center gap-1.5">
<span
aria-hidden="true"
className="inline-block h-3 w-3 shrink-0 rounded-sm border border-white/10"
style={{ background: `rgba(255,255,255,${QUADRANT_WASH[key] * 4})` }}
/>
<span className="text-gray-300">{name}</span> {gloss}
</span>
))}
<span className="text-gray-400 sm:col-span-2">Raw dated points grow toward today; the connecting line is smoothed. White ring = today.</span>
<span className="mt-1 flex flex-wrap items-center gap-x-3 gap-y-1 text-gray-400 sm:col-span-2">
<span>symbol + colour = fundamentals:</span>
{(['supportive', 'neutral', 'adverse', 'unknown'] as const).map((state) => (
<span key={state} className="inline-flex items-center gap-1.5">
<span aria-hidden="true" style={{ color: FUNDAMENTAL_VISUAL[state].color }}>{FUNDAMENTAL_SYMBOL[state]}</span>
{FUNDAMENTAL_VISUAL[state].label}
</span>
))}
</span>
</div>
)}
{crossesFreeze && (
<p className="mt-2 text-[11px] text-gray-600">
<p className="mt-2 text-xs text-gray-400">
History before {basketAsOf} is reconstructed against today's basket retrospective, not a live record.
</p>
)}
+109 -346
View File
@@ -9,35 +9,8 @@ import { Disclosure } from '../ui/Disclosure';
import { Dropdown } from '../ui/Dropdown';
import { Section } from '../ui/Section';
import { useToast } from '../ui/Toast';
import type { BacktestCurvePoint, BacktestPortfolioMonitorRun } from '../../lib/types';
function fmtR(v: number | null | undefined): string {
if (v === null || v === undefined) return '—';
return `${v > 0 ? '+' : ''}${v.toFixed(2)}R`;
}
function fmtPct(v: number | null): string {
return v === null ? '—' : `${v.toFixed(1)}%`;
}
function fmtMoney(v: number | null | undefined): string {
if (v === null || v === undefined) return '—';
return v.toLocaleString('en-US', { minimumFractionDigits: 2, maximumFractionDigits: 2 });
}
function fmtSignedPct(v: number | null | undefined): string {
if (v === null || v === undefined) return '—';
return `${v > 0 ? '+' : ''}${v.toFixed(1)}%`;
}
function fmtDrawdown(v: number | null | undefined): string {
return v === null || v === undefined ? '—' : `-${Math.abs(v).toFixed(1)}%`;
}
function fmtDays(v: number | null | undefined): string {
return v === null || v === undefined ? '—' : `${v.toFixed(1)}d`;
}
function rColor(v: number | null): string {
if (v === null) return 'text-gray-400';
if (v > 0) return 'text-emerald-400';
if (v < 0) return 'text-red-400';
return 'text-gray-300';
}
import { BacktestRecommendationCard } from './BacktestRecommendationCard';
import { PortfolioMonitorPanel } from './PortfolioMonitorPanel';
function timeAgo(iso: string): string {
const mins = Math.floor((Date.now() - new Date(iso).getTime()) / 60_000);
@@ -48,95 +21,14 @@ function timeAgo(iso: string): string {
return `${Math.floor(hrs / 24)}d ago`;
}
function Stat({ label, value, valueClass = 'text-gray-100', sub }: {
label: string; value: string; valueClass?: string; sub?: string;
}) {
return (
<div className="glass p-4">
<p className="section-index">{label}</p>
<p className={`num mt-1.5 text-2xl font-semibold ${valueClass}`}>{value}</p>
{sub && <p className="mt-1 text-xs text-gray-500">{sub}</p>}
</div>
);
}
function curvePath(
points: BacktestCurvePoint[],
min: number,
max: number,
w: number,
h: number,
pad: number,
startMs: number,
endMs: number,
): string {
if (points.length < 2) return '';
const span = Math.max(max - min, 1);
const timeSpan = Math.max(endMs - startMs, 1);
return points
.map((p, i) => {
const t = new Date(p.date).getTime();
const x = pad + ((t - startMs) / timeSpan) * (w - pad * 2);
const value = p.return_pct ?? 0;
const y = pad + (1 - (value - min) / span) * (h - pad * 2);
return `${i === 0 ? 'M' : 'L'}${x.toFixed(1)},${y.toFixed(1)}`;
})
.join(' ');
}
function EquityCurveChart({ run }: { run: BacktestPortfolioMonitorRun }) {
const portfolio = run.equity_curve ?? [];
const benchmark = run.benchmark_curve ?? [];
const values = [...portfolio, ...benchmark]
.map((p) => p.return_pct)
.filter((v): v is number => v !== null && v !== undefined);
if (portfolio.length < 2 || values.length === 0) {
return <Callout variant="empty">No equity curve points for this selection.</Callout>;
}
const min = Math.min(0, ...values);
const max = Math.max(0, ...values);
const times = [...portfolio, ...benchmark]
.map((p) => new Date(p.date).getTime())
.filter((v) => Number.isFinite(v));
if (times.length === 0) {
return <Callout variant="empty">No dated equity curve points for this selection.</Callout>;
}
const startMs = Math.min(...times);
const endMs = Math.max(...times);
const w = 720;
const h = 240;
const pad = 28;
const portfolioPath = curvePath(portfolio, min, max, w, h, pad, startMs, endMs);
const benchmarkPath = curvePath(benchmark, min, max, w, h, pad, startMs, endMs);
const lastPortfolio = portfolio[portfolio.length - 1]?.return_pct ?? null;
const lastBenchmark = benchmark[benchmark.length - 1]?.return_pct ?? run.spy_return_pct;
return (
<div className="glass overflow-hidden">
<div className="flex flex-wrap items-center justify-between gap-3 border-b border-white/[0.05] px-4 py-3">
<div>
<p className="text-sm font-semibold text-gray-100">{run.label}</p>
<p className="text-[11px] text-gray-500">{run.start_date} - {run.end_date}</p>
</div>
<div className="flex gap-4 text-xs">
<span className="text-blue-300">Portfolio {fmtSignedPct(lastPortfolio)}</span>
<span className="text-gray-400">S&P 500 {fmtSignedPct(lastBenchmark)}</span>
</div>
</div>
<svg viewBox={`0 0 ${w} ${h}`} className="h-64 w-full" role="img" aria-label="Portfolio return compared with S&P 500">
<line x1={pad} y1={h - pad} x2={w - pad} y2={h - pad} stroke="rgba(255,255,255,0.12)" />
<line x1={pad} y1={pad} x2={pad} y2={h - pad} stroke="rgba(255,255,255,0.12)" />
{benchmarkPath && (
<path d={benchmarkPath} fill="none" stroke="rgba(156,163,175,0.9)" strokeWidth="2" strokeDasharray="5 5" />
)}
<path d={portfolioPath} fill="none" stroke="rgb(96,165,250)" strokeWidth="3" />
<text x={pad} y={pad - 8} className="fill-gray-500 text-[10px]">{fmtSignedPct(max)}</text>
<text x={pad} y={h - 8} className="fill-gray-500 text-[10px]">{fmtSignedPct(min)}</text>
</svg>
</div>
);
}
const TARGET_MODEL_OPTIONS = [
{ value: 'production_gtl', label: 'Live GTL — production' },
{ value: 'structural_sr', label: 'Structural S/R — comparison' },
];
const CADENCE_OPTIONS = [
{ value: 'weekly', label: 'Weekly — default' },
{ value: 'daily', label: 'Daily — research' },
];
export function BacktestPanel() {
const { data: report, isLoading } = useBacktestReport();
@@ -150,8 +42,19 @@ export function BacktestPanel() {
const monitor = report?.portfolio_monitor ?? null;
const activeStrategy =
selectedStrategy || monitor?.production_strategy || monitor?.strategies[0]?.strategy || '';
// Default to the window the recommendation was computed on, so the tiles and
// the recommendation never open showing different numbers. They used to: the
// backend preferred "all" while this defaulted to "3y". The 3y fallback is
// only for reports predating basis_lookback.
const basisLookback = report?.recommendation?.basis_lookback ?? null;
const activeLookback =
selectedLookback || (monitor?.lookbacks.some((l) => l.lookback === '3y') ? '3y' : monitor?.lookbacks[0]?.lookback) || '';
selectedLookback ||
(basisLookback && monitor?.lookbacks.some((l) => l.lookback === basisLookback)
? basisLookback
: monitor?.lookbacks.some((l) => l.lookback === '3y')
? '3y'
: monitor?.lookbacks[0]?.lookback) ||
'';
const monitorRun = useMemo(
() =>
monitor?.runs.find((row) => row.strategy === activeStrategy && row.lookback === activeLookback) ??
@@ -178,7 +81,61 @@ export function BacktestPanel() {
return (
<Section title="Is the strategy working?" hint="portfolio simulation of the promoted strategy vs S&P 500">
<div className="space-y-4">
<div className="flex flex-wrap items-start justify-between gap-3">
{/* Run status and the controls that start a new run, on one line. The
explainer sits BELOW this row rather than beside it sharing a flex
row meant expanding it shoved every control down the page. */}
<div className="flex flex-wrap items-end justify-between gap-3">
<div className="min-w-0">
<p className="section-index">Last run</p>
{report ? (
<p className="mt-1 text-xs text-gray-400">
{timeAgo(report.generated_at)} · {report.tickers} tickers ·{' '}
{report.candidates} setups ({report.qualified} qualified) ·{' '}
{report.params.entry_cadence ?? 'weekly'},{' '}
{report.params.horizon_days}d horizon
{report.params.cost_per_side_pct != null && (
<> · net of {report.params.cost_per_side_pct}%/side</>
)}
{' · '}
<span className={report.params.is_production_target_model === false ? 'text-amber-300' : 'text-blue-300'}>
{report.params.target_model_label ?? 'Unknown (legacy report)'}
</span>
</p>
) : (
<p className="mt-1 text-xs text-gray-500">Never run</p>
)}
</div>
{/* flex-wrap is load-bearing: two dropdowns plus the button overflow a
narrow viewport otherwise. */}
<div className="flex flex-wrap items-end gap-2">
<div className="flex flex-col gap-1 text-[11px] uppercase tracking-wider text-gray-500">
<label htmlFor="backtest-target-model">Target model</label>
<Dropdown
id="backtest-target-model"
className="w-56 normal-case tracking-normal"
value={targetModel}
onChange={(v) => setTargetModel(v as BacktestTargetModel)}
options={TARGET_MODEL_OPTIONS}
/>
</div>
<div className="flex flex-col gap-1 text-[11px] uppercase tracking-wider text-gray-500">
<label htmlFor="backtest-cadence">Entry cadence</label>
<Dropdown
id="backtest-cadence"
className="w-44 normal-case tracking-normal"
value={cadence}
onChange={(v) => setCadence(v as BacktestCadence)}
options={CADENCE_OPTIONS}
/>
</div>
<Button onClick={() => run.mutate()} loading={run.isPending} className="shrink-0">
{run.isPending ? 'Starting…' : report ? 'Re-run' : 'Run backtest'}
</Button>
</div>
</div>
<div>
<Disclosure summary="How this is measured">
<p className="max-w-2xl text-xs text-gray-400">
The backtest replays the current config at the selected cadence at each point the setup is
@@ -187,113 +144,29 @@ export function BacktestPanel() {
fundamentals are held neutral (no point-in-time history). ~6 months is roughly one market regime,
so read it as directional.
</p>
<p className="mt-2 max-w-2xl text-xs text-gray-400">
<strong className="text-gray-300">Live GTL</strong> is the exact target path the scanner and the
scheduled backtest use; <strong className="text-gray-300">Structural S/R</strong> is a comparison
arm sourcing targets from chart structure. <strong className="text-gray-300">Weekly</strong> steps
five sessions at a time and is what the server runs; <strong className="text-gray-300">Daily</strong>
{' '}is roughly 5× the replay work.
</p>
</Disclosure>
<div className="flex w-full flex-col gap-3 sm:w-auto sm:items-end">
<fieldset className="grid w-full grid-cols-1 gap-2 sm:w-[34rem] sm:grid-cols-2">
<legend className="mb-1 text-[11px] font-medium uppercase tracking-wider text-gray-500">
Target model for this run
</legend>
<label
className={`cursor-pointer rounded-lg border px-3 py-2 transition-colors focus-within:ring-2 focus-within:ring-blue-400/60 ${
targetModel === 'production_gtl'
? 'border-blue-400/60 bg-blue-500/10'
: 'border-white/10 bg-white/[0.03] hover:border-white/20'
}`}
>
<input
className="sr-only"
type="radio"
name="backtest-target-model"
value="production_gtl"
checked={targetModel === 'production_gtl'}
onChange={() => setTargetModel('production_gtl')}
/>
<span className="flex items-center justify-between gap-2 text-sm font-medium text-gray-100">
Live GTL
<span className="rounded-full border border-blue-400/40 bg-blue-400/10 px-2 py-0.5 text-[9px] font-semibold uppercase tracking-widest text-blue-300">
Production
</span>
</span>
<span className="mt-1 block text-[11px] leading-4 text-gray-500">
Exact target path used by the live scanner and scheduled backtest.
</span>
</label>
<label
className={`cursor-pointer rounded-lg border px-3 py-2 transition-colors focus-within:ring-2 focus-within:ring-amber-400/60 ${
targetModel === 'structural_sr'
? 'border-amber-400/50 bg-amber-500/10'
: 'border-white/10 bg-white/[0.03] hover:border-white/20'
}`}
>
<input
className="sr-only"
type="radio"
name="backtest-target-model"
value="structural_sr"
checked={targetModel === 'structural_sr'}
onChange={() => setTargetModel('structural_sr')}
/>
<span className="text-sm font-medium text-gray-200">Structural S/R</span>
<span className="mt-1 block text-[11px] leading-4 text-gray-500">
Comparison only; uses chart structure as the target source.
</span>
</label>
</fieldset>
<fieldset className="grid w-full grid-cols-2 gap-2 sm:w-[34rem]">
<legend className="mb-1 text-[11px] font-medium uppercase tracking-wider text-gray-500">
Entry cadence
</legend>
<label
className={`cursor-pointer rounded-lg border px-3 py-2 transition-colors focus-within:ring-2 focus-within:ring-blue-400/60 ${
cadence === 'weekly'
? 'border-blue-400/60 bg-blue-500/10'
: 'border-white/10 bg-white/[0.03] hover:border-white/20'
}`}
>
<input
className="sr-only"
type="radio"
name="backtest-cadence"
value="weekly"
checked={cadence === 'weekly'}
onChange={() => setCadence('weekly')}
/>
<span className="flex items-center justify-between gap-2 text-sm font-medium text-gray-100">
Weekly
<span className="rounded-full border border-blue-400/40 bg-blue-400/10 px-2 py-0.5 text-[9px] font-semibold uppercase tracking-widest text-blue-300">
Default
</span>
</span>
<span className="mt-1 block text-[11px] leading-4 text-gray-500">
Resource-safe server run at five-session intervals.
</span>
</label>
<label
className={`cursor-pointer rounded-lg border px-3 py-2 transition-colors focus-within:ring-2 focus-within:ring-amber-400/60 ${
cadence === 'daily'
? 'border-amber-400/50 bg-amber-500/10'
: 'border-white/10 bg-white/[0.03] hover:border-white/20'
}`}
>
<input
className="sr-only"
type="radio"
name="backtest-cadence"
value="daily"
checked={cadence === 'daily'}
onChange={() => setCadence('daily')}
/>
<span className="text-sm font-medium text-gray-200">Daily</span>
<span className="mt-1 block text-[11px] leading-4 text-amber-300/80">
Research run: roughly 5× the replay work; prefer the offline snapshot runner.
</span>
</label>
</fieldset>
<Button onClick={() => run.mutate()} loading={run.isPending} className="shrink-0">
{run.isPending ? 'Starting…' : report ? 'Re-run backtest' : 'Run backtest'}
</Button>
</div>
{/* Only surfaced for non-default choices zero noise on the common path,
but a non-production selection still announces itself, which is what
the old always-amber cards were really for. */}
{(cadence === 'daily' || targetModel === 'structural_sr') && (
<div className="space-y-1 text-[11px] text-amber-300/80">
{cadence === 'daily' && (
<p>Daily replays ~5× the work prefer the offline snapshot runner.</p>
)}
{targetModel === 'structural_sr' && (
<p>Comparison arm not the live scanner's target path.</p>
)}
</div>
)}
{isLoading && <Callout variant="empty">Loading</Callout>}
@@ -306,131 +179,21 @@ export function BacktestPanel() {
{report && (
<>
<p className="text-[11px] text-gray-500">
Ran {timeAgo(report.generated_at)} · {report.tickers} tickers · {report.candidates} setups
({report.qualified} qualified) · {report.params.entry_cadence ?? 'weekly'} cadence,
{' '}{report.params.horizon_days}-day horizon
{report.params.cost_per_side_pct != null && (
<> · net of {report.params.cost_per_side_pct}%/side costs</>
)}
{' '}· target model:{' '}
<span className={report.params.is_production_target_model === false ? 'text-amber-300' : 'text-blue-300'}>
{report.params.target_model_label ?? 'Unknown (legacy report)'}
</span>
</p>
{monitor && monitorRun ? (
<div className="space-y-3">
<div className="flex flex-wrap items-end justify-between gap-3">
<div>
<p className="section-index">Portfolio monitor</p>
<p className="mt-1 text-xs text-gray-500">
Simulated book for the selected strategy and lookback, compared with the S&P 500.
</p>
</div>
<div className="flex flex-wrap gap-2">
<div className="flex flex-col gap-1 text-[11px] uppercase tracking-wider text-gray-500">
<label htmlFor="monitor-strategy">Strategy</label>
<Dropdown
id="monitor-strategy"
className="w-64 normal-case tracking-normal"
value={activeStrategy}
onChange={setSelectedStrategy}
options={monitor.strategies.map((s) => ({
value: s.strategy,
label: `${s.is_production ? 'Production: ' : ''}${s.label}`,
}))}
<PortfolioMonitorPanel
monitor={monitor}
monitorRun={monitorRun}
activeStrategy={activeStrategy}
activeLookback={activeLookback}
onStrategyChange={setSelectedStrategy}
onLookbackChange={setSelectedLookback}
basisLookback={basisLookback}
basisLookbackLabel={report.recommendation?.basis_lookback_label ?? null}
productionStrategy={monitor?.production_strategy ?? null}
/>
</div>
<div className="flex flex-col gap-1 text-[11px] uppercase tracking-wider text-gray-500">
<label htmlFor="monitor-lookback">Lookback</label>
<Dropdown
id="monitor-lookback"
className="w-36 normal-case tracking-normal"
value={activeLookback}
onChange={setSelectedLookback}
options={monitor.lookbacks.map((l) => ({ value: l.lookback, label: l.label }))}
/>
</div>
</div>
</div>
<div className="grid gap-3 sm:grid-cols-2 lg:grid-cols-5">
<Stat label="CAGR" value={fmtSignedPct(monitorRun.cagr_pct)} valueClass={rColor(monitorRun.cagr_pct)} />
<Stat label="Sharpe" value={monitorRun.sharpe == null ? '—' : monitorRun.sharpe.toFixed(2)} />
<Stat label="Max Drawdown" value={fmtDrawdown(monitorRun.max_drawdown_pct)} valueClass="text-amber-400" />
<Stat
label="Total Return"
value={fmtSignedPct(monitorRun.total_return_pct)}
valueClass={rColor(monitorRun.total_return_pct)}
sub={`vs S&P 500 ${fmtSignedPct(monitorRun.spy_return_pct)}`}
/>
<Stat label="Trades" value={String(monitorRun.trades)} sub={`${fmtPct(monitorRun.win_rate)} win rate`} />
</div>
<EquityCurveChart run={monitorRun} />
<p className="text-[11px] text-gray-500">
Avg hold {fmtDays(monitorRun.avg_hold_days)} · Best {fmtR(monitorRun.best_trade_r)} / Worst{' '}
{fmtR(monitorRun.worst_trade_r)} · Avg P&amp;L per trade {fmtMoney(monitorRun.avg_trade_pnl)}
{monitorRun.reentry_policy === 'gate_reset' ? (
<> · Re-entry after gate failure and fresh qualification</>
) : null}
</p>
{monitorRun.yearly_returns && monitorRun.yearly_returns.length > 0 && (
<div className="glass overflow-x-auto p-4">
<p className="section-index mb-2">Per-year returns</p>
<div className="flex flex-wrap gap-2">
{monitorRun.yearly_returns.map((y) => (
<div key={y.year} className="rounded border border-white/10 px-3 py-1.5">
<span className="num text-xs text-gray-500">{y.year}</span>{' '}
<span className={`num text-sm font-semibold ${rColor(y.return_pct)}`}>
{fmtSignedPct(y.return_pct)}
</span>
</div>
))}
</div>
</div>
{report.recommendation && (
<BacktestRecommendationCard recommendation={report.recommendation} />
)}
{monitor.note && <p className="text-[11px] text-gray-600">{monitor.note}</p>}
</div>
) : (
<Callout variant="empty">
This report predates the portfolio monitor re-run the backtest to populate it.
</Callout>
)}
{report.recommendation && report.recommendation.items.length > 0 && (
<div className="glass border border-blue-400/20 p-4">
<p className="section-index">What this backtest recommends</p>
{report.recommendation.headline && (
<p className="mt-1.5 text-sm font-semibold text-gray-100">
{report.recommendation.headline}
</p>
)}
<ul className="mt-2 space-y-1">
{report.recommendation.items.map((item) => (
<li
key={item.topic + item.text}
className={`text-xs ${item.text.includes('WARNING') || item.text.includes('LAGS') ? 'text-amber-400' : 'text-gray-400'}`}
>
{item.text}
</li>
))}
</ul>
{report.recommendation.note && (
<p className="mt-2 text-[11px] text-gray-600">{report.recommendation.note}</p>
)}
</div>
)}
<p className="text-[11px] text-gray-600">
Strategy research gate tuning, exit sweeps, factor rank-IC now runs locally against a
database snapshot (see README). This page keeps only what says whether the promoted strategy
is worth trading; your realized results up top show what it is actually delivering.
</p>
</>
)}
</div>
@@ -0,0 +1,139 @@
import { Disclosure } from '../ui/Disclosure';
import type { BacktestRecommendation } from '../../lib/types';
/**
* The verdict, ahead of the tuning detail.
*
* Two problems this solves. All eight findings used to render as equal-weight
* bullets, so "does this strategy work" sat in the same register as "which
* cutoff scored best". And the headline which is a *description of the
* config*, not a verdict was the loudest thing on the card while every actual
* finding was small grey text.
*
* So: findings first, each split into a label and its detail; the config
* description demoted to a footer where it belongs.
*/
const PRIMARY_TOPICS = new Set(['production', 'benchmark', 'robustness']);
/**
* Mirrors how the backend phrases a bad result `_build_recommendation` emits
* "Robustness WARNING: …" and "Book vs SPY: LAGS …". There is deliberately no
* `severity` field on the payload; if that changes, this is the one place to fix.
*/
function isWarning(text: string): boolean {
return text.includes('WARNING') || text.includes('LAGS');
}
/**
* Every backend string self-prefixes ("Gate: keep the R:R floor…"), so the
* prefix IS the label no need for a chip that would just repeat it, and no
* need to reword anything server-side. Split on the first colon; if a string
* ever stops carrying one, it renders whole as detail.
*/
function splitLabel(text: string): { label: string | null; detail: string } {
const at = text.indexOf(': ');
if (at === -1 || at > 48) return { label: null, detail: text };
return { label: text.slice(0, at), detail: text.slice(at + 2) };
}
function Finding({ text, primary }: { text: string; primary: boolean }) {
const warn = isWarning(text);
const { label, detail } = splitLabel(text);
return (
<li className="flex flex-col gap-0.5 sm:flex-row sm:gap-3">
{label && (
<span
className={`shrink-0 text-[11px] font-semibold uppercase tracking-wider sm:w-44 sm:pt-0.5 ${
warn ? 'text-amber-400' : 'text-gray-500'
}`}
>
{label}
</span>
)}
<span
className={`${primary ? 'text-sm' : 'text-xs'} ${
warn ? 'text-amber-300' : primary ? 'text-gray-200' : 'text-gray-400'
}`}
>
{detail}
</span>
</li>
);
}
export function BacktestRecommendationCard({
recommendation,
}: {
recommendation: BacktestRecommendation;
}) {
const items = recommendation.items;
if (items.length === 0) return null;
// A warning is always visible, whatever its topic — burying "the edge
// disappears without the top 5% of winners" behind a disclosure would defeat
// the point of surfacing it at all.
const primary = items.filter((i) => PRIMARY_TOPICS.has(i.topic) || isWarning(i.text));
const secondary = items.filter((i) => !PRIMARY_TOPICS.has(i.topic) && !isWarning(i.text));
const warningCount = items.filter((i) => isWarning(i.text)).length;
return (
<div className="space-y-2">
<div className="glass border border-blue-400/20 p-4">
<div className="flex flex-wrap items-center justify-between gap-2">
<p className="section-index">What this backtest recommends</p>
{/* No headline means the backend found no production monitor row, so
nothing here describes the production book. Zero keyword warnings
is then absence of data, not a clean bill of health a green chip
beside "this report predates the portfolio monitor" would be a
success badge for missing data. */}
{!recommendation.headline ? (
<span className="rounded-full border border-white/15 bg-white/[0.05] px-2 py-0.5 text-[10px] font-semibold uppercase tracking-wider text-gray-400">
baseline unavailable
</span>
) : warningCount > 0 ? (
<span className="rounded-full border border-amber-400/40 bg-amber-400/10 px-2 py-0.5 text-[10px] font-semibold uppercase tracking-wider text-amber-300">
{warningCount} warning{warningCount > 1 ? 's' : ''}
</span>
) : (
<span className="rounded-full border border-emerald-400/30 bg-emerald-400/10 px-2 py-0.5 text-[10px] font-semibold uppercase tracking-wider text-emerald-300">
no warnings
</span>
)}
</div>
{primary.length > 0 && (
<ul className="mt-3 space-y-2.5">
{primary.map((item) => (
<Finding key={item.topic + item.text} text={item.text} primary />
))}
</ul>
)}
{/* The config description, demoted: it says what the strategy IS, which
is context for the findings above rather than a finding itself. */}
{recommendation.headline && (
<div className="mt-3 border-t border-white/[0.06] pt-3">
<p className="section-index">Configuration under test</p>
<p className="mt-1 text-xs leading-relaxed text-gray-500">{recommendation.headline}</p>
</div>
)}
{recommendation.note && (
<p className="mt-2 text-[11px] text-gray-600">{recommendation.note}</p>
)}
</div>
{/* Outside the card body on purpose: Disclosure renders its own glass-sm
panel, so nesting it inside the bordered card double-frames it. */}
{secondary.length > 0 && (
<Disclosure summary={`Gate and cutoff detail (${secondary.length})`}>
<ul className="space-y-2">
{secondary.map((item) => (
<Finding key={item.topic + item.text} text={item.text} primary={false} />
))}
</ul>
</Disclosure>
)}
</div>
);
}
@@ -0,0 +1,89 @@
import { Callout } from '../ui/Callout';
import { fmtSignedPct } from '../../lib/format';
import type { BacktestCurvePoint, BacktestPortfolioMonitorRun } from '../../lib/types';
/**
* Portfolio return vs S&P 500 for one monitor run.
*
* Hand-rolled SVG on purpose: two polylines and two axis rules do not justify a
* charting dependency, and the shape is fixed. Lives in `signals/` rather than
* `ui/` because it is typed to the backtest payload generalising it for a
* single caller would be the wrong trade.
*/
function curvePath(
points: BacktestCurvePoint[],
min: number,
max: number,
w: number,
h: number,
pad: number,
startMs: number,
endMs: number,
): string {
if (points.length < 2) return '';
const span = Math.max(max - min, 1);
const timeSpan = Math.max(endMs - startMs, 1);
return points
.map((p, i) => {
const t = new Date(p.date).getTime();
const x = pad + ((t - startMs) / timeSpan) * (w - pad * 2);
const value = p.return_pct ?? 0;
const y = pad + (1 - (value - min) / span) * (h - pad * 2);
return `${i === 0 ? 'M' : 'L'}${x.toFixed(1)},${y.toFixed(1)}`;
})
.join(' ');
}
export function EquityCurveChart({ run }: { run: BacktestPortfolioMonitorRun }) {
const portfolio = run.equity_curve ?? [];
const benchmark = run.benchmark_curve ?? [];
const values = [...portfolio, ...benchmark]
.map((p) => p.return_pct)
.filter((v): v is number => v !== null && v !== undefined);
if (portfolio.length < 2 || values.length === 0) {
return <Callout variant="empty">No equity curve points for this selection.</Callout>;
}
const min = Math.min(0, ...values);
const max = Math.max(0, ...values);
const times = [...portfolio, ...benchmark]
.map((p) => new Date(p.date).getTime())
.filter((v) => Number.isFinite(v));
if (times.length === 0) {
return <Callout variant="empty">No dated equity curve points for this selection.</Callout>;
}
const startMs = Math.min(...times);
const endMs = Math.max(...times);
const w = 720;
const h = 240;
const pad = 28;
const portfolioPath = curvePath(portfolio, min, max, w, h, pad, startMs, endMs);
const benchmarkPath = curvePath(benchmark, min, max, w, h, pad, startMs, endMs);
const lastPortfolio = portfolio[portfolio.length - 1]?.return_pct ?? null;
const lastBenchmark = benchmark[benchmark.length - 1]?.return_pct ?? run.spy_return_pct;
return (
<div className="glass overflow-hidden">
<div className="flex flex-wrap items-center justify-between gap-3 border-b border-white/[0.05] px-4 py-3">
<div>
<p className="text-sm font-semibold text-gray-100">{run.label}</p>
<p className="text-[11px] text-gray-500">{run.start_date} - {run.end_date}</p>
</div>
<div className="flex gap-4 text-xs">
<span className="text-blue-300">Portfolio {fmtSignedPct(lastPortfolio)}</span>
<span className="text-gray-400">S&P 500 {fmtSignedPct(lastBenchmark)}</span>
</div>
</div>
<svg viewBox={`0 0 ${w} ${h}`} className="h-64 w-full" role="img" aria-label="Portfolio return compared with S&P 500">
<line x1={pad} y1={h - pad} x2={w - pad} y2={h - pad} stroke="rgba(255,255,255,0.12)" />
<line x1={pad} y1={pad} x2={pad} y2={h - pad} stroke="rgba(255,255,255,0.12)" />
{benchmarkPath && (
<path d={benchmarkPath} fill="none" stroke="rgba(156,163,175,0.9)" strokeWidth="2" strokeDasharray="5 5" />
)}
<path d={portfolioPath} fill="none" stroke="rgb(96,165,250)" strokeWidth="3" />
<text x={pad} y={pad - 8} className="fill-gray-500 text-[10px]">{fmtSignedPct(max)}</text>
<text x={pad} y={h - 8} className="fill-gray-500 text-[10px]">{fmtSignedPct(min)}</text>
</svg>
</div>
);
}
@@ -5,8 +5,7 @@ import { triggerJob, resetTrackRecord } from '../../api/admin';
import { Button } from '../ui/Button';
import { Disclosure } from '../ui/Disclosure';
import { useToast } from '../ui/Toast';
import { BacktestPanel } from './BacktestPanel';
import { MyTradesPanel } from './MyTradesPanel';
import { fmtR, rColor } from '../../lib/format';
// Need at least this many matured setups before the pipeline check means anything;
// below it the live sample is too noisy to compare.
@@ -16,18 +15,6 @@ const DRIFT_TOLERANCE_R = 0.2;
type PipelineStatus = 'building' | 'tracking' | 'drift' | 'no-backtest';
function fmtR(value: number | null): string {
if (value === null) return '—';
return `${value > 0 ? '+' : ''}${value.toFixed(2)}R`;
}
function rColor(value: number | null): string {
if (value === null) return 'text-gray-400';
if (value > 0) return 'text-emerald-400';
if (value < 0) return 'text-red-400';
return 'text-gray-300';
}
function StatusChip({ status }: { status: PipelineStatus }) {
const styles: Record<PipelineStatus, { cls: string; label: string }> = {
tracking: { cls: 'border-emerald-500/30 bg-emerald-500/15 text-emerald-300', label: '✓ in sync' },
@@ -39,7 +26,7 @@ function StatusChip({ status }: { status: PipelineStatus }) {
return <span className={`shrink-0 rounded-full border px-2.5 py-1 text-xs font-medium ${s.cls}`}>{s.label}</span>;
}
export function TrackRecordPanel() {
export function EvaluationPanel() {
const queryClient = useQueryClient();
const toast = useToast();
@@ -101,19 +88,14 @@ export function TrackRecordPanel() {
return (
<div className="space-y-6">
{/* Your real, realized results come first; the strategy simulation follows. */}
<MyTradesPanel />
<div className="border-t border-white/[0.06]" />
<BacktestPanel />
<Disclosure summary="Track-record maintenance">
<Disclosure summary="Setup-grading diagnostic & maintenance">
<div className="space-y-4 pt-1">
<p className="max-w-2xl text-xs text-gray-500">
<span className="text-amber-300/90">Diagnostic only not production P&amp;L.</span>{' '}
Grades gate-level touch vs stop (the rejected take-profit model). Production exits are
initial stop / ATR trail / max hold see paper trades and the portfolio monitor above.
Target before stop = win, stop first = loss (same-bar both = loss), neither in 30 trading
days = expired at 0R. Only matured windows count. Scores{' '}
initial stop / ATR trail / max hold see the Paper Trades tab and the portfolio monitor
above. Target before stop = win, stop first = loss (same-bar both = loss), neither in 30
trading days = expired at 0R. Only matured windows count. Scores{' '}
<span className="text-gray-300">all</span> setups as a control group; runs nightly.
</p>
@@ -2,22 +2,10 @@ import { useMemo } from 'react';
import { Link } from 'react-router-dom';
import { usePaperTrades } from '../../hooks/usePaperTrades';
import { tradePnl } from '../../lib/paperTrade';
import { formatPrice } from '../../lib/format';
import { formatPrice, fmtR, fmtSignedMoney, rColor } from '../../lib/format';
import { Section } from '../ui/Section';
import { Callout } from '../ui/Callout';
function money(v: number): string {
return `${v >= 0 ? '+' : ''}$${Math.abs(v).toFixed(2)}`;
}
function fmtR(v: number | null): string {
return v === null ? '—' : `${v > 0 ? '+' : ''}${v.toFixed(2)}R`;
}
function color(v: number | null): string {
if (v === null) return 'text-gray-400';
if (v > 0) return 'text-emerald-400';
if (v < 0) return 'text-red-400';
return 'text-gray-300';
}
import { StatTile } from '../ui/StatTile';
// How the trade was closed — useful context on real trades at almost no cost.
function reasonMeta(reason: string | null): { label: string; cls: string } {
@@ -31,18 +19,6 @@ function reasonMeta(reason: string | null): { label: string; cls: string } {
}
}
function Stat({ label, value, valueClass = 'text-gray-100', sub }: {
label: string; value: string; valueClass?: string; sub?: string;
}) {
return (
<div className="glass p-4">
<p className="section-index">{label}</p>
<p className={`num mt-1.5 text-2xl font-semibold ${valueClass}`}>{value}</p>
{sub && <p className="mt-1 text-xs text-gray-500">{sub}</p>}
</div>
);
}
export function MyTradesPanel() {
const { data: closed, isLoading } = usePaperTrades('closed');
@@ -70,7 +46,10 @@ export function MyTradesPanel() {
if (isLoading) return null;
return (
<Section title="My Trades" hint="your realized paper-trading results">
<Section
title="Closed Trades"
hint="realized paper-trading results — open positions are on the Dashboard"
>
{stats.total === 0 ? (
<Callout variant="empty">
No closed trades yet. Take setups as paper trades and theyll resolve here when price hits
@@ -79,11 +58,11 @@ export function MyTradesPanel() {
) : (
<div className="space-y-4">
<div className="grid gap-3 sm:grid-cols-2 lg:grid-cols-5">
<Stat label="Hit Rate" value={stats.hitRate != null ? `${stats.hitRate.toFixed(1)}%` : '—'} sub={`${stats.wins}W / ${stats.losses}L`} />
<Stat label="Expectancy" value={fmtR(stats.avgR)} valueClass={color(stats.avgR)} sub="avg R per closed trade" />
<Stat label="Total R" value={fmtR(stats.totalR)} valueClass={color(stats.totalR)} sub={`${stats.total} closed`} />
<Stat label="Total P&L" value={money(stats.totalPnl)} valueClass={color(stats.totalPnl)} sub="realized, all closed" />
<Stat label="Alpha vs S&P 500" value={stats.totalAlpha != null ? money(stats.totalAlpha) : '—'} valueClass={color(stats.totalAlpha)} sub="realized vs buy-and-hold SPY" />
<StatTile label="Hit Rate" value={stats.hitRate != null ? `${stats.hitRate.toFixed(1)}%` : '—'} sub={`${stats.wins}W / ${stats.losses}L`} />
<StatTile label="Expectancy" value={fmtR(stats.avgR)} valueClass={rColor(stats.avgR)} sub="avg R per closed trade" />
<StatTile label="Total R" value={fmtR(stats.totalR)} valueClass={rColor(stats.totalR)} sub={`${stats.total} closed`} />
<StatTile label="Total P&L" value={fmtSignedMoney(stats.totalPnl)} valueClass={rColor(stats.totalPnl)} sub="realized, all closed" />
<StatTile label="Alpha vs S&P 500" value={stats.totalAlpha != null ? fmtSignedMoney(stats.totalAlpha) : '—'} valueClass={rColor(stats.totalAlpha)} sub="realized vs buy-and-hold SPY" />
</div>
<div className="glass overflow-x-auto">
@@ -112,9 +91,9 @@ export function MyTradesPanel() {
</td>
<td className="num px-4 py-2.5 text-right text-gray-300">{formatPrice(t.entry_price)}</td>
<td className="num px-4 py-2.5 text-right text-gray-300">{t.close_price != null ? formatPrice(t.close_price) : '—'}</td>
<td className={`num px-4 py-2.5 text-right font-semibold ${p ? color(p.pnl) : 'text-gray-500'}`}>{p ? money(p.pnl) : '—'}</td>
<td className={`num px-4 py-2.5 text-right ${p?.r != null ? color(p.r) : 'text-gray-500'}`}>{p?.r != null ? fmtR(p.r) : '—'}</td>
<td className={`num px-4 py-2.5 text-right ${t.alpha_pct != null ? color(t.alpha_pct) : 'text-gray-500'}`} title="Return vs. S&P 500 over the holding period">{t.alpha_pct != null ? `${t.alpha_pct >= 0 ? '+' : ''}${t.alpha_pct.toFixed(1)}%` : '—'}</td>
<td className={`num px-4 py-2.5 text-right font-semibold ${p ? rColor(p.pnl) : 'text-gray-500'}`}>{p ? fmtSignedMoney(p.pnl) : '—'}</td>
<td className={`num px-4 py-2.5 text-right ${p?.r != null ? rColor(p.r) : 'text-gray-500'}`}>{p?.r != null ? fmtR(p.r) : '—'}</td>
<td className={`num px-4 py-2.5 text-right ${t.alpha_pct != null ? rColor(t.alpha_pct) : 'text-gray-500'}`} title="Return vs. S&P 500 over the holding period">{t.alpha_pct != null ? `${t.alpha_pct >= 0 ? '+' : ''}${t.alpha_pct.toFixed(1)}%` : '—'}</td>
<td className="px-4 py-2.5">
<span className={`num text-[10px] font-semibold uppercase tracking-wider ${reasonMeta(t.close_reason).cls}`} title="How the trade was closed">
{reasonMeta(t.close_reason).label}
@@ -0,0 +1,219 @@
import { Callout } from '../ui/Callout';
import { Dropdown } from '../ui/Dropdown';
import { StatTile } from '../ui/StatTile';
import { EquityCurveChart } from './EquityCurveChart';
import {
fmtDays,
fmtDrawdown,
fmtPct,
fmtR,
fmtRatio,
fmtSignedMoney,
fmtSignedPct,
rColor,
} from '../../lib/format';
import type {
BacktestPortfolioMonitor,
BacktestPortfolioMonitorRun,
} from '../../lib/types';
/**
* The simulated book for one strategy/lookback selection, against the S&P 500.
*
* Selection state deliberately stays in BacktestPanel it also resolves which
* run this panel receives, so splitting it here would mean resolving twice.
*/
export function PortfolioMonitorPanel({
monitor,
monitorRun,
activeStrategy,
activeLookback,
onStrategyChange,
onLookbackChange,
basisLookback = null,
basisLookbackLabel = null,
productionStrategy = null,
}: {
monitor: BacktestPortfolioMonitor | null | undefined;
monitorRun: BacktestPortfolioMonitorRun | null | undefined;
activeStrategy: string;
activeLookback: string;
onStrategyChange: (v: string) => void;
onLookbackChange: (v: string) => void;
/** The window the recommendation below was computed on. */
basisLookback?: string | null;
basisLookbackLabel?: string | null;
productionStrategy?: string | null;
}) {
if (!monitor || !monitorRun) {
return (
<Callout variant="empty">
This report predates the portfolio monitor re-run the backtest to populate it.
</Callout>
);
}
// Key ABSENT (not null) means the cached report predates these metrics.
// Gated on sortino specifically: calmar and avg_trade_pnl have always been
// emitted, so testing those would half-populate the row with dashes.
const isLegacyRun = monitorRun.sortino === undefined;
return (
<div className="space-y-3">
<div className="flex flex-wrap items-end justify-between gap-3">
<div>
<p className="section-index">Portfolio monitor</p>
<p className="mt-1 text-xs text-gray-500">
Simulated book for the selected strategy and lookback, compared with the S&P 500.
</p>
</div>
<div className="flex flex-wrap gap-2">
<div className="flex flex-col gap-1 text-[11px] uppercase tracking-wider text-gray-500">
<label htmlFor="monitor-strategy">Strategy</label>
<Dropdown
id="monitor-strategy"
className="w-64 normal-case tracking-normal"
value={activeStrategy}
onChange={onStrategyChange}
options={monitor.strategies.map((s) => ({
value: s.strategy,
// "Production: " prefix dropped — a bullet costs one character
// instead of twelve, and the full config is spelled out under
// the chart anyway.
label: `${s.is_production ? '● ' : ''}${s.label}`,
}))}
/>
</div>
<div className="flex flex-col gap-1 text-[11px] uppercase tracking-wider text-gray-500">
<label htmlFor="monitor-lookback">Lookback</label>
<Dropdown
id="monitor-lookback"
className="w-36 normal-case tracking-normal"
value={activeLookback}
onChange={onLookbackChange}
options={monitor.lookbacks.map((l) => ({ value: l.lookback, label: l.label }))}
/>
</div>
</div>
</div>
{/* The recommendation below is baked into the report and cannot follow a
dropdown. On load the two agree by construction; say so plainly the
moment a selection moves off that basis. */}
{((basisLookback && activeLookback !== basisLookback) ||
(productionStrategy && activeStrategy !== productionStrategy)) && (
<p className="text-[11px] text-amber-300/80">
Showing{' '}
{productionStrategy && activeStrategy !== productionStrategy
? 'a comparison strategy'
: 'a different window'}
. The recommendation below is computed on the production strategy over{' '}
{basisLookbackLabel ?? basisLookback} these tiles will not match it.
</p>
)}
{/* Tier 1 — what the book returned. */}
<div className="grid gap-3 sm:grid-cols-2 lg:grid-cols-5">
<StatTile
label="Total Return"
value={fmtSignedPct(monitorRun.total_return_pct)}
valueClass={rColor(monitorRun.total_return_pct)}
sub={`vs S&P 500 ${fmtSignedPct(monitorRun.spy_return_pct)}`}
/>
<StatTile label="CAGR" value={fmtSignedPct(monitorRun.cagr_pct)} valueClass={rColor(monitorRun.cagr_pct)} />
<StatTile label="Max Drawdown" value={fmtDrawdown(monitorRun.max_drawdown_pct)} valueClass="text-amber-400" />
<StatTile
label="EV / trade"
value={fmtSignedMoney(monitorRun.avg_trade_pnl)}
valueClass={rColor(monitorRun.avg_trade_pnl)}
title="Average realized P&L per closed trade. Scales with position size, so it carries no quality band."
/>
<StatTile label="Trades" value={String(monitorRun.trades)} sub={`${fmtPct(monitorRun.win_rate)} win rate`} />
</div>
{/* Tier 2 how good that return was. Smaller and labelled on purpose:
ten equal tiles would read as ten equally important facts. */}
{isLegacyRun ? (
<p className="text-[11px] text-gray-600">
Risk-adjusted quality metrics appear after the next backtest run.
</p>
) : (
<div className="space-y-2">
<div className="flex flex-wrap items-baseline justify-between gap-2">
<p className="section-index">Risk-adjusted quality</p>
<p className="text-[11px] text-gray-600">
Bands are set stricter than textbook ranges this universe is today's
survivors replayed backward, which flatters every ratio.
</p>
</div>
<div className="grid gap-3 sm:grid-cols-2 lg:grid-cols-5">
<StatTile
label="Sharpe"
value={fmtRatio(monitorRun.sharpe)}
metric="sharpe"
raw={monitorRun.sharpe}
title="Return per unit of total volatility (annualized). Penalizes upside swings as well as downside."
/>
<StatTile
label="Sortino"
value={fmtRatio(monitorRun.sortino)}
metric="sortino"
raw={monitorRun.sortino}
title="Return per unit of downside deviation (annualized). Punishes losing days only, unlike Sharpe."
/>
<StatTile
label="Calmar (MAR)"
value={fmtRatio(monitorRun.calmar)}
metric="calmar"
raw={monitorRun.calmar}
title="CAGR divided by maximum drawdown — return earned per unit of worst-case pain."
/>
<StatTile
label="Gain / Pain"
value={fmtRatio(monitorRun.gain_to_pain)}
metric="gain_to_pain"
raw={monitorRun.gain_to_pain}
title="Sum of monthly returns divided by the absolute sum of the negative ones (Schwager)."
/>
<StatTile
label="Profit Factor ($)"
value={fmtRatio(monitorRun.profit_factor)}
metric="profit_factor"
raw={monitorRun.profit_factor}
title="Gross winning dollars divided by gross losing dollars, across closed trades."
/>
</div>
</div>
)}
<EquityCurveChart run={monitorRun} />
{/* avg_trade_pnl is a tile now (EV / trade) — not repeated here. */}
<p className="text-[11px] text-gray-500">
Avg hold {fmtDays(monitorRun.avg_hold_days)} · Best {fmtR(monitorRun.best_trade_r)} / Worst{' '}
{fmtR(monitorRun.worst_trade_r)}
{monitorRun.reentry_policy === 'gate_reset' ? (
<> · Re-entry after gate failure and fresh qualification</>
) : null}
</p>
{monitorRun.yearly_returns && monitorRun.yearly_returns.length > 0 && (
<div className="glass overflow-x-auto p-4">
<p className="section-index mb-2">Per-year returns</p>
<div className="flex flex-wrap gap-2">
{monitorRun.yearly_returns.map((y) => (
<div key={y.year} className="rounded border border-white/10 px-3 py-1.5">
<span className="num text-xs text-gray-500">{y.year}</span>{' '}
<span className={`num text-sm font-semibold ${rColor(y.return_pct)}`}>
{fmtSignedPct(y.return_pct)}
</span>
</div>
))}
</div>
</div>
)}
{monitor.note && <p className="text-[11px] text-gray-600">{monitor.note}</p>}
</div>
);
}
+1 -1
View File
@@ -16,7 +16,7 @@ const sizeClasses: Record<Size, string> = {
md: 'px-4 py-2 text-sm',
};
export function Spinner({ className = 'h-4 w-4' }: { className?: string }) {
function Spinner({ className = 'h-4 w-4' }: { className?: string }) {
return (
<svg className={`animate-spin ${className}`} viewBox="0 0 24 24" fill="none" aria-hidden="true">
<circle className="opacity-25" cx="12" cy="12" r="10" stroke="currentColor" strokeWidth="4" />
+1 -1
View File
@@ -9,7 +9,7 @@ interface DisclosureProps {
export function Disclosure({ summary, children }: DisclosureProps) {
return (
<details className="glass-sm group">
<summary className="flex cursor-pointer select-none items-center gap-2 px-4 py-2.5 text-xs font-medium text-gray-400 transition-colors hover:text-gray-200 [&::-webkit-details-marker]:hidden">
<summary className="flex min-h-11 cursor-pointer select-none items-center gap-2 px-4 py-2.5 text-xs font-medium text-gray-400 transition-colors hover:text-gray-200 [&::-webkit-details-marker]:hidden">
<span className="inline-block transition-transform duration-200 group-open:rotate-90"></span>
{summary}
</summary>
+6 -1
View File
@@ -86,7 +86,12 @@ export function Dropdown({
onClick={() => setOpen((v) => !v)}
className="input-glass flex w-full items-center justify-between gap-2 px-3 py-1.5 text-left text-sm"
>
<span className={selected ? 'text-gray-200' : 'text-gray-500'}>
{/* truncate, not wrap: a long option name used to push the trigger to
three lines and shove the whole control row out of alignment. */}
<span
className={`truncate ${selected ? 'text-gray-200' : 'text-gray-500'}`}
title={selected ? selected.label : undefined}
>
{selected ? selected.label : placeholder}
</span>
<svg
-4
View File
@@ -1,9 +1,5 @@
const pulse = 'animate-pulse rounded-lg bg-white/[0.05]';
export function SkeletonLine({ className = '' }: { className?: string }) {
return <div className={`${pulse} h-4 w-full ${className}`} />;
}
export function SkeletonCard({ className = '' }: { className?: string }) {
return <div className={`${pulse} h-32 w-full ${className}`} />;
}
+71
View File
@@ -0,0 +1,71 @@
import {
BAND_STYLE,
bandTicks,
classifyMetric,
meterFraction,
} from '../../lib/metricBands';
/**
* One labelled metric.
*
* Optionally carries a quality meter: pass `metric` (a key in METRIC_BANDS) and
* the numeric `raw` value. The meter is the answer to "2.72 — is that good?"
* a track showing where the value sits, ticks at the band edges, and the band
* word. Colour never travels alone; the word is always rendered beside it.
*
* Every tile is the same size. Hierarchy comes from grouping and section
* labels, not from shrinking one row two sizes read as inconsistent rather
* than as a deliberate ranking.
*/
export function StatTile({
label,
value,
valueClass = 'text-gray-100',
sub,
title,
metric,
raw,
}: {
label: string;
value: string;
valueClass?: string;
sub?: string;
/** Native tooltip — how the metric is defined. */
title?: string;
/** Key into METRIC_BANDS; enables the quality meter. */
metric?: string;
/** Numeric value the meter reads (the formatted `value` is display-only). */
raw?: number | null;
}) {
const band = metric ? classifyMetric(metric, raw) : null;
const style = band ? BAND_STYLE[band] : null;
return (
<div className="glass flex flex-col p-4" title={title}>
<p className="section-index">{label}</p>
<p className={`num mt-1.5 text-2xl font-semibold ${valueClass}`}>{value}</p>
{style && metric && (
<div className="mt-2.5">
<div className="relative h-1.5 overflow-hidden rounded-full bg-white/[0.07]">
<div
className={`h-full rounded-full ${style.fill}`}
style={{ width: `${meterFraction(metric, raw) * 100}%` }}
/>
{/* Band edges — where "fair" becomes "good", and so on. */}
{bandTicks(metric).map((t) => (
<span
key={t}
className="absolute top-0 h-full w-px bg-black/50"
style={{ left: `${t * 100}%` }}
/>
))}
</div>
<p className={`mt-1.5 text-[11px] font-medium ${style.text}`}>{style.label}</p>
</div>
)}
{sub && <p className="mt-1 text-xs text-gray-500">{sub}</p>}
</div>
);
}
-146
View File
@@ -1,146 +0,0 @@
/* Dev-only visual harness for FundamentalsPanel. Served at /harness.html by
* `vite`. Not imported by the app. Renders the three key states so desktop and
* mobile can be eyeballed with representative fixtures. */
import { createRoot } from 'react-dom/client';
import '../styles/globals.css';
import { FundamentalsPanel } from '../components/ticker/FundamentalsPanel';
import type { FundamentalResponse, MetricItem } from '../lib/types';
function h(period: string, value: number | null) {
return { period_end: period, value };
}
const P = ['2025-06-30', '2025-09-30', '2025-12-31', '2026-03-28'];
function dateFromToday(days: number): string {
const date = new Date();
date.setHours(12, 0, 0, 0);
date.setDate(date.getDate() + days);
return [
date.getFullYear(),
String(date.getMonth() + 1).padStart(2, '0'),
String(date.getDate()).padStart(2, '0'),
].join('-');
}
function metric(key: string, value: number | null, hist: (number | null)[],
industry: MetricItem['industry'] = null,
caveat: string | null = null): MetricItem {
return {
key: key as MetricItem['key'], value,
history: hist.map((v, i) => h(P[i], v)),
industry, period_end: '2026-03-28', filed_date: '2026-05-01', caveat,
source: 'sec',
};
}
const ind = (median: number, favorable_percentile: number) =>
({ label: 'SIC 35 peers', median, favorable_percentile, peer_count: 12 });
const legacy = {
pe_ratio: null, revenue_growth: null, earnings_surprise: null, market_cap: null,
next_earnings_date: null, fetched_at: null, unavailable_fields: {},
setup_eligible: true, setup_block_code: null, setup_block_reason: null,
};
const full: FundamentalResponse = {
symbol: 'AAPL', ...legacy,
earnings: {
next: { date: dateFromToday(12), session: 'amc', days_until: 12 },
recent: [
{ announce_date: '2025-08-01', period_end: '2025-06-30', eps_estimate: 1.4, eps_actual: 1.6, surprise_pct: 14.3 },
{ announce_date: '2025-11-01', period_end: '2025-09-30', eps_estimate: 1.7, eps_actual: 1.9, surprise_pct: 11.8 },
{ announce_date: '2026-02-01', period_end: '2025-12-31', eps_estimate: 2.6, eps_actual: 2.4, surprise_pct: -7.7 },
{ announce_date: '2026-05-01', period_end: '2026-03-28', eps_estimate: 1.5, eps_actual: 1.65, surprise_pct: 10.0 },
],
},
metrics: [
metric('revenue_growth_yoy', 18, [8, 11, 15, 18], ind(11, 82)),
metric('eps_growth_yoy', 24, [10, 18, 22, 24], ind(15, 70)),
metric('operating_margin', 32, [30, 31, 31, 32], ind(22, 88)),
metric('fcf_margin', 28, [24, 25, 27, 28], ind(18, 80)),
metric('net_debt', 16.2e9, [46e9, 44e9, 24e9, 16.2e9], null),
metric('net_debt_to_ebitda', 1.4, [1.9, 1.7, 1.5, 1.4], ind(2.1, 68)),
metric('share_count_change_yoy', -1.7, [-2.4, -2.2, -2.3, -1.7], null),
],
valuation: {
pe: 29.2, fcf_yield: 3.8, market_cap_est: 3.2e12,
pe_industry: ind(23.5, 30), fcf_yield_industry: ind(3.1, 70), price_date: '2026-05-01',
},
reads: {
header: 'growth accelerating · margins improving · valuation priced above peers',
by_key: {
revenue_growth_yoy: 'accelerating', eps_growth_yoy: 'accelerating',
operating_margin: 'improving', fcf_margin: 'improving',
share_count_change_yoy: 'buying back', net_debt_to_ebitda: 'conservative leverage',
pe: 'priced above peers', fcf_yield: 'above peers', net_debt: null,
},
},
};
const partial: FundamentalResponse = {
symbol: 'NEWCO', ...legacy,
earnings: { next: { date: dateFromToday(0), session: 'unknown', days_until: 0 }, recent: [] },
metrics: [
metric('revenue_growth_yoy', 12, [null, 8, 10, 12], null),
metric(
'eps_growth_yoy',
null,
[null, null, null, null],
null,
'Not comparable: share count changed at least 25%; possible split or corporate action.',
),
metric('operating_margin', 25, [24, 24, 25, 25], null),
metric('fcf_margin', null, [null, null, null, null], null),
metric('net_debt', null, [], null),
metric('net_debt_to_ebitda', 1.9, [1.7, 1.8, 1.9, 1.9], null),
metric(
'share_count_change_yoy',
null,
[1.8, 2.0, 2.0, null],
null,
'Not comparable: share count changed at least 25%; possible split or corporate action.',
),
],
valuation: {
pe: 15.2, fcf_yield: null, market_cap_est: 5.4e8,
pe_industry: null, fcf_yield_industry: null, price_date: '2026-05-01',
},
reads: {
header: 'growth steady · margins stable',
by_key: {
revenue_growth_yoy: 'steady', operating_margin: 'stable',
share_count_change_yoy: '2.1% dilution', net_debt_to_ebitda: null,
pe: null, fcf_yield: null, eps_growth_yoy: null, fcf_margin: null, net_debt: null,
},
},
};
const empty: FundamentalResponse = {
symbol: 'ADR', ...legacy,
earnings: { next: null, recent: [] },
metrics: [
'revenue_growth_yoy', 'eps_growth_yoy', 'operating_margin', 'fcf_margin',
'net_debt', 'net_debt_to_ebitda', 'share_count_change_yoy',
].map((k) => metric(k, null, [])),
valuation: null,
reads: { header: null, by_key: {} },
};
function Case({ title, data }: { title: string; data: FundamentalResponse }) {
return (
<div>
<div className="mb-1.5 text-[11px] uppercase tracking-widest text-gray-500">{title}</div>
<FundamentalsPanel data={data} />
</div>
);
}
createRoot(document.getElementById('root')!).render(
<div className="mx-auto max-w-3xl space-y-8 p-6">
<p className="text-[11px] uppercase tracking-widest text-gray-500">
Desktop width (~768px, two columns). Resize the browser to ~390px to check mobile (single column).
</p>
<Case title="Full" data={full} />
<Case title="Partial · insufficient peers" data={partial} />
<Case title="Empty" data={empty} />
</div>,
);
+1 -1
View File
@@ -21,7 +21,7 @@ import type { ExitPolicy, TradeSetup } from './types';
* Guarded by test_prod_strategy_parity.py so a backend change can't silently
* desync this.
*/
export const SETUP_STOP_ATR_MULTIPLIER = 1.5;
const SETUP_STOP_ATR_MULTIPLIER = 1.5;
export interface ExitPlan {
mode: ExitPolicy['mode'];
+55
View File
@@ -72,3 +72,58 @@ export function formatDateTime(d: string): string {
hour12: true,
})}`;
}
// ── Metric display helpers ─────────────────────────────────────────────────
// Shared by the Signals backtest/paper-trade panels. Dashboard and
// OpenTradesPanel deliberately still carry their own copies — migrating them is
// a separate change, not drive-by scope.
/** R-multiple with an explicit sign. e.g. 1.2 → "+1.20R", null → "—" */
export function fmtR(v: number | null | undefined): string {
if (v === null || v === undefined) return '—';
return `${v > 0 ? '+' : ''}${v.toFixed(2)}R`;
}
/** e.g. 12.34 → "12.3%" */
export function fmtPct(v: number | null | undefined): string {
return v === null || v === undefined ? '—' : `${v.toFixed(1)}%`;
}
/** e.g. 12.34 → "+12.3%" */
export function fmtSignedPct(v: number | null | undefined): string {
if (v === null || v === undefined) return '—';
return `${v > 0 ? '+' : ''}${v.toFixed(1)}%`;
}
/** Always rendered negative, whatever sign the source uses. 17.3 → "-17.3%" */
export function fmtDrawdown(v: number | null | undefined): string {
return v === null || v === undefined ? '—' : `-${Math.abs(v).toFixed(1)}%`;
}
/** e.g. 15.3 → "15.3d" */
export function fmtDays(v: number | null | undefined): string {
return v === null || v === undefined ? '—' : `${v.toFixed(1)}d`;
}
/** Unitless ratios — Sharpe, Sortino, Calmar, Gain/Pain, profit factor. */
export function fmtRatio(v: number | null | undefined): string {
return v === null || v === undefined ? '—' : v.toFixed(2);
}
/**
* Signed currency, using U+2212 for negatives. e.g. -12.3 "$12.30"
* Use wherever a value can go negative and the unit is money.
* (For a bare unsigned amount there is already `formatPrice` above.)
*/
export function fmtSignedMoney(v: number | null | undefined): string {
if (v === null || v === undefined) return '—';
return `${v >= 0 ? '+' : ''}$${Math.abs(v).toFixed(2)}`;
}
/** Green above zero, red below, neutral at zero or null. */
export function rColor(v: number | null | undefined): string {
if (v === null || v === undefined) return 'text-gray-400';
if (v > 0) return 'text-emerald-400';
if (v < 0) return 'text-red-400';
return 'text-gray-300';
}
-112
View File
@@ -1,112 +0,0 @@
/**
* Fundamental dimension readouts for the Fundamentals tab.
*
* Scoring mirrors app/services/scoring_service.py _compute_fundamental_score:
* equal-weighted average of available P/E, revenue growth, and earnings
* surprise (need 2 metrics). Market cap is display-only, not scored.
*/
export interface FundamentalMetrics {
pe_ratio: number | null;
revenue_growth: number | null;
earnings_surprise: number | null;
market_cap?: number | null;
}
export interface StatusRead {
text: string;
tone: string;
}
const clamp = (v: number, lo = 0, hi = 100) => Math.max(lo, Math.min(hi, v));
/** P/E sub-score: lower is better. PE 15 → 100, 30 → 50, 45 → 0. */
export function peSubScore(pe: number): number {
return clamp(100 - (pe - 15) * (100 / 30));
}
/** Revenue growth sub-score: 0% → 50, +20% → 100, 20% → 0. */
export function revenueGrowthSubScore(growthPct: number): number {
return clamp(50 + growthPct * 2.5);
}
/** Earnings surprise sub-score: 0% → 50, +10% → 100, 10% → 0. */
export function earningsSurpriseSubScore(surprisePct: number): number {
return clamp(50 + surprisePct * 5);
}
/**
* Overall fundamental score, or null when fewer than 2 scored metrics.
* Matches backend MIN_METRICS = 2.
*/
export function fundamentalScore(m: FundamentalMetrics): number | null {
const parts: number[] = [];
if (m.pe_ratio != null && m.pe_ratio > 0) parts.push(peSubScore(m.pe_ratio));
if (m.revenue_growth != null) parts.push(revenueGrowthSubScore(m.revenue_growth));
if (m.earnings_surprise != null) parts.push(earningsSurpriseSubScore(m.earnings_surprise));
if (parts.length < 2) return null;
return parts.reduce((a, b) => a + b, 0) / parts.length;
}
export function overallFundamentalStatus(score: number | null): StatusRead {
if (score == null) {
return { text: 'incomplete data', tone: 'text-amber-300' };
}
if (score >= 70) return { text: 'strong fundamentals', tone: 'text-emerald-300' };
if (score >= 55) return { text: 'healthy', tone: 'text-emerald-300' };
if (score >= 45) return { text: 'mixed / average', tone: 'text-gray-400' };
if (score >= 30) return { text: 'soft', tone: 'text-amber-300' };
return { text: 'weak fundamentals', tone: 'text-red-300' };
}
export function peStatus(pe: number | null): StatusRead | null {
if (pe == null || !(pe > 0)) return null;
if (pe <= 15) return { text: 'cheap / attractive', tone: 'text-emerald-300' };
if (pe <= 25) return { text: 'fair', tone: 'text-gray-400' };
if (pe <= 35) return { text: 'expensive', tone: 'text-amber-300' };
return { text: 'rich', tone: 'text-red-300' };
}
export function revenueGrowthStatus(growthPct: number | null): StatusRead | null {
if (growthPct == null) return null;
if (growthPct >= 20) return { text: 'strong growth', tone: 'text-emerald-300' };
if (growthPct >= 5) return { text: 'solid growth', tone: 'text-emerald-300' };
if (growthPct >= -5) return { text: 'flat', tone: 'text-gray-400' };
if (growthPct >= -20) return { text: 'contracting', tone: 'text-amber-300' };
return { text: 'deep contraction', tone: 'text-red-300' };
}
export function earningsSurpriseStatus(surprisePct: number | null): StatusRead | null {
if (surprisePct == null) return null;
if (surprisePct >= 10) return { text: 'beat (large)', tone: 'text-emerald-300' };
if (surprisePct >= 2) return { text: 'beat', tone: 'text-emerald-300' };
if (surprisePct >= -2) return { text: 'in line', tone: 'text-gray-400' };
if (surprisePct >= -10) return { text: 'miss', tone: 'text-amber-300' };
return { text: 'miss (large)', tone: 'text-red-300' };
}
/** Size band only — not good/bad, not part of the score. */
export function marketCapStatus(marketCap: number | null): StatusRead | null {
if (marketCap == null || !(marketCap > 0)) return null;
if (marketCap >= 200e9) return { text: 'mega cap', tone: 'text-gray-400' };
if (marketCap >= 10e9) return { text: 'large cap', tone: 'text-gray-400' };
if (marketCap >= 2e9) return { text: 'mid cap', tone: 'text-gray-400' };
if (marketCap >= 300e6) return { text: 'small cap', tone: 'text-gray-400' };
return { text: 'micro cap', tone: 'text-gray-400' };
}
export function metricStatus(
key: 'pe_ratio' | 'revenue_growth' | 'earnings_surprise' | 'market_cap',
value: number | null,
): StatusRead | null {
switch (key) {
case 'pe_ratio':
return peStatus(value);
case 'revenue_growth':
return revenueGrowthStatus(value);
case 'earnings_surprise':
return earningsSurpriseStatus(value);
case 'market_cap':
return marketCapStatus(value);
}
}
+77
View File
@@ -0,0 +1,77 @@
/**
* Quality bands for the risk-adjusted metrics.
*
* A tile reading "Sortino 2.72" answers nothing on its own. These bands turn
* each ratio into weak / fair / good / strong so the tile says whether the
* number is any good.
*
* The bands are deliberately STRICTER than the textbook ranges. This backtest
* replays today's ~512 tracked tickers backward, so every name that failed or
* was acquired inside the window is missing and every ratio here is flattered.
* Standard thresholds would print "strong" on numbers survivorship inflated.
* Treat a band as a claim about this book relative to itself, not a claim that
* the live strategy will reproduce it.
*
* Edges are lower-inclusive: a value exactly on an edge takes the higher band.
*/
export type BandName = 'weak' | 'fair' | 'good' | 'strong';
export interface MetricBand {
/** Lower edges for fair / good / strong. Below the first edge is weak. */
edges: [number, number, number];
/** Where the meter track ends. Values above clamp to full. */
max: number;
}
export const METRIC_BANDS: Record<string, MetricBand> = {
sharpe: { edges: [0.8, 1.5, 2.5], max: 3.5 },
sortino: { edges: [1.2, 2.0, 3.0], max: 4.0 },
calmar: { edges: [0.5, 1.0, 2.5], max: 3.5 },
gain_to_pain: { edges: [1.0, 1.5, 2.5], max: 3.5 },
profit_factor: { edges: [1.3, 1.8, 2.5], max: 3.5 },
};
const BAND_ORDER: BandName[] = ['weak', 'fair', 'good', 'strong'];
export function classifyMetric(
key: keyof typeof METRIC_BANDS | string,
value: number | null | undefined,
): BandName | null {
const band = METRIC_BANDS[key];
if (!band || value === null || value === undefined || !Number.isFinite(value)) {
return null;
}
const passed = band.edges.filter((edge) => value >= edge).length;
return BAND_ORDER[passed];
}
/** Fraction of the meter track a value fills, clamped to 0..1. */
export function meterFraction(
key: keyof typeof METRIC_BANDS | string,
value: number | null | undefined,
): number {
const band = METRIC_BANDS[key];
if (!band || value === null || value === undefined || !Number.isFinite(value)) {
return 0;
}
return Math.max(0, Math.min(1, value / band.max));
}
/** Band edges as track fractions, for drawing the tick marks. */
export function bandTicks(key: keyof typeof METRIC_BANDS | string): number[] {
const band = METRIC_BANDS[key];
if (!band) return [];
return band.edges.map((edge) => edge / band.max);
}
/**
* Status colours, not the categorical palette these encode state, so they are
* reserved and always paired with the band word rather than standing alone.
*/
export const BAND_STYLE: Record<BandName, { fill: string; text: string; label: string }> = {
weak: { fill: 'bg-red-400/70', text: 'text-red-400', label: 'weak' },
fair: { fill: 'bg-amber-400/70', text: 'text-amber-400', label: 'fair' },
good: { fill: 'bg-emerald-400/70', text: 'text-emerald-400', label: 'good' },
strong: { fill: 'bg-emerald-300/80', text: 'text-emerald-300', label: 'strong' },
};
+64 -13
View File
@@ -1,17 +1,68 @@
import type { MarketRegime } from './types';
import type { FundamentalState, MarketRegime } from './types';
export function regimeColor(label: MarketRegime['label']): string {
switch (label) {
case 'bullish':
return 'text-emerald-400';
case 'bearish':
return 'text-red-400';
case 'neutral':
return 'text-amber-400';
default:
return 'text-gray-400';
}
}
/** One visual vocabulary for the three-channel regime monitor. Keep chart SVG
* literals and DOM text in sync rather than letting Tailwind aliases and
* hard-coded colours describe the same state differently.
*
* **Hue identifies the channel, and only the channel.** Two collisions made
* that false and both are fixed here:
*
* - `supportive` was literally `state`, so teal meant "the State score" in the
* Time view and "fundamentals supportive" in the Path view of the same card.
* - `adverse` sat 19 degrees from `warning`, which is inside deuteranope
* confusion range for two channels that appear on adjacent tooltip lines.
*
* The market pair now sits at 27/190 degrees and the fundamental pair at
* 0/158, so every *cross-channel* pair is at least 27 degrees apart. All six
* clear 4.5:1 against `--surface`. Fundamentals additionally carry a glyph, so
* colour is never the sole encoding for the categorical channel.
*
* `neutral` and `unknown` are deliberately the same hue: they are two states of
* one channel, both meaning "no directional signal", separated by lightness
* (7.1:1 vs 5.4:1) and by glyph (filled circle vs ring). Do not "fix" their
* proximity by giving `unknown` a hue that would make an absence of evidence
* look like a reading.
*
* These deliberately do *not* reuse `--up-text`/`--down-text`: those are the
* app's directional tokens, and `--up-text` is already this chart's State
* colour, which is how the first collision happened.
*/
export const REGIME_VISUAL = {
// Market channels — continuous scores, drawn as lines and positions.
state: '#6ec9db',
warning: '#fb923c',
// Fundamental channel — categorical, drawn as glyphs.
supportive: '#34d399',
neutral: '#9aa0b0',
adverse: '#f87171',
unknown: '#848a9c',
} as const;
/** Market quadrant severity as an opacity ramp on one neutral never a hue.
*
* The quadrants are a State x Warning construct, so colouring them borrowed
* hues that already meant something else: "Healthy" was painted in the
* fundamental supportive colour and "Stabilizing" in the adverse one, which put
* an adverse glyph on an adverse-coloured background while meaning roughly the
* opposite (damage receding). Opacity carries how many axes are elevated, the
* position and labels carry which, and hue stays free to mean channel.
*/
export const QUADRANT_WASH = {
healthy: 0.015,
early_warning: 0.05,
stabilizing: 0.05,
active_stress: 0.085,
} as const;
export const FUNDAMENTAL_VISUAL: Record<
FundamentalState,
{ label: string; color: string; glyph: 'up' | 'circle' | 'diamond' | 'ring' }
> = {
supportive: { label: 'Supportive', color: REGIME_VISUAL.supportive, glyph: 'up' },
neutral: { label: 'Neutral', color: REGIME_VISUAL.neutral, glyph: 'circle' },
adverse: { label: 'Adverse', color: REGIME_VISUAL.adverse, glyph: 'diamond' },
unknown: { label: 'Unknown', color: REGIME_VISUAL.unknown, glyph: 'ring' },
};
export function regimeDot(label: MarketRegime['label']): string {
switch (label) {
+137 -24
View File
@@ -196,6 +196,8 @@ export interface ScheduleConfig {
schedule_near_close_pipeline_cron: string;
schedule_after_close_pipeline_cron: string;
schedule_intraday_pipeline_cron: string;
schedule_backtest_cron: string;
schedule_ticker_universe_cron: string;
}
// Runtime sentiment LLM configuration
@@ -293,6 +295,20 @@ export interface BacktestPortfolioPolicy {
cagr_pct: number | null;
max_drawdown_pct: number;
sharpe: number | null;
sharpe_se?: number | null;
psr?: number | null;
/** CAGR / max drawdown — the same number commonly called MAR. */
calmar?: number | null;
/**
* Optional because reports cached before these landed lack the keys entirely.
* An ABSENT `sortino` is how the UI detects such a report distinct from
* `null`, which means "computed, undefined for this run".
*/
sortino?: number | null;
/** Schwager, on monthly returns. */
gain_to_pain?: number | null;
/** DOLLAR-based. Not the R-based profit_factor on BacktestBucket. */
profit_factor?: number | null;
trades: number;
win_rate: number | null;
avg_trade_pnl: number | null;
@@ -320,6 +336,13 @@ export interface BacktestCurvePoint {
export interface BacktestRecommendation {
headline: string | null;
items: { topic: string; text: string }[];
/**
* The monitor lookback every production/benchmark figure was read from. The
* page defaults its selector to this so the tiles and the recommendation
* cannot open on different windows. Absent on reports predating the field.
*/
basis_lookback?: string | null;
basis_lookback_label?: string | null;
note?: string;
}
@@ -486,9 +509,24 @@ export interface RegimeReading {
trend?: { delta_7: number | null; delta_30: number | null };
}
/** Qualitative capex / earnings-reaction context. Not part of either score. */
export interface RegimeFundamentalOverlay {
export type FundamentalState = 'supportive' | 'neutral' | 'adverse' | 'unknown';
export type EvidenceQuality = 'complete' | 'partial' | 'stale' | 'manual' | 'unavailable';
/** The third channel: capex / earnings-reaction context, read alongside State
* and Warning by confluence. Deliberately never a term in either score see
* the methodology doc on why no fusion weight is measurable yet. */
export interface RegimeFundamentalContext {
/** Derived from the stored facts by fixed rules, not by an LLM's judgement. */
state: FundamentalState;
evidence_quality: EvidenceQuality;
capex_signal: FundamentalState;
reaction_signal: FundamentalState;
/** Timing only: there is an effective, non-stale record to display. */
available: boolean;
/** Content too: it is available *and* actually determined something. A
* collected observation whose extraction failed is available but not usable,
* and only `usable` may confirm anything or count as study exposure. */
usable: boolean;
pending: boolean;
stale: boolean;
effective_date: string | null;
@@ -501,7 +539,7 @@ export interface RegimeFundamentalOverlay {
source: string | null;
fetched_at: string | null;
/** Whether anything was actually collected. Live reading only; the snapshot's
* point-in-time overlay omits it. */
* point-in-time record omits it. */
observed?: boolean;
observed_in_snapshot?: boolean;
}
@@ -510,6 +548,11 @@ export interface RegimeHistoryPoint {
date: string;
state: number | null;
warning: number | null;
/** The fundamental channel as recorded that day drives the Path dot colour.
* Rows written before the channel existed read as "unknown", which is correct:
* nothing was observed then either. */
fundamental_state: FundamentalState;
evidence_quality: EvidenceQuality;
state_coverage: number | null;
warning_coverage: number | null;
basket_hash: string | null;
@@ -522,10 +565,12 @@ export interface RegimeMonitor {
date?: string;
state?: RegimeReading;
warning?: RegimeReading;
/** Point-in-time overlay recorded in the snapshot. */
fundamental_overlay?: RegimeFundamentalOverlay;
/** Current observation, even when it is not effective until the next session. */
fundamental_context?: RegimeFundamentalOverlay;
/** The channel as recorded in the snapshot — point-in-time, effective-date gated. */
fundamental_context?: RegimeFundamentalContext;
/** What we know right now, even when it is not effective until the next
* session. Separate from the above so a just-collected observation cannot
* look as though it had been backdated into the record. */
fundamental_live?: RegimeFundamentalContext;
inputs?: {
vix: number | null;
vix_date: string | null;
@@ -560,7 +605,7 @@ export interface RegimeMonitor {
}
export interface RegimeFundamentals {
methodology: 'v3';
methodology: 'v4';
f1_score: number | null;
f3_score: number | null;
locked: boolean;
@@ -573,7 +618,7 @@ export interface RegimeFundamentals {
}
export type CapexState = 'raising' | 'holding' | 'cutting' | 'unknown';
export type GoodNewsReaction = 'yes' | 'no' | 'mixed';
export type GoodNewsReaction = 'yes' | 'no' | 'mixed' | 'unknown';
export interface RegimeFundamentalsUpdate {
capex?: Record<string, CapexState>;
@@ -588,9 +633,22 @@ export interface RegimeConfig {
}
// Event study — measured lead time of early-warning indicators vs. drawdowns
export interface EventStudyMetrics {
events: number;
events_warned: number;
events_missed: number;
alarm_episodes: number;
false_alarms: number;
/** null when the rule had no eligible sessions — undefined, not zero. */
false_alarms_per_year: number | null;
median_lead_days: number | null;
}
export interface EventStudyReport {
available: boolean;
reason?: string;
/** Report shape, independent of methodology. Mismatched reports are discarded. */
schema?: number;
methodology?: string;
generated_at?: string;
evaluation?: 'exploratory' | 'holdout';
@@ -601,14 +659,11 @@ export interface EventStudyReport {
event_threshold_pct: number;
event_cooldown_days: number;
horizon_days: number;
train_fraction: number;
warn_percentile: number;
warn_threshold: number;
basket_hash: string;
basket_asof: string;
credit_sensor_from?: string | null;
};
/** How far the headline metrics can be trusted. See _reliability(). */
/** How far the *fitted* variant's metrics can be trusted. See _reliability(). */
reliability?: {
events_detected: number;
events_in_holdout: number;
@@ -622,21 +677,75 @@ export interface EventStudyReport {
sample?: {
start: string;
end: string;
train_end: string;
test_start: string;
/** Where the quadrant baseline seeds — not a holdout boundary. */
evaluable_from: string;
sessions: number;
holdout_sessions: number;
evaluable_sessions: number;
events_detected: number;
events_evaluable: number;
};
metrics?: {
/** The quadrant-change rule that actually reaches Telegram. The headline. */
shipped?: {
rule: {
state_divider: number;
warning_divider: number;
margin: number;
confirm_sessions: number;
cooldown_days: number;
entry: string;
};
metrics: EventStudyMetrics;
events: { date: string; warned: boolean; lead_days: number | null }[];
quadrant_changes: number;
/** Debugging payload: every change the replay would have alerted on. Not rendered. */
fires: { index: number; date: string; from: string; to: string; state: number; warning: number }[];
/** Credit history starts partway through, so Warning is W1+W2 before it. */
by_era?: {
credit_from: string;
pre_credit: EventStudyMetrics & { label: string; start: string; end: string; sessions: number };
full_coverage: EventStudyMetrics & { label: string; start: string; end: string; sessions: number };
} | null;
};
/** The fundamental channel's actual exposure its rows are scored on this
* window, not on the market rows' full sample. */
fundamental_coverage?: {
observations: number;
/** Sessions with usable (observed, effective, non-stale) context. */
sessions_eligible: number;
evaluable_sessions: number;
/** Corrections whose warning horizon had usable context. */
events_covered: number;
events_evaluable: number;
minimum_events: number;
/** False until enough corrections are covered: the fundamental rows are
* untested, not failed, and must not render as a 0/N result. */
measurable: boolean;
};
/** Ablations, external baselines, and the fundamental channel all on fixed
* (unfitted) rules, so every row is scored on the same events. */
comparison?: (EventStudyMetrics & {
id: string;
label: string;
kind: 'ablation' | 'baseline' | 'fundamental';
note: string;
measurable: boolean;
})[];
null_model?: {
draws: number;
alarms_per_draw: number;
events: number;
events_warned: number;
events_missed: number;
alarm_episodes: number;
false_alarms: number;
false_alarms_per_year: number;
median_lead_days: number | null;
mean_warned: number;
sd_warned: number;
observed_warned: number;
p_at_least_observed: number;
} | null;
/** The original 70/30 fitted-threshold study, kept for continuity. */
fitted?: {
params: { train_fraction: number; warn_percentile: number; warn_threshold: number };
sample: { train_end: string; test_start: string; holdout_sessions: number };
metrics: EventStudyMetrics;
events: { date: string; warned: boolean; lead_days: number | null }[];
};
events?: { date: string; warned: boolean; lead_days: number | null }[];
recent_breadth?: { date: string; breadth: number; warning: number | null }[];
}
@@ -856,6 +965,10 @@ export interface Ticker {
symbol: string;
name: string | null;
created_at: string;
/** Set once the symbol stopped trading: excluded from signals, history kept. */
delisted_on: string | null;
/** How the delisting was learned: "form_25" (SEC confirmed) | "manual". */
delisted_reason: string | null;
}
// Admin
+437 -107
View File
@@ -6,6 +6,7 @@ import { Disclosure } from '../components/ui/Disclosure';
import { Badge } from '../components/ui/Badge';
import { SkeletonCard, SkeletonTable } from '../components/ui/Skeleton';
import { useAuthStore } from '../stores/authStore';
import { FUNDAMENTAL_VISUAL } from '../lib/regime';
import {
getEventStudy,
getRegimeConfig,
@@ -17,11 +18,13 @@ import {
} from '../api/regime';
import type {
CapexState,
FundamentalState,
EventStudyMetrics,
EventStudyReport,
GoodNewsReaction,
RegimeBand,
RegimeConfig,
RegimeFundamentalOverlay,
RegimeFundamentalContext,
RegimeFundamentals,
RegimeFundamentalsUpdate,
RegimeMonitor,
@@ -39,7 +42,7 @@ const BAND_STYLES: Record<RegimeBand, { text: string; bar: string; ring: string;
function TrendChip({ label, delta }: { label: string; delta: number | null | undefined }) {
if (delta == null) {
return <span className="rounded-lg bg-white/[0.04] px-2.5 py-1 text-xs text-gray-500">{label}: n/a</span>;
return <span className="rounded-lg bg-white/[0.04] px-2.5 py-1 text-xs text-gray-400">{label}: n/a</span>;
}
const color = delta === 0 ? 'text-gray-400' : delta > 0 ? 'text-red-400' : 'text-emerald-400';
const arrow = delta === 0 ? '→' : delta > 0 ? '↑' : '↓';
@@ -68,21 +71,21 @@ function ScoreGauge({
// shared set would mislabel one of them. Render none rather than wrong ones.
const ticks = bands ? [bands.watch, bands.elevated, bands.breaking] : [];
return (
<div className={`glass border p-6 ${style?.ring ?? 'border-white/[0.06]'}`}>
<div className={`glass h-full border p-5 ${style?.ring ?? 'border-white/[0.06]'}`}>
<div className="flex flex-wrap items-end justify-between gap-3">
<div>
<div className="text-[11px] uppercase tracking-wider text-gray-500">{label}</div>
<div className="text-xs uppercase tracking-wider text-gray-400">{label}</div>
<div className="mt-1 flex items-baseline gap-2">
<span className={`font-display text-6xl font-bold ${style?.text ?? 'text-gray-500'}`}>
<span className={`font-display text-5xl font-bold ${style?.text ?? 'text-gray-500'}`}>
{score == null ? '—' : Math.round(score)}
</span>
{score != null && <span className="text-sm text-gray-500">/ 100</span>}
{score != null && <span className="text-sm text-gray-400">/ 100</span>}
</div>
<div className="mt-1 flex flex-wrap items-center gap-2">
<span className={`text-sm font-medium ${style?.text ?? 'text-gray-500'}`}>
{style?.label ?? 'Incomplete'}
</span>
<span className="text-xs text-gray-600">coverage {Math.round(reading?.coverage ?? 0)}%</span>
<span className="text-xs text-gray-400">coverage {Math.round(reading?.coverage ?? 0)}%</span>
</div>
</div>
<div className="flex gap-2">
@@ -102,7 +105,7 @@ function ScoreGauge({
/>
</div>
{/* Thresholds come from the reading: the two axes no longer share them. */}
<div className="relative mt-1.5 h-4 text-[10px] uppercase tracking-wider text-gray-600">
<div className="relative mt-1.5 h-4 text-xs uppercase tracking-wider text-gray-400">
<span className="absolute left-0">0</span>
{ticks.map((tick) => (
<span key={tick} className="absolute -translate-x-1/2 num" style={{ left: `${tick}%` }}>
@@ -113,86 +116,153 @@ function ScoreGauge({
</div>
</>
)}
<p className="mt-4 text-xs text-gray-500">{footnote}</p>
<p className="mt-4 text-xs leading-relaxed text-gray-400">{footnote}</p>
</div>
);
}
const CAPEX_TONE: Record<CapexState, string> = {
raising: 'text-emerald-400',
holding: 'text-amber-400',
cutting: 'text-red-400',
unknown: 'text-gray-500',
/** Mirrors `_capex_signal`: holding is the neutral case, so it takes the neutral
* colour rather than an amber that reads as a third severity and sits close to
* the Warning channel's orange. */
const CAPEX_COLOR: Record<CapexState, string> = {
raising: FUNDAMENTAL_VISUAL.supportive.color,
holding: FUNDAMENTAL_VISUAL.neutral.color,
cutting: FUNDAMENTAL_VISUAL.adverse.color,
unknown: FUNDAMENTAL_VISUAL.unknown.color,
};
const OVERLAY_TITLE = 'Fundamental overlay · context, not scored';
function FundamentalOverlayCard({ overlay }: { overlay: RegimeFundamentalOverlay }) {
const capex = overlay.capex ?? {};
const reaction = overlay.good_news_stock_down;
// Nothing collected: the stored default is "unknown" for every hyperscaler
// and "mixed" for the reaction, which are placeholders, not a reading.
if (overlay.observed === false) {
return (
<div className="glass border border-white/[0.06] p-5">
<div className="text-[11px] uppercase tracking-wider text-gray-500">{OVERLAY_TITLE}</div>
<p className="mt-3 text-xs text-gray-500">
No observation collected yet. An admin can collect one under Admin · Monitor settings. It is
context only it never enters State or Warning.
</p>
</div>
);
function sentenceCase(value: string): string {
const text = value.replace(/_/g, ' ');
return text.charAt(0).toUpperCase() + text.slice(1);
}
function reactionReading(reaction: GoodNewsReaction | null): { label: string; color: string } {
switch (reaction) {
case 'yes':
return { label: 'Yes · good news sold', color: FUNDAMENTAL_VISUAL.adverse.color };
case 'no':
return { label: 'No · ordinary reactions', color: FUNDAMENTAL_VISUAL.supportive.color };
case 'mixed':
return { label: 'Mixed · no clear pattern', color: FUNDAMENTAL_VISUAL.neutral.color };
default:
return { label: 'Unknown · not observed', color: FUNDAMENTAL_VISUAL.unknown.color };
}
}
function FundamentalSummaryCard({ overlay }: { overlay: RegimeFundamentalContext }) {
const tone = FUNDAMENTAL_VISUAL[overlay.state] ?? FUNDAMENTAL_VISUAL.unknown;
const observed = overlay.observed ?? Boolean(overlay.fetched_at);
const status = !observed
? 'No usable observation. This channel remains Unknown.'
: overlay.pending
? `Collected now; enters the point-in-time record ${overlay.effective_date ?? 'next session'}.`
: overlay.stale
? 'The last state is retained for context, but stale evidence cannot confirm alerts.'
: !overlay.usable
? 'An observation was collected, but no signal could be determined.'
: null;
return (
<div className="glass border border-white/[0.06] p-5">
<div className="flex flex-wrap items-baseline justify-between gap-2">
<div className="text-[11px] uppercase tracking-wider text-gray-500">{OVERLAY_TITLE}</div>
<div className="flex flex-wrap items-center gap-2 text-[11px] text-gray-500">
{overlay.source && <span>{overlay.source}</span>}
{/* When pending, the line below is the single carrier of this date. */}
{overlay.effective_date && !overlay.pending && <span>· effective {overlay.effective_date}</span>}
<div className="glass h-full border p-5" style={{ borderColor: `${tone.color}33` }}>
<div className="flex flex-wrap items-start justify-between gap-2">
<div className="text-[11px] uppercase tracking-wider text-gray-400">Fundamentals · context</div>
<div className="flex flex-wrap gap-1.5">
{overlay.pending && <Badge label="pending" variant="manual" />}
{overlay.stale && <Badge label="stale" variant="manual" />}
</div>
</div>
{/* A pending observation is still shown it is the freshest read we
have, and nothing here is scored. The date says when the stored
point-in-time record picks it up. */}
{overlay.pending && (
<p className="mt-3 text-xs text-amber-400/90">
Shown as collected. The point-in-time record picks it up{' '}
{overlay.effective_date ?? 'next session'} observations are never backdated.
<div className="mt-2 flex flex-wrap items-baseline gap-3">
<span className="font-display text-4xl font-bold" style={{ color: tone.color }}>{tone.label}</span>
<span className="rounded-lg bg-white/[0.04] px-2.5 py-1 text-xs text-gray-300">
{sentenceCase(overlay.evidence_quality)} evidence
</span>
</div>
<div className="mt-4 grid grid-cols-2 gap-2 text-xs">
<div className="rounded-lg bg-white/[0.025] px-3 py-2">
<div className="text-gray-400">Capex</div>
<div className="mt-0.5 font-medium" style={{ color: FUNDAMENTAL_VISUAL[overlay.capex_signal].color }}>
{sentenceCase(overlay.capex_signal)}
</div>
</div>
<div className="rounded-lg bg-white/[0.025] px-3 py-2">
<div className="text-gray-400">Reaction</div>
<div className="mt-0.5 font-medium" style={{ color: FUNDAMENTAL_VISUAL[overlay.reaction_signal].color }}>
{sentenceCase(overlay.reaction_signal)}
</div>
</div>
</div>
{status && <p className="mt-3 text-xs leading-relaxed text-gray-400">{status}</p>}
{(overlay.source || overlay.effective_date) && (
<p className="mt-3 text-[11px] text-gray-400">
{overlay.source ?? 'stored observation'}
{overlay.effective_date && ` · effective ${overlay.effective_date}`}
</p>
)}
<div className="mt-4 grid gap-4 sm:grid-cols-2">
<div>
<div className="mb-2 flex items-baseline justify-between text-xs">
<span className="font-medium text-gray-300">Hyperscaler capex guidance</span>
<span className="num text-gray-500">{overlay.capex_stress ?? 'n/a'}</span>
</div>
<div className="space-y-1">
);
}
function FundamentalEvidence({ overlay }: { overlay: RegimeFundamentalContext }) {
const observed = overlay.observed ?? Boolean(overlay.fetched_at);
if (!observed) return null;
const capex = overlay.capex ?? {};
const reaction = reactionReading(overlay.good_news_stock_down);
return (
<Disclosure summary="Fundamental evidence · capex and earnings reaction">
<div className="grid gap-5 pt-1 sm:grid-cols-2">
<div>
<div className="mb-2 text-xs font-medium text-gray-200">Hyperscaler capex guidance</div>
{Object.keys(capex).length === 0 ? (
<p className="text-xs text-gray-400">No company-level observation.</p>
) : (
<div className="space-y-1.5">
{Object.entries(capex).map(([symbol, state]) => (
<div key={symbol} className="flex items-center justify-between text-xs">
<span className="font-mono text-gray-400">{symbol}</span>
<span className={CAPEX_TONE[state] ?? 'text-gray-500'}>{state}</span>
<span className="font-mono text-gray-300">{symbol}</span>
<span className="font-medium" style={{ color: CAPEX_COLOR[state] }}>{sentenceCase(state)}</span>
</div>
))}
</div>
)}
</div>
<div>
<div className="mb-2 flex items-baseline justify-between text-xs">
<span className="font-medium text-gray-300">Good news, stock down</span>
<span className="num text-gray-500">{overlay.earnings_stress ?? 'n/a'}</span>
</div>
<div className={`text-sm font-medium ${reaction === 'yes' ? 'text-red-400' : reaction === 'no' ? 'text-emerald-400' : 'text-gray-500'}`}>
{reaction === 'yes' ? 'Yes — beats sold into' : reaction === 'no' ? 'No — ordinary reactions' : 'Mixed'}
<div className="mb-2 text-xs font-medium text-gray-200">Good news, stock down</div>
<div className="text-sm font-medium" style={{ color: reaction.color }}>{reaction.label}</div>
<p className="mt-2 text-xs text-gray-400">
Derived context: capex {overlay.capex_signal} · reaction {overlay.reaction_signal}.
</p>
</div>
</div>
{overlay.reasoning && (
<details className="mt-4 border-t border-white/[0.06] pt-3">
<summary className="cursor-pointer text-xs font-medium text-gray-400 hover:text-gray-200">
Source reasoning
</summary>
<p className="mt-2 text-xs leading-relaxed text-gray-300">{overlay.reasoning}</p>
</details>
)}
</Disclosure>
);
}
function ConfluenceStrip({ warning, context }: { warning: RegimeReading; context?: RegimeFundamentalContext }) {
const warningElevated = warning.band === 'elevated' || warning.band === 'breaking';
if (!warningElevated || !context?.usable || context.state !== 'adverse') return null;
return (
<div className="glass-sm relative overflow-hidden px-4 py-3" role="status">
<span className="absolute inset-y-0 left-0 w-1 bg-gradient-to-b from-orange-400 to-red-400" aria-hidden="true" />
<div className="flex flex-wrap items-baseline gap-x-3 gap-y-1 pl-1">
<span className="text-xs font-semibold uppercase tracking-wider text-orange-300">Confluence active</span>
<span className="text-sm text-gray-200">
Warning is {warning.band}; point-in-time fundamentals are adverse with {context.evidence_quality} evidence.
</span>
<span className="text-xs text-gray-400">Condition only · never a combined score</span>
</div>
{overlay.reasoning && <p className="mt-4 text-xs leading-relaxed text-gray-400">{overlay.reasoning}</p>}
</div>
);
}
@@ -209,7 +279,7 @@ function PillarTable({ state, warning }: { state: RegimeReading; warning: Regime
<div className="overflow-x-auto rounded-lg border border-white/[0.06]">
<table className="w-full text-sm">
<thead>
<tr className="border-b border-white/[0.06] text-left text-xs uppercase tracking-wider text-gray-500">
<tr className="border-b border-white/[0.06] text-left text-xs uppercase tracking-wider text-gray-400">
<th className="px-4 py-3 font-medium">Pillar / sensor</th>
<th className="px-4 py-3 text-right font-medium">Score</th>
<th className="px-4 py-3 text-right font-medium">Weight</th>
@@ -221,7 +291,7 @@ function PillarTable({ state, warning }: { state: RegimeReading; warning: Regime
<tr className="border-b border-white/[0.06] bg-white/[0.02]">
<td colSpan={4} className="px-4 py-2 text-[11px] uppercase tracking-wider text-gray-400">
{title}
<span className="ml-2 normal-case tracking-normal text-gray-600">
<span className="ml-2 normal-case tracking-normal text-gray-400">
{reading.score ?? '—'} · {Math.round(reading.coverage)}% coverage
</span>
</td>
@@ -232,8 +302,8 @@ function PillarTable({ state, warning }: { state: RegimeReading; warning: Regime
<div className="font-medium text-gray-200">{pillar.label}</div>
<div className="mt-1 space-y-0.5">
{pillar.sensors.map((sensor) => (
<div key={sensor.id} className="text-xs text-gray-500">
<span className="font-mono text-gray-600">{sensor.id}</span> {sensor.label}:{' '}
<div key={sensor.id} className="text-xs text-gray-400">
<span className="font-mono text-gray-400">{sensor.id}</span> {sensor.label}:{' '}
<span className="num text-gray-400">{sensor.score == null ? 'n/a' : sensor.score}</span>
</div>
))}
@@ -256,7 +326,7 @@ function PillarTable({ state, warning }: { state: RegimeReading; warning: Regime
function MetaChip({ label, value, title }: { label: string; value: ReactNode; title?: string }) {
return (
<span className="rounded-lg bg-white/[0.03] px-2.5 py-1 text-[11px] text-gray-500" title={title}>
<span className="rounded-lg bg-white/[0.03] px-2.5 py-1 text-xs text-gray-400" title={title}>
{label} <span className="num text-gray-400">{value}</span>
</span>
);
@@ -288,56 +358,273 @@ function MetaStrip({ data }: { data: RegimeMonitor }) {
);
}
function EventStudyBody({ report }: { report: EventStudyReport }) {
const metrics = report.metrics;
function StatTiles({ metrics }: { metrics: EventStudyMetrics }) {
return (
<div className="space-y-4">
<div className="flex flex-wrap items-center gap-2">
<Badge label={report.evaluation ?? 'exploratory'} variant={report.evaluation === 'holdout' ? 'auto' : 'manual'} />
{report.generated_at && <span className="text-xs text-gray-500">generated {new Date(report.generated_at).toLocaleDateString()}</span>}
{report.sample && <span className="text-xs text-gray-500">test {report.sample.test_start} {report.sample.end}</span>}
</div>
<p className="text-sm leading-relaxed text-gray-300">{report.summary}</p>
{metrics && (
<div className="grid grid-cols-2 gap-2 sm:grid-cols-4">
{[
['Warned', `${metrics.events_warned}/${metrics.events}`],
['Missed', metrics.events_missed],
['False alarms/year', metrics.false_alarms_per_year.toFixed(1)],
['False alarms/year', metrics.false_alarms_per_year?.toFixed(1) ?? '—'],
['Median lead', metrics.median_lead_days == null ? '—' : `${metrics.median_lead_days}d`],
].map(([label, value]) => (
<div key={String(label)} className="rounded-lg border border-white/[0.06] bg-white/[0.02] px-3 py-2">
<div className="text-[11px] text-gray-500">{label}</div>
<div className="text-xs text-gray-400">{label}</div>
<div className="mt-0.5 text-lg font-semibold text-gray-200">{value}</div>
</div>
))}
</div>
)}
{report.events && report.events.length > 0 && (
);
}
function EventTable({ events }: { events: { date: string; warned: boolean; lead_days: number | null }[] }) {
return (
<div className="overflow-x-auto rounded-lg border border-white/[0.06]">
<table className="w-full text-xs">
<thead><tr className="border-b border-white/[0.06] text-left text-gray-500">
<thead><tr className="border-b border-white/[0.06] text-left text-gray-400">
<th className="px-3 py-2 font-medium">Correction</th>
<th className="px-3 py-2 text-right font-medium">Warned</th>
<th className="px-3 py-2 text-right font-medium">Lead</th>
</tr></thead>
<tbody>{report.events.map((event) => (
<tbody>{events.map((event) => (
<tr key={event.date} className="border-b border-white/[0.03] last:border-0">
<td className="px-3 py-2 num text-gray-300">{event.date}</td>
<td className={`px-3 py-2 text-right ${event.warned ? 'text-emerald-400' : 'text-gray-500'}`}>{event.warned ? 'yes' : 'no'}</td>
<td className={`px-3 py-2 text-right ${event.warned ? 'text-emerald-400' : 'text-gray-400'}`}>{event.warned ? 'yes' : 'no'}</td>
<td className="px-3 py-2 text-right num text-gray-300">{event.lead_days == null ? '—' : `${event.lead_days}d`}</td>
</tr>
))}</tbody>
</table>
</div>
);
}
/** Shipped rule against ablations, external baselines, and chance.
*
* The two kinds answer different questions and must not be read as one list:
* an ablation asks whether the quadrant machinery earns its place, a baseline
* asks whether the score earns its complexity.
*/
function ComparisonTable({ report }: { report: EventStudyReport }) {
const shipped = report.shipped;
if (!shipped || !report.comparison?.length) return null;
const rows = [
{
id: 'shipped',
label: 'Quadrant alert (shipped)',
kind: 'shipped' as const,
note: shipped.rule.entry,
measurable: true,
...shipped.metrics,
},
...report.comparison,
];
const KIND_LABEL: Record<string, string> = {
shipped: 'shipped',
ablation: 'ablation',
baseline: 'baseline',
fundamental: 'fundamental',
};
return (
<div className="space-y-2">
<div className="overflow-x-auto rounded-lg border border-white/[0.06]">
<table className="w-full text-xs">
<thead><tr className="border-b border-white/[0.06] text-left text-gray-400">
<th className="px-3 py-2 font-medium">Rule</th>
<th className="px-3 py-2 text-right font-medium">Warned</th>
<th className="px-3 py-2 text-right font-medium">FA/yr</th>
<th className="px-3 py-2 text-right font-medium">Median lead</th>
</tr></thead>
<tbody>{rows.map((row) => (
<tr
key={row.id}
className={`border-b border-white/[0.03] last:border-0 ${row.kind === 'shipped' ? 'bg-white/[0.03]' : ''}`}
title={row.note}
>
<td className={`px-3 py-2 ${row.kind === 'shipped' ? 'font-medium text-gray-200' : 'text-gray-400'}`}>
{row.label}
<span className="ml-2 text-xs uppercase tracking-wide text-gray-400">{KIND_LABEL[row.kind]}</span>
</td>
{/* A rule whose input does not exist yet scores 0/N, and printing
that would read as tested-and-failed. Say "not measurable". */}
{row.measurable === false ? (
<td className="px-3 py-2 text-right text-xs italic text-gray-400" colSpan={3}>
{/* Not "no observations yet": once some exist but fewer than
the minimum are covered, that is simply false. Matches the
callout below. */}
insufficient exposure not measurable
</td>
) : (
<>
<td className="px-3 py-2 text-right num text-gray-300">{row.events_warned}/{row.events}</td>
<td className="px-3 py-2 text-right num text-gray-300">{row.false_alarms_per_year?.toFixed(1) ?? '—'}</td>
<td className="px-3 py-2 text-right num text-gray-300">{row.median_lead_days == null ? '—' : `${row.median_lead_days}d`}</td>
</>
)}
</tr>
))}
{report.null_model && (
<tr className="border-t border-white/[0.06] text-gray-400">
<td className="px-3 py-2">
Random alarms, same firing rate
<span className="ml-2 text-xs uppercase tracking-wide text-gray-400">null</span>
</td>
<td className="px-3 py-2 text-right num">
{report.null_model.mean_warned.toFixed(1)} ± {report.null_model.sd_warned.toFixed(1)}
</td>
<td className="px-3 py-2 text-right num"></td>
<td className="px-3 py-2 text-right num"></td>
</tr>
)}</tbody>
</table>
</div>
</div>
);
}
function StudyVerdict({ report }: { report: EventStudyReport }) {
const model = report.null_model;
if (!model) return null;
const chancePct = (model.p_at_least_observed * 100).toFixed(0);
const indistinguishable = model.p_at_least_observed >= 0.1;
// The number carries the claim, not the adjective. At ~10 corrections a p of
// 0.09 is not evidence of anything, so "beats the null" would over-state a
// result this panel is otherwise careful never to over-state.
return (
<Callout variant={indistinguishable ? 'warning' : 'info'}>
<strong>
{indistinguishable
? `Not distinguishable from chance (p = ${model.p_at_least_observed.toFixed(2)}).`
: `Above the firing-rate null (p = ${model.p_at_least_observed.toFixed(2)}).`}
</strong>{' '}
Random alarms match or beat {model.observed_warned}/{model.events} warned corrections in {chancePct}% of{' '}
{model.draws} draws placing {model.alarms_per_draw} alarms over the same sessions. Corrections cluster and random
placement does not, so this is the floor, not the bar.
</Callout>
);
}
/** The credit sensor starts partway through, so Warning is a different
* construct either side of it. The share is derived, never asserted: if one era
* carries no corrections there is no comparison to draw and the per-era ratios
* would be noise dressed up as a finding. */
function EraDisclosure({
eras,
divider,
}: {
eras: NonNullable<NonNullable<EventStudyReport['shipped']>['by_era']>;
divider: number | undefined;
}) {
const { pre_credit: pre, full_coverage: full } = eras;
const total = pre.sessions + full.sessions;
const share = total > 0 ? Math.round((pre.sessions / total) * 100) : 0;
// An era holding one or two corrections has a recall of 0/1 or 1/2, which is
// not a rate. Below this the eras get their false-alarm rates compared and
// nothing else.
const comparable = pre.events >= 3 && full.events >= 3;
return (
<Disclosure summary={`Sensor-era caveat · ${share}% of sessions predate credit`}>
<p className="text-xs leading-relaxed text-gray-400">
<strong>{share}% of the evaluated sessions predate the credit sensor.</strong> W3 begins {eras.credit_from}, so
before that Warning renormalises to W1+W2 and the fixed {divider} divider is applied to a different construct
than it was reasoned about. Dropping the training split makes every correction evaluable; it does not make the
coverage gap go away, it moves it from the threshold to the score.
{comparable ? (
<>
{' '}Two sensors:{' '}
<strong className="text-gray-300">{pre.events_warned}/{pre.events}</strong> at{' '}
{pre.false_alarms_per_year?.toFixed(1) ?? '—'} FA/yr. All three:{' '}
<strong className="text-gray-300">{full.events_warned}/{full.events}</strong> at{' '}
{full.false_alarms_per_year?.toFixed(1) ?? '—'} FA/yr.
</>
) : (
<>
{' '}The corrections do not straddle that boundary ({pre.events} before, {full.events} after), so the two
eras cannot be compared on recall only the false-alarm rates are meaningful ({pre.false_alarms_per_year?.toFixed(1) ?? '—'}{' '}
vs {full.false_alarms_per_year?.toFixed(1) ?? '—'} per year).
</>
)}
</p>
</Disclosure>
);
}
function EventStudyBody({ report }: { report: EventStudyReport }) {
const shipped = report.shipped;
const eras = shipped?.by_era;
return (
<div className="space-y-4">
<div className="flex flex-wrap items-center gap-2">
<Badge label={report.evaluation ?? 'exploratory'} variant={report.evaluation === 'holdout' ? 'auto' : 'manual'} />
{report.generated_at && <span className="text-xs text-gray-400">generated {new Date(report.generated_at).toLocaleDateString()}</span>}
{report.sample && <span className="text-xs text-gray-400">{report.sample.evaluable_from} {report.sample.end}</span>}
</div>
<StudyVerdict report={report} />
<p className="text-sm leading-relaxed text-gray-300">{report.summary}</p>
{shipped && <StatTiles metrics={shipped.metrics} />}
<ComparisonTable report={report} />
{shipped && shipped.events.length > 0 && (
<Disclosure summary={`Correction details · ${shipped.events.length} events`}>
<EventTable events={shipped.events} />
</Disclosure>
)}
{report.null_model && (
<Disclosure summary="Null-model interpretation">
<p className="text-xs leading-relaxed text-gray-400">
The null places {report.null_model.alarms_per_draw} alarms at random over the same sessions and at the
shipped rule's firing rate. Corrections cluster while random placement does not, so this is a floor rather
than a demanding benchmark: a clustering rule could beat it without genuine foresight.
</p>
</Disclosure>
)}
{report.fundamental_coverage && !report.fundamental_coverage.measurable && (
<Disclosure
summary={`Fundamental exposure · ${report.fundamental_coverage.events_covered}/${report.fundamental_coverage.events_evaluable} corrections covered`}
>
<p className="text-xs leading-relaxed text-gray-400">
<strong>Insufficient exposure the fundamental rows are untested, not failed.</strong>{' '}
The channel had usable context on{' '}
<strong className="text-gray-300">
{report.fundamental_coverage.sessions_eligible} of{' '}
{report.fundamental_coverage.evaluable_sessions}
</strong>{' '}
evaluated sessions, covering{' '}
<strong className="text-gray-300">
{report.fundamental_coverage.events_covered} of{' '}
{report.fundamental_coverage.events_evaluable}
</strong>{' '}
corrections ({report.fundamental_coverage.minimum_events} needed;{' '}
{report.fundamental_coverage.observations} observation
{report.fundamental_coverage.observations === 1 ? '' : 's'} recorded). Those rows are
scored only on that window, never on the market rows' full sample otherwise a
fortnight of data would render as a 0/10 and read as a failed test. Read the market rows
as a verdict on the technical sensors and the alert machinery only.
</p>
</Disclosure>
)}
{eras && <EraDisclosure eras={eras} divider={shipped?.rule.warning_divider} />}
{report.fitted && (
<Disclosure summary={`Fitted-threshold variant · ${report.fitted.metrics.events_warned}/${report.fitted.metrics.events} on the 30% holdout`}>
<div className="space-y-3 pt-1">
<p className="text-xs leading-relaxed text-gray-400">
The original study, kept because it is what the methodology document reports: an{' '}
{report.fitted.params.warn_percentile}th-percentile Warning threshold (
{report.fitted.params.warn_threshold}) frozen on the first{' '}
{(report.fitted.params.train_fraction * 100).toFixed(0)}% of sessions and measured on the rest. Nothing
consumes this rule the shipped alert uses fixed dividers with hysteresis, confirmation and a cooldown.
</p>
<StatTiles metrics={report.fitted.metrics} />
{report.fitted.events.length > 0 && <EventTable events={report.fitted.events} />}
{report.reliability && (report.reliability.underpowered || report.reliability.sensor_coverage_mismatch) && (
<Callout variant="warning">
<div className="space-y-1.5">
{report.reliability.underpowered && (
<p>
<strong>Underpowered.</strong> Only {report.reliability.events_in_holdout} of{' '}
{report.reliability.events_detected} detected corrections fall in the test period (
{report.reliability.events_detected} detected corrections fall in the holdout (
{report.reliability.minimum_events}+ needed). Read the direction, not the ratio.
</p>
)}
@@ -355,6 +642,9 @@ function EventStudyBody({ report }: { report: EventStudyReport }) {
</Callout>
)}
</div>
</Disclosure>
)}
</div>
);
}
@@ -394,15 +684,26 @@ function FundamentalsEditor({
}) {
const [capex, setCapex] = useState<Record<string, CapexState>>(() => ({ ...data.capex }));
const [reaction, setReaction] = useState<GoodNewsReaction>(data.good_news_stock_down);
const knownCapex = Object.values(capex).filter((state) => state !== 'unknown');
// Mirrors _CAPEX_STATE_SCORES: raising 0, holding 50, cutting 100. Holding is
// the deceleration case and used to score identically to raising.
const capexPoints = knownCapex.reduce((sum, state) => sum + (state === 'cutting' ? 100 : state === 'holding' ? 50 : 0), 0);
const derivedF1 = knownCapex.length >= 3 ? Math.round((capexPoints / knownCapex.length) * 10) / 10 : null;
const derivedF3 = reaction === 'yes' ? 100 : reaction === 'no' ? 0 : null;
const values = Object.values(capex);
const counts = {
cutting: values.filter((s) => s === 'cutting').length,
holding: values.filter((s) => s === 'holding').length,
raising: values.filter((s) => s === 'raising').length,
unknown: values.filter((s) => s === 'unknown').length,
};
// Mirrors _capex_signal: any cut is adverse on partial evidence, any hold is
// neutral, all-known-raising is supportive, nothing known is unknown. No
// average — an average would let cuts and unknowns land on "neutral".
const capexSignal: FundamentalState =
counts.cutting > 0 ? 'adverse'
: counts.holding > 0 ? 'neutral'
: counts.raising > 0 ? 'supportive'
: 'unknown';
const reactionSignal: FundamentalState =
reaction === 'yes' ? 'adverse' : reaction === 'no' ? 'supportive' : reaction === 'mixed' ? 'neutral' : 'unknown';
return (
<div className="space-y-4">
<div className="flex flex-wrap items-center gap-2 text-xs text-gray-500">
<div className="flex flex-wrap items-center gap-2 text-xs text-gray-400">
<span>Source: {data.source}</span>
{data.fetched_at && <span>· fetched {new Date(data.fetched_at).toLocaleDateString()}</span>}
{data.effective_date && <span>· effective {data.effective_date}</span>}
@@ -411,8 +712,8 @@ function FundamentalsEditor({
{data.reasoning && <p className="text-xs leading-relaxed text-gray-400">{data.reasoning}</p>}
<div>
<div className="mb-2 flex items-center justify-between gap-3 text-xs">
<span className="font-medium text-gray-300">F1 · Capex guidance by hyperscaler</span>
<span className="num text-gray-500">score {derivedF1 ?? 'n/a'}</span>
<span className="font-medium text-gray-300">Capex guidance by hyperscaler</span>
<span style={{ color: FUNDAMENTAL_VISUAL[capexSignal].color }}>{capexSignal}</span>
</div>
<div className="grid grid-cols-2 gap-2">
{Object.entries(capex).map(([symbol, state]) => (
@@ -428,17 +729,24 @@ function FundamentalsEditor({
</label>
))}
</div>
<p className="mt-1.5 text-[11px] text-gray-600">Raising = 0, holding = 50, cutting = 100; at least three known names required.</p>
<p className="mt-1.5 text-xs text-gray-400">
{counts.cutting} cutting · {counts.holding} holding · {counts.raising} raising · {counts.unknown} unknown.
Any cut reads adverse on partial evidence; supportive needs every known name raising.
</p>
</div>
<label className="flex items-center justify-between gap-3 text-xs text-gray-400">
<span>
<span className="font-medium text-gray-300">F3 · Good news, stock down</span>
<span className="ml-2 num text-gray-600">score {derivedF3 ?? 'n/a'}</span>
<span className="font-medium text-gray-300">Good news, stock down</span>
<span className="ml-2" style={{ color: FUNDAMENTAL_VISUAL[reactionSignal].color }}>{reactionSignal}</span>
</span>
{/* "Mixed" is an observed mixed reaction; "unknown" is nobody looked or
the extraction failed. Collapsing them made a parse error read as
neutral evidence. */}
<select className={SELECT_CLASS} value={reaction} onChange={(event) => setReaction(event.target.value as GoodNewsReaction)}>
<option value="yes">Yes · stress</option>
<option value="no">No · ordinary</option>
<option value="mixed">Mixed · unavailable</option>
<option value="yes">Yes · good news sold</option>
<option value="no">No · reacting normally</option>
<option value="mixed">Mixed · observed, no clear pattern</option>
<option value="unknown">Unknown · not observed</option>
</select>
</label>
<div className="flex flex-wrap gap-2">
@@ -465,7 +773,7 @@ function ConfigEditor({ data, onSave, saving }: { data: RegimeConfig; onSave: (u
<input type="number" min={30} max={180} value={staleness} onChange={(event) => setStaleness(Number(event.target.value))} className="w-20 rounded-md border border-white/[0.08] bg-white/[0.03] px-2 py-1 text-right num text-gray-200" />
<span>days</span>
</label>
<p className="text-[11px] text-gray-600">Changing the basket resets its freeze date and silently reseeds quadrant alerts.</p>
<p className="text-xs text-gray-400">Changing the basket resets its freeze date and silently reseeds quadrant alerts.</p>
<button className="btn-primary px-3 py-1.5 text-sm disabled:opacity-50" disabled={saving || symbols.length < 20} onClick={() => onSave({ breadth_basket: symbols, fundamental_staleness_days: staleness })}>Save basket &amp; freshness</button>
</div>
);
@@ -483,13 +791,13 @@ function AdminControls() {
<Disclosure summary="Admin · Monitor settings">
<div className="grid gap-5 xl:grid-cols-2 xl:gap-6">
<section className="border-b border-white/[0.06] pb-5 xl:border-b-0 xl:border-r xl:pb-0 xl:pr-6">
<div className="mb-3 text-[11px] uppercase tracking-wider text-gray-500">Fundamental observations</div>
<div className="mb-3 text-xs uppercase tracking-wider text-gray-400">Fundamental observations</div>
{fundamentals.isLoading && <SkeletonCard className="h-36" />}
{fundamentals.data && <FundamentalsEditor key={fundamentals.dataUpdatedAt} data={fundamentals.data} onSave={(body) => saveFundamentals.mutate(body)} onRefresh={() => refresh.mutate()} saving={saveFundamentals.isPending} refreshing={refresh.isPending} />}
{refresh.isError && <Callout variant="error">Refresh failed: {(refresh.error as Error).message}</Callout>}
</section>
<section>
<div className="mb-3 text-[11px] uppercase tracking-wider text-gray-500">Fixed basket &amp; freshness</div>
<div className="mb-3 text-xs uppercase tracking-wider text-gray-400">Fixed basket &amp; freshness</div>
{config.isLoading && <SkeletonCard className="h-36" />}
{config.data && <ConfigEditor key={config.dataUpdatedAt} data={config.data} onSave={(updates) => saveConfig.mutate(updates)} saving={saveConfig.isPending} />}
{saveConfig.isError && <Callout variant="error">Save failed: {(saveConfig.error as Error).message}</Callout>}
@@ -508,7 +816,19 @@ export default function RegimePage() {
<div className="space-y-6 animate-slide-up">
<PageHeader
title="AI/Tech Risk Monitor"
subtitle="AI/Tech risk thermometer — observational only, feeds no entry, exit, or sizing decision"
subtitle="Market stress, early warning, and fundamental context"
actions={
<div className="flex flex-wrap items-center justify-end gap-2">
<Badge label="observational only" variant="default" />
{data?.date && <span className="num text-xs text-gray-400">as of {data.date}</span>}
{data?.available && (
<Badge
label={data.data_quality?.is_fresh ? 'fresh' : 'check data'}
variant={data.data_quality?.is_fresh ? 'auto' : 'manual'}
/>
)}
</div>
}
/>
{monitor.isLoading && <><SkeletonCard className="h-44" /><SkeletonTable rows={6} cols={4} /></>}
@@ -524,7 +844,7 @@ export default function RegimePage() {
</Callout>
)}
<div className="grid gap-4 lg:grid-cols-2">
<div className="grid gap-4 lg:grid-cols-3">
<ScoreGauge
label="State · stress right now"
reading={data.state}
@@ -542,17 +862,27 @@ export default function RegimePage() {
<ScoreGauge
label="Warning · deterioration & divergence"
reading={data.warning}
footnote="Breadth divergence, SMH/SPY rollover, and HY credit impulse. Missing sensors reduce coverage; they never default to 50."
footnote={
<>
Breadth divergence · SMH/SPY rollover · HY credit impulse.
{data.warning.coverage < 100 && ' Missing sensors are omitted rather than filled.'}
</>
}
/>
{data.fundamental_live && <FundamentalSummaryCard overlay={data.fundamental_live} />}
</div>
<ConfluenceStrip warning={data.warning} context={data.fundamental_context} />
<Suspense fallback={<SkeletonCard className="h-80" />}><RegimeChart /></Suspense>
{data.fundamental_live && <FundamentalEvidence overlay={data.fundamental_live} />}
<PillarTable state={data.state} warning={data.warning} />
{data.fundamental_context && <FundamentalOverlayCard overlay={data.fundamental_context} />}
<Disclosure summary="Data provenance · coverage and history">
<MetaStrip data={data} />
</Disclosure>
</>
)}
+39 -9
View File
@@ -1,31 +1,61 @@
import type { ReactNode } from 'react';
import { useSearchParams } from 'react-router-dom';
import { PageHeader } from '../components/ui/PageHeader';
import { Tabs } from '../components/ui/Tabs';
import { SetupsPanel } from '../components/signals/SetupsPanel';
import { TrackRecordPanel } from '../components/signals/TrackRecordPanel';
import { MyTradesPanel } from '../components/signals/MyTradesPanel';
import { BacktestPanel } from '../components/signals/BacktestPanel';
import { EvaluationPanel } from '../components/signals/EvaluationPanel';
const tabs = ['Setups', 'Track Record'] as const;
const tabs = ['Setups', 'Paper Trades', 'Backtest'] as const;
type Tab = (typeof tabs)[number];
// `track` stays the Paper Trades slug: App.tsx redirects the legacy /performance
// route to ?tab=track, and that is where realized results live.
const SLUG_TO_TAB: Record<string, Tab> = {
track: 'Paper Trades',
backtest: 'Backtest',
};
const TAB_TO_SLUG: Record<Tab, string> = {
Setups: '',
'Paper Trades': 'track',
Backtest: 'backtest',
};
const SUBTITLE: Record<Tab, string> = {
Setups: 'Detected trade setups from the latest scan',
'Paper Trades': 'What the strategy actually delivered on trades you took',
Backtest: 'Whether the promoted strategy is worth trading, replayed over history',
};
export default function SignalsPage() {
const [searchParams, setSearchParams] = useSearchParams();
const activeTab: Tab = searchParams.get('tab') === 'track' ? 'Track Record' : 'Setups';
const activeTab: Tab = SLUG_TO_TAB[searchParams.get('tab') ?? ''] ?? 'Setups';
const setTab = (tab: Tab) => {
setSearchParams(tab === 'Track Record' ? { tab: 'track' } : {}, { replace: true });
const slug = TAB_TO_SLUG[tab];
setSearchParams(slug ? { tab: slug } : {}, { replace: true });
};
const body: Record<Tab, ReactNode> = {
Setups: <SetupsPanel />,
'Paper Trades': <MyTradesPanel />,
// The backtest and the diagnostic that checks it against live outcomes.
Backtest: (
<div className="space-y-6">
<BacktestPanel />
<EvaluationPanel />
</div>
),
};
return (
<div className="space-y-6 animate-slide-up">
<PageHeader
title="Signals"
subtitle="Detected trade setups and how past signals actually performed"
/>
<PageHeader title="Signals" subtitle={SUBTITLE[activeTab]} />
<Tabs tabs={tabs} active={activeTab} onChange={setTab} />
<div className="animate-fade-in" key={activeTab}>
{activeTab === 'Setups' ? <SetupsPanel /> : <TrackRecordPanel />}
{body[activeTab]}
</div>
</div>
);
+19
View File
@@ -42,6 +42,14 @@
appearance: textfield;
}
/* --ink, not --up-text: a focus ring must not carry a semantic colour. The
directional token reads as "up/positive" and lands at poor contrast on the
controls that are already that colour. */
:where(button, a, input, select, textarea, summary, [tabindex]):focus-visible {
outline: 2px solid var(--ink);
outline-offset: 3px;
}
/* Atmosphere: faint starfield + soft rim-cyan / ember glows + film grain */
#root {
position: relative;
@@ -83,6 +91,17 @@
}
}
@media (prefers-reduced-motion: reduce) {
*,
*::before,
*::after {
animation-duration: 0.01ms !important;
animation-iteration-count: 1 !important;
scroll-behavior: auto !important;
transition-duration: 0.01ms !important;
}
}
@layer components {
/* Mars horizon — fixed at the viewport bottom, atmosphere only, never data */
.app-horizon {
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,88 @@
# Regime Monitor v4 calibration
Generated 2026-08-08T23:16:44 at `43ee619`, 2024-12-05 → 2026-07-24.
Source hashes (sha256, first 16):
- `app/services/regime_monitor_service.py``3b307e3c2045f35b`
- `app/services/breadth_service.py``b9ceb93d68d8f01b`
- `scripts/run_regime_monitor_calibration.py``9bc925256cc6f567`
## Hard gates
| gate | expected | measured | |
|---|---|---|---|
| symbols_fetched | 33 | 33 | ok |
| per_symbol_warmup_252_bars | all | 33 | ok |
| per_symbol_reaches_last_session | 2026-07-24 | 33 | ok |
| breadth_counts_full_basket | 30 | 408/408 sessions | ok |
| sessions_scored | 408 | 408 | ok |
| last_scored_date | 2026-07-24 | 2026-07-24 | ok |
| w1_available_every_session | 408 | 408 | ok |
| state_coverage_100_every_row | 0 | 0 | ok |
| no_stale_inputs | 0 | 0 | ok |
| first_scored_date | 2024-12-05 | 2024-12-05 | ok |
| state_v4_le_v3_every_row | 0 | 0 | ok |
## Distributions
| variant | avg | median | p80 | p90 | max |
|---|---|---|---|---|---|
| v2_reconstruction | 22.68 | 16.15 | 35.1 | 65.63 | 91.2 |
| v2_reconstruction_oas400 | 26.54 | 18.65 | 42.52 | 81.3 | 100.0 |
| v3 | 18.13 | 9.1 | 31.36 | 65.0 | 87.4 |
| v4 | 14.78 | 8.35 | 21.7 | 43.63 | 83.5 |
| v4-vix-only | 16.64 | 8.35 | 28.6 | 61.59 | 86.6 |
| v4-p1-only | 16.28 | 9.1 | 25.44 | 45.63 | 84.0 |
| v4-vix-b | 15.24 | 8.55 | 22.62 | 44.33 | 83.6 |
| v4-p1-capped | 14.73 | 8.35 | 21.7 | 43.63 | 80.1 |
## Saturation census (sessions pegged at 100)
| variant | P1 | P2 | P3 | V1 |
|---|---|---|---|---|
| v2_reconstruction | 46 | 0 | 39 | 14 |
| v2_reconstruction_oas400 | 46 | 0 | 39 | 14 |
| v3 | 46 | 0 | 0 | 14 |
| v4 | 0 | 0 | 0 | 0 |
| v4-vix-only | 46 | 0 | 0 | 0 |
| v4-p1-only | 0 | 0 | 0 | 14 |
| v4-vix-b | 0 | 0 | 0 | 0 |
| v4-p1-capped | 0 | 0 | 0 | 0 |
## Reproduction gates — v2_reconstruction
| figure | published | measured | |
|---|---|---|---|
| v2_state_avg | 22.6 | 22.68 | ok |
| v2_state_p80 | 35.1 | 35.1 | ok |
| v2_state_max | 91.2 | 91.2 | ok |
| v2_p3_pegged | 39 | 39 | ok |
| w1_live_sessions | 108 | 108 | ok |
## Reproduction gates — v3
| figure | published | measured | |
|---|---|---|---|
| v3_state_max | 87.4 | 87.4 | ok |
## v4 band-share grid (watch 20 / elevated 50)
| breaking | stable | watch | elevated | breaking |
|---|---|---|---|---|
| 60 | 78.9 | 13.0 | 2.9 | 5.1 |
| 65 | 78.9 | 13.0 | 4.7 | 3.4 |
| 70 | 78.9 | 13.0 | 6.9 | 1.2 |
## Scenarios (pillar arithmetic, explicit sensor scores)
| scenario | price | breadth | C1 | V1 | State |
|---|---|---|---|---|---|
| S1 ordinary tape | 7.5 | 0.0 | 0.0 | 4.0 | **3.6** |
| S2 10% correction, calm credit | 31.25 | 62.5 | 0.0 | 34.4 | **33.28** |
| S3a 2022-style, calm credit, no death cross | 90.83 | 100.0 | 0.0 | 60.0 | **70.33** |
| S3b 2022-style, calm credit, death cross | 100.0 | 100.0 | 0.0 | 60.0 | **74.0** |
| S4 credit event on top | 100.0 | 100.0 | 75.0 | 86.67 | **93.0** |
| S5 March 2020, everything pegged | 100.0 | 100.0 | 100.0 | 100.0 | **100.0** |
Recommendation: `{'state_bands_candidate': [20.0, 50.0, 65.0], 'provisional': False, 'note': 'confirm against band_grid + scenarios before shipping'}`
@@ -0,0 +1,242 @@
# SEC fundamentals alerts, 2026-08-21
Two `sec_facts` warnings, investigated against live SEC data. Both originate in SEC's
own published data — a stale per-company Company-Facts file (1) and a stale
ticker→CIK mapping (2) — and neither is a parser defect: no stored fundamental value
is wrong. Every SEC-side probe below reproduces offline from public endpoints; the
four database facts used are quoted where they appear.
## 1. `filing_gap_aged` — 43 gaps, all `not_in_companyfacts`
**Root cause: SEC's per-company Company-Facts files are stale for these issuers,
while the same filings are present in SEC's own `frames` aggregation.**
All ten named filings are real 10-Qs filed 2026-07-28/29, present in the issuer's
`submissions` with `isXBRL=1`, with complete R-files and XBRL in the EDGAR archive
— and absent from `companyfacts/CIK*.json`:
| CIK | issuer | accession | filed | in `companyfacts` | newest fact in file |
|---|---|---|---|---|---|
| 0000001800 | Abbott | 0001628280-26-050134 | 2026-07-28 | no | 2026-04-29 |
| 0000021344 | Coca-Cola | 0001628280-26-050503 | 2026-07-29 | no | 2026-04-30 |
| 0000024741 | Corning | 0000024741-26-000255 | 2026-07-29 | no | 2026-05-01 |
| 0000029989 | Omnicom | 0000029989-26-000019 | 2026-07-29 | no | 2026-04-29 |
| 0000037996 | Ford | 0000037996-26-000156 | 2026-07-29 | no | 2026-04-30 |
| 0000040533 | General Dynamics | 0000040533-26-000032 | 2026-07-29 | no | 2026-07-01 |
| 0000048898 | Hubbell | 0001628280-26-050405 | 2026-07-29 | no | 2026-06-04 |
| 0000049071 | Humana | 0000049071-26-000050 | 2026-07-29 | no | 2026-04-29 |
| 0000049196 | Huntington Bancshares | 0000049196-26-000066 | 2026-07-28 | no | 2026-04-30 |
| 0000062996 | Masco | 0000062996-26-000027 | 2026-07-29 | no | 2026-04-22 |
Ruled out, with evidence:
- **Not a global SEC outage.** Company Facts is current for other issuers filing the
same days — MSFT `0001193125-26-323660` @2026-07-29, AAPL @2026-07-31, P&G
@2026-08-04, Chevron @2026-08-06, JPMorgan @2026-08-20.
- **Not a CDN/cache artifact.** A cache-busted request with `Cache-Control: no-cache`
returns the identical stale 3.39 MB payload; the response carries no cache headers.
- **Not our filter.** The scan covers every taxonomy/concept/unit in the payload.
- **Not a metadata discriminator.** Gap and non-gap filings are identical on
`isXBRL`, `isInlineXBRL`, `reportDate`, `primaryDocDescription`.
- **SEC does have the facts.** `frames/us-gaap/Assets/USD/CY2026Q2I.json` lists
Abbott at exactly the missing accession `0001628280-26-050134`, and Coca-Cola and
Ford at theirs. The per-company endpoints are the degraded ones:
`companyconcept/CIK0000001800/us-gaap/Assets.json` returns `"units":{"USD":{}}`.
**Consequence, and why the gate changed.** Retrying `companyfacts` cannot recover
these — Abbott's file has been stale since April. And because `active_gaps`
supersedes a gap only on a *successfully ingested later* filing, a stale file also
swallows Q3: the pause was open-ended, not seasonal, on 43 large caps.
**Fix** (`app/services/fundamentals_quality_service.py`): once `filing_gap_aged` has
escalated a gap (`escalated_at`), it stops pausing setups **if** the issuer's own
newest stored 10-K/10-Q is under `GAP_GATE_RECENT_FILING_DAYS` (180) old. Pause hands
off to the alert; an issuer with nothing that recent stays paused. `active_gaps` is
deliberately untouched, so `_retry_backlog` keeps retrying and a recovered filing
still resolves normally. The bound is applied to the queue path *and* the
`validation_json` summary path, which mirrors the same filings — bounding only one
leaves the behaviour unchanged in production.
**This is a bounded reprieve, not a removal — know the two ways it ends.** Abbott's
newest ingested filing is `0001628280-26-028357`, filed 2026-04-29, so its recency
window closes around **2026-10-26**; most of the 43 sit on late-April filings and
turn back to paused within days of each other. That crossing is **silent**: the
importer escalates only gaps with `escalated_at IS NULL`, so `filing_gap_aged` does
not re-fire for a gap it has already reported. Separately, a Q3 10-Q that also fails
to ingest creates a *new* un-escalated gap on the same CIK, which re-pauses it at
once (that one does raise its own `filing_gap_aged` 14 days later). Whether the
silent re-block deserves a re-escalation signal is an open call, deliberately not
made here — "one actionable escalation rather than a daily warning" is the existing
design intent.
**Not done, with reasons.** A `frames`-backed recovery source was considered and
rejected: frames are calendar-aligned with a tolerance (off-fiscal filers drop out)
and carry one fact per issuer per period, so amendment/restatement semantics differ
from Company Facts — lossy as a snapshot source, not merely expensive. Parsing the
filing's own inline-XBRL instance is the authoritative alternative but is a new
subsystem (contexts, dimensions, unit refs) duplicating the parser's fact model.
## 2. `snapshot_discrepancy` — 0000906107-15-000012 / -000016
**Root cause: two tracked tickers claim the same filing, because SEC's
`company_tickers.json` still points the old symbol at a non-traded co-registrant.
No stored value is wrong and no reparse is warranted.**
CIK 0000906107 is **Vivmark Residential** (VMRK, formerly Equity Residential). Both
alerted accessions are **combined EQR + ERP Operating LP 10-Qs** — one accession, two
registrants (0000906107 and 0000931182) — the pattern behind the existing
co-registrant recovery path.
The stored rows are **byte-identical** to what the current parser reconstructs from
EQR's own Company Facts — every column, verified: `cik` (`0000906107`), `form`,
`filed_date`, `accepted_at`, both period dates, `fiscal_year`/`fiscal_period`,
`revenue`, `net_income`, `operating_income`, `diluted_eps`, `cfo`, the two nulls,
`cash_and_st_investments`, `total_debt` (340,900,000 / null),
`shares_outstanding`, `shares_outstanding_date`, `weighted_avg_diluted_shares`. Both
carry `import_run_id = 6`, and CIK 0000906107 holds all 69 of its filings across runs
630, so the issuer's own history is complete.
Run 63 (2026-08-19) recorded
`fields: ["cik"]` for both accessions, and the universe explains it:
```
tickers: VMRK -> 0000906107 (Vivmark Residential, ex-Equity Residential)
EQR -> 0000931182 (ERP Operating Ltd Partnership)
```
SEC's own `company_tickers.json` carries `{"cik_str": 931182, "ticker": "EQR",
"title": "ERP OPERATING LTD PARTNERSHIP"}` — after the rename, the old symbol stayed
attached to the **non-traded operating partnership**, the co-registrant on those
combined 10-Qs. `resolve_ciks` reads `active_only` tickers and follows SEC, so
0000931182 is tracked. Its Company Facts holds 7 accessions, exactly 2 of them
EQR-prefixed, so its backfill reconstructs exactly those two rows, stamps them
`cik=0000931182`, and collides with the rows already stored under 0000906107 —
identical in every fact, differing only in attribution.
It cannot self-heal. The collision loser never stores a row (the insert is skipped as
immutable), so `_ciks_with_snapshots` never sees 0000931182, and it is full-history
backfilled — refetching every submissions shard and its companyfacts — **on every
run**, re-raising the warning each time. `fundamental_snapshots` for 0000906107 holds
all 69 filings across runs 630, so the issuer's own history is complete and correct.
### Fixes
**Code** (`sec_fundamentals_importer.py`): a `cik`-only difference is no longer
reported as a reconstruction discrepancy. It raises `accession_cik_collision`, naming
both CIKs and pointing at `sec_cik_overrides`, because the fix is the universe, not
the parser. The reparse path also excludes these from its rewrite set — rewriting a
cik-only difference would re-stamp the filing onto the co-registrant and take it from
the issuer that filed it. (A reparse run while both CIKs are tracked fails validation
on `duplicate accession in staged snapshots` instead, which is a safe stop.)
**Data — needs an operator, and the alert repeats daily until then.** `EQR` is a stale
symbol: the security now trades as `VMRK`, which is already tracked at the correct
CIK. Retiring the `EQR` ticker ends the loop. A `sec_cik_overrides` pin of
`EQR -> 906107` would silence the collision but leave two tickers on one security,
double-counting the issuer in scans — retirement is the right action.
**Not fixed, deliberately:** the permanent-backfill loop itself. A tracked CIK whose
only parseable filings belong to another CIK is re-backfilled every run; ending that
in code means teaching `_ciks_with_snapshots` about foreign-owned accessions, which is
more state for a condition that is now loudly and specifically reported.
### Separate observation: `total_debt` on this issuer looks wrong
Independent of the alert, and unchanged by any fix here: the parser reconstructs
`total_debt = 340,900,000` for EQR's 2015 Q1 and `null` for Q2, while the REIT carried
roughly $10bn of debt. `_compose_debt` returns the short-term component alone when
every `_LONG_TERM_DEBT_AGG` concept **and** the `LongTermDebtNoncurrent`/`Current`
pair miss — which is what happened here, and Q2 matched neither. Worth checking
against a current REIT filer before trusting `total_debt` for that sector.
---
## 3. Follow-ups from the two alerts above
### 3a. The reprieve in (1) ended silently — now it doesn't
The hand-off in section 1 is a **bounded** reprieve. It ends two ways, and neither
said anything: the issuer's stored filings age past `GAP_GATE_RECENT_FILING_DAYS`
(for the 43, their last good filings are late April, so ~2026-10-26), or a newer
filing gap arrives and the all-escalated condition fails. `filing_gap_aged` cannot
report either, because it only escalates gaps whose `escalated_at` is NULL and so
never fires twice for the same gap.
`sec_filing_gaps.exempted_at` (migration `034`) makes the transition observable: set
quietly while the issuer is exempt, cleared when the exemption lapses, and the clear
is what raises `filing_gap_repaused`. Once per lapse, re-arming if the issuer's data
recovers and ages out again. A gap that was never exempt has no transition and stays
silent — it is simply still paused, which `filing_gap_aged` already said.
The exemption rule itself is not duplicated: `fundamentals_quality_service.gap_exempt_ciks`
is now public and the importer alerts on membership changes in exactly the set the
gate reads.
### 3b. `total_debt` was materially wrong for a third of large caps
The EQR observation in section 2 was not a REIT edge case. Measured over 19 large
caps, the old composition — `LongTermDebt`, else `LongTermDebtNoncurrent`/`Current`,
plus one of `ShortTermBorrowings`/`CommercialPaper` — missed two whole tagging styles:
| issuer | before | after | what was missed |
|---|---:|---:|---|
| T | None | 143.95b | `LongTermDebtAndCapitalLeaseObligations` |
| XOM | None | 47.66b | same |
| VZ | 21.78b | 165.23b | same (read only the current maturities) |
| KO | 0.25b | 39.31b | same (read only commercial paper) |
| HD | 3.50b | 48.33b | same |
| O | 1.40b | 26.53b | REIT parts (`NotesPayable` + `SecuredDebt`) |
| VMRK | 1.50b | 9.09b | same |
| CVX | 0.40b | **None** | partial suppressed — see below |
| PFE | 63.10b | 63.19b | `DebtCurrent` is the completer current side |
| 10 others | — | unchanged | already composed correctly |
`total_debt` feeds `net_debt``net_debt_to_ebitda` → the peer percentile and the
categorical leverage read, so Coca-Cola at 0.25bn of debt was not a missing value —
it was a confident *"conservative leverage"* on an issuer carrying ~39bn.
The composition now spans four mutually exclusive styles, with each concept's span
respected: `LongTermDebt` already includes current maturities (Apple tags all three
and 71.34 + 11.01 = 82.30 confirms it), `LongTermDebtAndCapitalLeaseObligations` is
noncurrent and needs a current complement, and `DebtCurrent` *is* that whole
complement rather than an addition to it.
**A short-term component alone is no longer reported as a total.** Chevron tags full
debt only in its 10-K, so its 10-Q carries 0.40bn of short-term borrowing and nothing
else. `_net_debt` needs both sides and yields nothing when either is missing, so None
costs a leverage read where the partial value produced a confidently wrong one.
The REIT branch needed disambiguating, because `NotesPayable` does not mean the same
thing across issuers (measured over 14 REITs): MAA tags `NotesPayable` 5.66bn =
`UnsecuredDebt` 5.30bn + `SecuredDebt` 0.36bn **exactly**, so there it is the total and
adding the secured side double-counts — while EQR tags it alongside a *larger*
`SecuredDebt` (5.38bn vs 6.38bn in 2013), where it is only the unsecured component.
`UnsecuredDebt`'s presence separates the two: where tagged it is the unambiguous
unsecured side and `NotesPayable` is ignored; where absent, `NotesPayable` is that
side. Both sides are required, which is also what stops the branch inventing a total
from a fragment.
| REIT | before | after | |
|---|---:|---:|---|
| MAA | None | 5.66b | matches its own `NotesPayable` total exactly |
| KIM | None | 8.74b | |
| O / VMRK | 1.40b / 1.50b | 26.53b / 9.09b | |
| BXP | 0.75b | **None** | tagged only `SecuredDebt` + paper against ~15bn real debt |
| VTR | 0.27b | **None** | same shape |
| 8 others | — | unchanged | already composed correctly |
Known limit: where EQR tags both the parts and the aggregate, the parts sum 2.612.2%
*below* it, so this branch approximates. It is last in line — any issuer tagging an
aggregate never reaches it — and the alternative there is no value at all.
### Sequencing the history fix
Snapshots are immutable, so **3b corrects new filings only**; every stored quarter
keeps its old `total_debt`. `scripts/reparse_fundamentals.py` exists for exactly this
("after a parser fix, keeping the stored row is preserving a stale cache").
**Retire the `EQR` ticker before reparsing.** A reparse backfills every tracked CIK,
so while both 0000906107 and 0000931182 are tracked, both stage the same two 2015
accessions and the run fails validation on `duplicate accession in staged snapshots`.
That is a safe stop — nothing is written — but the reparse will not complete until the
collision is gone.
+886
View File
@@ -0,0 +1,886 @@
"""Offline replay of the AI/Tech Risk Monitor, for calibrating a methodology cut.
Reproduces the State/Warning series session by session from the same inputs the
live job uses -- Alpaca for prices, FRED for VIX and HY OAS -- with no database,
so a sensor change can be measured against real history before it ships.
v3 was calibrated this way ad-hoc and the harness was never committed, which is
why its published numbers cannot be re-derived today. This is that harness.
**It never reimplements an unchanged live sensor.** ``_compute_index``,
``_score_pillars``, breadth, divergence, P2, P4 and the Warning sensors are
imported and called. Only *candidate* formulas (proposed for v4) and *retired*
ones (v2, no longer in the codebase) are defined here and patched onto the
service for the duration of a variant. Once a candidate ships, delete it here and
import the shipped function instead, or the two will drift.
The script refuses to emit a band recommendation unless every hard gate passes.
That is deliberate: it must be structurally impossible to read a calibration
result out of a run whose pipeline did not validate.
**Fundamental channel note.** The sourced capex / earnings read is a separate
categorical channel and is never a term in State or Warning, so every variant and
gate below is unaffected by it. This harness passes no observation, which means
the ``fundamental_context`` on each replayed row reads ``unknown`` -- correct, and
the same thing production reports for a session nobody observed. Calibrating
anything *about* that channel needs an observation series passed through
``_compute_index(..., observations=...)``, and enough history to be worth
calibrating against.
Research branch only. Example:
.\\.venv\\Scripts\\python.exe scripts\\run_regime_monitor_calibration.py ^
--end 2026-07-24 --sessions 408 --methodology v3,v4 ^
--cache-dir .calib-cache
"""
from __future__ import annotations
import argparse
import asyncio
import contextlib
import json
import math
import statistics
import subprocess
import sys
from collections.abc import Iterator
from copy import deepcopy
from datetime import date, datetime, timedelta
from pathlib import Path
from typing import Any, Callable
ROOT = Path(__file__).resolve().parents[1]
if str(ROOT) not in sys.path:
sys.path.insert(0, str(ROOT))
from app.config import settings # noqa: E402
from app.providers.alpaca import AlpacaOHLCVProvider # noqa: E402
from app.services import breadth_service # noqa: E402
from app.services import regime_monitor_service as rms # noqa: E402
# Published v3/v2 figures from docs/research/regime-monitor-v3.md. The v3 pair is
# DERIVED there (22.6 - 0.4; 91.2 - 3.8), not measured, so its tolerance is loose
# on purpose -- anything tighter would be false precision.
PUBLISHED = {
"window_end": "2026-07-24",
# The session count alone is tautological -- the harness slices the tail of
# leader_series, so it can only ever equal what was asked for. The start date
# is what actually validates the calendar.
"window_first": "2024-12-05",
"sessions": 408,
"w1_live_sessions": 108,
"v2_state_avg": 22.6,
"v2_state_p80": 35.1,
"v2_state_max": 91.2,
"v2_p3_pegged": 39,
# No v3 average is published: the doc's "-0.4" is measured against
# v3-with-the-percentile-leg, not against v2, so only the max is checkable.
"v3_state_max": 87.4,
}
# Every symbol comes from one source on one split-adjustment basis. Mixing a
# sqlite snapshot for the basket with Alpaca for the leaders would splice two
# adjustment bases mid-200-DMA for any symbol that split in between.
LEADER, CONFIRM, MARKET = "SMH", "QQQ", "SPY"
# ---------------------------------------------------------------------------
# Candidate formulas (proposed for v4) -- patched in, never shipped from here
# ---------------------------------------------------------------------------
P5_VIX_ANCHORS_A = ((15.0, 0.0), (20.0, 20.0), (25.0, 38.0), (30.0, 55.0), (40.0, 80.0), (55.0, 100.0))
P5_VIX_ANCHORS_B = ((15.0, 0.0), (20.0, 25.0), (25.0, 45.0), (30.0, 65.0), (40.0, 85.0), (55.0, 100.0))
P1_TREND_BREAK_ANCHORS = ((0.0, 20.0), (3.0, 35.0), (8.0, 55.0), (15.0, 75.0), (25.0, 100.0))
def _candidate_under_200(closes: list[float], anchors=P1_TREND_BREAK_ANCHORS) -> float | None:
"""Graduated trend break: 0 above the 200-DMA, else scaled by depth below it.
The live version returns a bare 0/100, which pins the price pillar's max()
at 100 through any real selloff and stops P3's ladder resolving. The step at
the crossing (0 -> 20) is kept deliberately: the break itself is a genuine
binary event and deserves a floor; only the depth past it is graduated.
"""
sma200 = rms._sma(closes, 200)
if sma200 is None or sma200 <= 0:
return None
pct_below = (sma200 - closes[-1]) / sma200 * 100.0
if pct_below <= 0:
return 0.0
return rms._clamp(rms._interpolate(pct_below, anchors))
def _candidate_p1(anchors=P1_TREND_BREAK_ANCHORS) -> Callable:
def p1_trend_break(smh, qqq, leader_weight: float = 2.0):
return rms._blend(
_candidate_under_200(smh, anchors), _candidate_under_200(qqq, anchors), leader_weight
)
return p1_trend_break
def _candidate_p5(anchors) -> Callable:
def p5_volatility(vix: float | None) -> float | None:
if vix is None:
return None
return rms._clamp(rms._interpolate(vix, anchors))
return p5_volatility
def _capped(fn: Callable, cap: float) -> Callable:
"""P1_SCORE_CAP fallback: cap the sensor score after the blend, before max()."""
def wrapped(smh, qqq, leader_weight: float = 2.0):
value = fn(smh, qqq, leader_weight)
return None if value is None else min(value, cap)
return wrapped
# ---------------------------------------------------------------------------
# Retired formulas -- reconstructed, no longer in the codebase
# ---------------------------------------------------------------------------
def _v3_under_200(closes: list[float]) -> float | None:
"""v3's binary trend break, retired when v4 graduated it."""
sma200 = rms._sma(closes, 200)
if sma200 is None:
return None
return 100.0 if closes[-1] < sma200 else 0.0
def _v3_p1_trend_break(smh, qqq, leader_weight: float = 2.0) -> float | None:
return rms._blend(_v3_under_200(smh), _v3_under_200(qqq), leader_weight)
def _v3_p5_volatility(vix: float | None) -> float | None:
"""v3's linear VIX ramp, retired when v4 anchored it. Saturated at 30."""
if vix is None:
return None
return rms._clamp((vix - 15.0) / 15.0 * 100.0)
def _v2_drawdown(closes: list[float]) -> float | None:
if len(closes) < 30:
return None
peak = max(closes[-252:])
if peak <= 0:
return None
return rms._clamp((peak - closes[-1]) / peak * 100.0 * 5.0)
def _v2_p3_drawdown(smh, qqq, leader_weight: float = 2.0) -> float | None:
"""v2 took max() across the legs, so the more volatile leader always won."""
vals = [v for v in (_v2_drawdown(smh), _v2_drawdown(qqq)) if v is not None]
return max(vals) if vals else None
def _v2_divergence_series(breadth, benchmark_closes, lookback: int = 20):
"""v2's hard price gate: the sensor ZEROED during any decline.
v3 replaced this with a taper, which is why v2 shows W1 nonzero on only 108
of 408 sessions while v3 shows it nonzero far more often. Reconstructing it
is the only way to check that published figure.
"""
bench = {d: c for d, c in benchmark_closes}
common = sorted(d for d in bench if d in breadth)
out: dict[date, float] = {}
for i in range(lookback, len(common)):
d, d0 = common[i], common[i - lookback]
if bench[d0] <= 0:
continue
price_ret = (bench[d] / bench[d0] - 1.0) * 100.0
deterioration = max(0.0, -(breadth[d] - breadth[d0]))
score = deterioration * 5.0 if price_ret >= 0 else 0.0
out[d] = max(0.0, min(100.0, round(score, 2)))
return out
def _v2_f2_credit_spreads(oas_values: list[float]) -> float | None:
"""70% named anchors + 30% upper-tail percentile over whatever window it got."""
if not oas_values:
return None
latest = oas_values[-1]
absolute = rms._oas_absolute_score(latest)
if len(oas_values) < 30:
return round(absolute, 2)
less = sum(1 for v in oas_values if v < latest)
equal = sum(1 for v in oas_values if v == latest)
percentile = (less + 0.5 * equal) / len(oas_values) * 100.0
relative = rms._clamp((percentile - 50.0) / 45.0 * 100.0)
return round(absolute * 0.7 + relative * 0.3, 2)
# ---------------------------------------------------------------------------
# Variants
# ---------------------------------------------------------------------------
VARIANTS: dict[str, dict[str, Callable]] = {
# Retired since the v4 cutover -- "nothing patched" is now v4, so v3 has to
# be reconstructed like v2 to stay comparable.
"v3": {
"p1_trend_break": _v3_p1_trend_break,
"p5_volatility": _v3_p5_volatility,
},
# SHIPPED as of v4 -- nothing patched, so this variant exercises live code.
# Keeping a private copy here would let the harness and the service drift.
"v4": {},
"v4-vix-b": {
"p1_trend_break": _candidate_p1(),
"p5_volatility": _candidate_p5(P5_VIX_ANCHORS_B),
},
"v4-p1-capped": {
"p1_trend_break": _capped(_candidate_p1(), 50.0),
"p5_volatility": _candidate_p5(P5_VIX_ANCHORS_A),
},
"v4-vix-only": {"p1_trend_break": _v3_p1_trend_break},
"v4-p1-only": {"p5_volatility": _v3_p5_volatility},
# (4) v2 as production actually fetched it: a 400-calendar-day OAS source,
# which left the oldest rows with no credit at all. Truncating the SERIES is
# the only faithful simulation -- patching the per-session window is not,
# because the data was simply absent.
"v2_reconstruction_oas400": {
"p1_trend_break": _v3_p1_trend_break,
"p5_volatility": _v3_p5_volatility,
"p3_drawdown": _v2_p3_drawdown,
"f2_credit_spreads": _v2_f2_credit_spreads,
"HY_OAS_WINDOW_DAYS": 3653,
},
# v2 State sensors + the v2 divergence gate that feeds W1. The v2 *Warning
# composition* (F1/F3 fundamentals, 20 of 100 points) is NOT reconstructed,
# so only State statistics and the W1 census are comparable to the published
# v2 figures -- not the Warning score.
"v2_reconstruction": {
# v2 shared v3's binary trend break and linear VIX ramp verbatim, so both
# are retired now and must be restored here too -- otherwise a "v2" replay
# silently picks up v4's graded sensors.
"p1_trend_break": _v3_p1_trend_break,
"p5_volatility": _v3_p5_volatility,
"p3_drawdown": _v2_p3_drawdown,
"f2_credit_spreads": _v2_f2_credit_spreads,
# v2 sliced HY_OAS_REFERENCE_YEARS = 10.0 per session. The percentile leg
# ranks the current spread against that window, so replaying it against
# a 700-day slice gives systematically different mid-distribution scores.
"HY_OAS_WINDOW_DAYS": 3653,
},
}
# Variants needing the retired divergence formula rather than the live one.
V2_DIVERGENCE_VARIANTS = {"v2_reconstruction", "v2_reconstruction_oas400"}
# Variants whose OAS *source series* is truncated before replay, in calendar days.
OAS_SOURCE_TRUNCATION = {"v2_reconstruction_oas400": 400}
# A v4 recommendation is meaningless without both of these: the row-wise
# state_v4 <= state_v3 invariant needs them, and it is a hard gate.
# v2_reconstruction is required too: it carries every published figure the
# reproduction rests on (avg/p80/max/P3-pegged/W1-live). Without it a run could
# emit a confident recommendation having checked nothing against v2 at all,
# while the methodology doc claims v2 and v3 are reproduced first.
REQUIRED_VARIANTS = ("v2_reconstruction", "v3", "v4")
@contextlib.contextmanager
def patched(overrides: dict[str, Callable]) -> Iterator[None]:
"""Swap functions on the service module, then restore exactly."""
original = {name: getattr(rms, name) for name in overrides}
try:
for name, fn in overrides.items():
setattr(rms, name, fn)
yield
finally:
for name, fn in original.items():
setattr(rms, name, fn)
# ---------------------------------------------------------------------------
# Inputs
# ---------------------------------------------------------------------------
def _basket(config: dict) -> list[str]:
return list(config["breadth_basket"])
def _all_symbols(config: dict) -> list[str]:
return list(dict.fromkeys(_basket(config) + [LEADER, CONFIRM, MARKET]))
async def _load_prices(
symbols: list[str], start: date, end: date, cache_dir: Path | None, quiet: bool
) -> dict[str, list[tuple[date, float]]]:
cache = None
if cache_dir:
cache_dir.mkdir(parents=True, exist_ok=True)
cache = cache_dir / f"prices-{start}-{end}.json"
if cache.exists():
raw = json.loads(cache.read_text(encoding="utf-8"))
if set(raw) >= set(symbols):
if not quiet:
print(f"prices: cache hit ({len(raw)} symbols)", flush=True)
return {
s: [(date.fromisoformat(d), float(c)) for d, c in raw[s]] for s in symbols
}
provider = AlpacaOHLCVProvider(settings.alpaca_api_key, settings.alpaca_api_secret)
out: dict[str, list[tuple[date, float]]] = {}
for index, symbol in enumerate(symbols, 1):
bars = await provider.fetch_ohlcv(symbol, start, end)
out[symbol] = sorted((b.date, float(b.close)) for b in bars)
if not quiet:
print(f" [{index}/{len(symbols)}] {symbol}: {len(out[symbol])} bars", flush=True)
if cache:
cache.write_text(
json.dumps({s: [[d.isoformat(), c] for d, c in v] for s, v in out.items()}),
encoding="utf-8",
)
return out
async def _load_fred(series_id: str, start: date, end: date, cache_dir: Path | None):
cache = cache_dir / f"{series_id}-{start}-{end}.json" if cache_dir else None
if cache and cache.exists():
raw = json.loads(cache.read_text(encoding="utf-8"))
return [(date.fromisoformat(d), float(v)) for d, v in raw]
series = await rms._fetch_fred_series(series_id, start, end)
if cache and series:
cache.write_text(
json.dumps([[d.isoformat(), v] for d, v in series]), encoding="utf-8"
)
return series
# ---------------------------------------------------------------------------
# Replay
# ---------------------------------------------------------------------------
def _replay(
variant: str,
prices: dict[str, list[tuple[date, float]]],
vix, oas, config: dict, sessions: list[date],
breadth_series, divergence_by_variant: dict[str, Any], breadth_counts,
) -> list[dict]:
# Divergence is computed outside _compute_index, so the retired v2 gate has
# to be selected here rather than patched onto the module.
divergence_series = divergence_by_variant[
"v2" if variant in V2_DIVERGENCE_VARIANTS else "live"
]
truncate_days = OAS_SOURCE_TRUNCATION.get(variant)
if truncate_days is not None and oas:
cutoff = max(d for d, _ in oas) - timedelta(days=truncate_days)
oas = [(d, v) for d, v in oas if d >= cutoff]
rows: list[dict] = []
with patched(VARIANTS[variant]):
for as_of in sessions:
rows.append(
rms._compute_index(
prices, vix, oas, {}, deepcopy(config), as_of,
breadth_series, divergence_series, breadth_counts,
)
)
return rows
def _percentile(values: list[float], pct: float) -> float | None:
if not values:
return None
ordered = sorted(values)
k = (len(ordered) - 1) * pct / 100.0
lo, hi = math.floor(k), math.ceil(k)
if lo == hi:
return ordered[int(k)]
return ordered[lo] + (ordered[hi] - ordered[lo]) * (k - lo)
def _sensor_score(row: dict, pillar_id: str, sensor_id: str) -> float | None:
"""Search both axes: W1/W2/W3 live under ``warning``, P*/B1/C1/V1 under ``state``."""
for axis in ("state", "warning"):
for pillar in row[axis]["pillars"]:
if pillar["id"] == pillar_id:
for sensor in pillar["sensors"]:
if sensor["id"] == sensor_id:
return sensor["score"]
return None
def _band_shares(scores: list[float], bands: tuple[float, float, float]) -> dict[str, float]:
if not scores:
return {}
counts = {"stable": 0, "watch": 0, "elevated": 0, "breaking": 0}
for score in scores:
counts[rms.band_for(score, bands)] += 1 # bands passed explicitly -- see module docstring
return {k: round(v / len(scores) * 100.0, 1) for k, v in counts.items()}
def _stats(rows: list[dict], label: str) -> dict:
states = [r["state"]["score"] for r in rows if r["state"]["score"] is not None]
warnings = [r["warning"]["score"] for r in rows if r["warning"]["score"] is not None]
def pegged(pillar: str, sensor: str) -> int:
return sum(1 for r in rows if (_sensor_score(r, pillar, sensor) or 0) >= 100.0)
argmax_sole, argmax_tied, ties = {"P1": 0, "P2": 0, "P3": 0}, {"P1": 0, "P2": 0, "P3": 0}, 0
# The P1_SCORE_CAP rule is "sole argmax on >80% of sessions with State >= 40",
# so the all-session count does not evaluate it. Track the conditional
# population separately rather than deciding off the wrong denominator.
stressed_sole, stressed_total = {"P1": 0, "P2": 0, "P3": 0}, 0
for row in rows:
legs = {s: _sensor_score(row, "price", s) for s in ("P1", "P2", "P3")}
live = {k: v for k, v in legs.items() if v is not None}
if not live:
continue
top = max(live.values())
winners = [k for k, v in live.items() if v == top]
if len(winners) > 1:
ties += 1
for w in winners:
argmax_tied[w] += 1
if len(winners) == 1:
argmax_sole[winners[0]] += 1
if (row["state"]["score"] or 0) >= 40.0:
stressed_total += 1
if len(winners) == 1:
stressed_sole[winners[0]] += 1
return {
"label": label,
"sessions": len(rows),
"state": {
"avg": round(statistics.fmean(states), 2) if states else None,
"median": round(statistics.median(states), 2) if states else None,
"p80": round(_percentile(states, 80), 2) if states else None,
"p90": round(_percentile(states, 90), 2) if states else None,
"max": round(max(states), 2) if states else None,
"scored": len(states),
},
"warning_avg": round(statistics.fmean(warnings), 2) if warnings else None,
"saturation_census": {
"p3_pegged": pegged("price", "P3"),
"p2_pegged": pegged("price", "P2"),
"v1_pegged": pegged("volatility", "V1"),
"p1_pegged": pegged("price", "P1"),
},
"price_argmax_sole": argmax_sole,
"price_argmax_tie_inclusive": argmax_tied,
"price_argmax_ties": ties,
"price_argmax_when_state_ge_40": {
"sessions": stressed_total,
"sole": stressed_sole,
"p1_sole_share_pct": round(stressed_sole["P1"] / stressed_total * 100.0, 1)
if stressed_total else None,
"cap_rule": "P1_SCORE_CAP warranted if p1_sole_share_pct > 80",
},
"w1_nonzero_sessions": sum(
1 for r in rows if (_sensor_score(r, "breadth_divergence", "W1") or 0) > 0
),
"band_shares_current": _band_shares(states, rms.STATE_BANDS),
}
# ---------------------------------------------------------------------------
# Gates
# ---------------------------------------------------------------------------
def _completeness_gates(
prices: dict, symbols: list[str], sessions: list[date], breadth_counts: dict, basket_size: int
) -> list[dict]:
"""Without these, every gate below can pass on partial data.
_breadth_with_counts publishes on min_tickers=20, so 20 of 30 basket names
still yields 100% State coverage and a plausible W1 count.
"""
first, last = sessions[0], sessions[-1]
missing = [s for s in symbols if not prices.get(s)]
thin = [
s for s in symbols
if len([d for d, _ in prices.get(s, []) if d < first]) < 252
]
truncated = [s for s in symbols if not prices.get(s) or prices[s][-1][0] < last]
short_basket = sorted(
d.isoformat() for d in sessions if breadth_counts.get(d, 0) != basket_size
)
return [
{"gate": "symbols_fetched", "expected": len(symbols),
"measured": len(symbols) - len(missing), "passed": not missing, "detail": missing},
{"gate": "per_symbol_warmup_252_bars", "expected": "all",
"measured": len(symbols) - len(thin), "passed": not thin, "detail": thin},
{"gate": "per_symbol_reaches_last_session", "expected": last.isoformat(),
"measured": len(symbols) - len(truncated), "passed": not truncated, "detail": truncated},
{"gate": "breadth_counts_full_basket", "expected": basket_size,
"measured": f"{len(sessions) - len(short_basket)}/{len(sessions)} sessions",
"passed": not short_basket, "detail": short_basket[:20]},
]
def _pipeline_gates(rows: list[dict], sessions: list[date], expected_first: str) -> list[dict]:
coverage_bad = [
r["date"] for r in rows if (r["state"]["coverage"] or 0) < 100.0
]
w1_available = sum(
1 for r in rows if _sensor_score(r, "breadth_divergence", "W1") is not None
)
stale = [r["date"] for r in rows if r["data_quality"]["stale_inputs"]]
gates = [
{"gate": "sessions_scored", "expected": PUBLISHED["sessions"],
"measured": len(rows), "passed": len(rows) == PUBLISHED["sessions"]},
{"gate": "last_scored_date", "expected": PUBLISHED["window_end"],
"measured": rows[-1]["date"], "passed": rows[-1]["date"] == PUBLISHED["window_end"]},
# Availability, not the published "W1 live 108" -- that figure counts
# NONZERO sessions under v2's hard price gate and is checked there.
{"gate": "w1_available_every_session", "expected": len(rows),
"measured": w1_available, "passed": w1_available == len(rows)},
{"gate": "state_coverage_100_every_row", "expected": 0,
"measured": len(coverage_bad), "passed": not coverage_bad, "detail": coverage_bad[:20]},
{"gate": "no_stale_inputs", "expected": 0,
"measured": len(stale), "passed": not stale, "detail": stale[:20]},
]
# Unconditional: sessions_scored is tautological when the harness slices the
# tail of leader_series, so the start date is the only real calendar check.
# An optional gate is not a gate.
gates.append({
"gate": "first_scored_date", "expected": expected_first,
"measured": rows[0]["date"], "passed": rows[0]["date"] == expected_first,
})
return gates
def _invariant_gate(v3_rows: list[dict], v4_rows: list[dict]) -> dict:
"""state_v4 <= state_v3 on every aligned row.
Provable, not heuristic: graduated P1 never exceeds binary P1, anchored VIX
never exceeds (vix-15)/15*100, max() is monotone in its arguments, and no
other State sensor or weight changes. A violation means the harness is
mis-wired, not that the calibration is interesting.
"""
violations = []
for a, b in zip(v3_rows, v4_rows):
assert a["date"] == b["date"], "row misalignment"
s3, s4 = a["state"]["score"], b["state"]["score"]
if s3 is not None and s4 is not None and s4 > s3 + 1e-9:
violations.append({"date": a["date"], "v3": s3, "v4": s4})
return {
"gate": "state_v4_le_v3_every_row", "expected": 0,
"measured": len(violations), "passed": not violations, "detail": violations[:20],
}
def _soft_gates(stats: dict, variant: str) -> list[dict]:
if variant == "v3":
# The doc states no v3 average: its "-0.4" is measured against
# v3-with-the-percentile-leg, not against v2. Only the max is checkable.
pairs = [("v3_state_max", stats["state"]["max"], 0.5)]
else:
pairs = [("v2_state_avg", stats["state"]["avg"], 0.3),
("v2_state_p80", stats["state"]["p80"], 0.5),
("v2_state_max", stats["state"]["max"], 0.5),
("v2_p3_pegged", stats["saturation_census"]["p3_pegged"], 0),
("w1_live_sessions", stats["w1_nonzero_sessions"], 0)]
out = []
for key, measured, tol in pairs:
expected = PUBLISHED[key]
ok = measured is not None and abs(measured - expected) <= tol
out.append({"gate": key, "expected": expected, "measured": measured,
"tolerance": tol, "passed": ok})
return out
# ---------------------------------------------------------------------------
# Scenarios -- pillar arithmetic, stated as explicit sensor scores
# ---------------------------------------------------------------------------
def _scenarios(vix_anchors, p1_anchors) -> list[dict]:
"""The meaning anchors for the band choice, machine-checked rather than prose.
Stated as explicit sensor scores because drawdown + VIX + OAS does not
determine State: the price pillar is max(P1, P2, P3) and P2 is set by the
50/200-DMA gap, which no drawdown figure implies.
"""
def state(p1, p2, p3, breadth, c1, v1):
price = max(p1, p2, p3)
return round((price * 40 + breadth * 25 + c1 * 20 + v1 * 15) / 100, 2)
def p1_at(pct_below):
return round(rms._interpolate(pct_below, p1_anchors), 2) if pct_below > 0 else 0.0
def p3_at(dd):
return round(rms._interpolate(dd, rms.P3_DRAWDOWN_ANCHORS), 2)
def v1_at(vix):
return round(rms._interpolate(vix, vix_anchors), 2)
rows = [
("S1 ordinary tape", 0.0, 0.0, p3_at(3), rms.breadth_level_score(65), 0.0, v1_at(16)),
("S2 10% correction, calm credit", p1_at(2), 0.0, p3_at(10), rms.breadth_level_score(35), 0.0, v1_at(24)),
("S3a 2022-style, calm credit, no death cross", p1_at(20), 0.0, p3_at(35), rms.breadth_level_score(8), 0.0, v1_at(32)),
("S3b 2022-style, calm credit, death cross", p1_at(20), 100.0, p3_at(35), rms.breadth_level_score(8), 0.0, v1_at(32)),
("S4 credit event on top", p1_at(25), 100.0, p3_at(40), rms.breadth_level_score(5), rms.f2_credit_spreads([6.0]), v1_at(45)),
("S5 March 2020, everything pegged", 100.0, 100.0, 100.0, 100.0, 100.0, 100.0),
]
return [
{"scenario": name, "P1": p1, "P2": p2, "P3": p3, "price": max(p1, p2, p3),
"breadth": br, "C1": c1, "V1": v1, "state": state(p1, p2, p3, br, c1, v1)}
for name, p1, p2, p3, br, c1, v1 in rows
]
def _markdown(report: dict) -> str:
"""Scannable sibling to the JSON. Never written over a curated docs/ file."""
lines = ["# Regime Monitor v4 calibration", ""]
src = report["source"]
dirty = " **(dirty working tree)**" if src.get("git_dirty") else ""
lines.append(f"Generated {report['generated_at']} at `{src['git_rev']}`{dirty}, "
f"{report['provenance']['scored_range'][0]}{report['provenance']['scored_range'][1]}.")
lines += ["", "Source hashes (sha256, first 16):", ""]
for rel, digest in src["source_sha256"].items():
lines.append(f"- `{rel}` — `{digest}`")
lines += ["", "## Hard gates", "", "| gate | expected | measured | |", "|---|---|---|---|"]
for g in report["hard_gates"]:
lines.append(f"| {g['gate']} | {g['expected']} | {g['measured']} | {'ok' if g['passed'] else '**FAIL**'} |")
lines += ["", "## Distributions", "", "| variant | avg | median | p80 | p90 | max |", "|---|---|---|---|---|---|"]
for name, v in report["variants"].items():
st = v["state"]
lines.append(f"| {name} | {st['avg']} | {st['median']} | {st['p80']} | {st['p90']} | {st['max']} |")
lines += ["", "## Saturation census (sessions pegged at 100)", "",
"| variant | P1 | P2 | P3 | V1 |", "|---|---|---|---|---|"]
for name, v in report["variants"].items():
c = v["saturation_census"]
lines.append(f"| {name} | {c['p1_pegged']} | {c['p2_pegged']} | {c['p3_pegged']} | {c['v1_pegged']} |")
for name, v in report["variants"].items():
if v.get("soft_gates"):
lines += ["", f"## Reproduction gates — {name}", "",
"| figure | published | measured | |", "|---|---|---|---|"]
for g in v["soft_gates"]:
lines.append(f"| {g['gate']} | {g['expected']} | {g['measured']} | {'ok' if g['passed'] else '**miss**'} |")
if "v4" in report["variants"]:
lines += ["", "## v4 band-share grid (watch 20 / elevated 50)", "",
"| breaking | stable | watch | elevated | breaking |", "|---|---|---|---|---|"]
for row in report["variants"]["v4"]["band_grid"]:
w, e, b = row["bands"]
if w == 20.0 and e == 50.0:
sh = row["shares"]
lines.append(f"| {b:.0f} | {sh['stable']} | {sh['watch']} | {sh['elevated']} | {sh['breaking']} |")
lines += ["", "## Scenarios (pillar arithmetic, explicit sensor scores)", "",
"| scenario | price | breadth | C1 | V1 | State |", "|---|---|---|---|---|---|"]
for sc in report["scenarios"]["vix_a"]:
lines.append(f"| {sc['scenario']} | {sc['price']} | {sc['breadth']} | {sc['C1']} | {sc['V1']} | **{sc['state']}** |")
lines += ["", f"Recommendation: `{report['v4_recommendation']}`", ""]
return "\n".join(lines)
def _band_grid(states: list[float]) -> list[dict]:
grid = []
for watch in (15.0, 20.0, 25.0):
for elevated in (40.0, 50.0):
for breaking in (60.0, 65.0, 70.0):
grid.append({
"bands": [watch, elevated, breaking],
"shares": _band_shares(states, (watch, elevated, breaking)),
})
return grid
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def _parse_args() -> argparse.Namespace:
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--end", default=PUBLISHED["window_end"])
p.add_argument("--sessions", type=int, default=PUBLISHED["sessions"])
p.add_argument("--history-days", type=int, default=1200, help="matches production _fetch_prices")
p.add_argument("--oas-window-days", type=int, default=rms.HY_OAS_WINDOW_DAYS,
help="per-session slice for the canonical run")
p.add_argument("--oas-fetch-days", type=int, default=int(365.25 * 13),
help="fetch range; matches v2's request (ICE truncates to ~3y)")
p.add_argument("--methodology", default=",".join(REQUIRED_VARIANTS))
p.add_argument("--expected-first-session", default=None,
help="assert the first replayed date. Defaults to the published "
"window's start; REQUIRED when --end/--sessions are overridden, "
"since the count alone is tautological.")
p.add_argument("--cache-dir", default=None)
p.add_argument("--out", default=None)
p.add_argument("--quiet", action="store_true")
return p.parse_args()
def _git_rev() -> str:
try:
return subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=ROOT,
capture_output=True, text=True, check=True).stdout.strip()
except Exception:
return "unknown"
def _source_state() -> dict:
"""Identify the code that produced this run, not just the commit HEAD names.
A run from a dirty tree is not reproducible by checking out git_rev -- which
is exactly how the first v4 artifact was generated, with HEAD still on the
harness commit while the v4 sensors lived only in the working tree. The
hashes make that visible instead of implied.
"""
import hashlib
tracked = [
"app/services/regime_monitor_service.py",
"app/services/breadth_service.py",
"scripts/run_regime_monitor_calibration.py",
]
digests = {}
for rel in tracked:
path = ROOT / rel
digests[rel] = hashlib.sha256(path.read_bytes()).hexdigest()[:16] if path.exists() else None
try:
dirty = bool(subprocess.run(["git", "status", "--porcelain"], cwd=ROOT,
capture_output=True, text=True, check=True).stdout.strip())
except Exception:
dirty = None
return {"git_rev": _git_rev(), "git_dirty": dirty, "source_sha256": digests}
async def _main() -> int:
args = _parse_args()
end = date.fromisoformat(args.end)
cache_dir = Path(args.cache_dir) if args.cache_dir else None
config = deepcopy(rms.DEFAULT_CONFIG)
symbols = _all_symbols(config)
variants = [v.strip() for v in args.methodology.split(",") if v.strip()]
expected_first = args.expected_first_session
if not expected_first:
if args.end != PUBLISHED["window_end"] or args.sessions != PUBLISHED["sessions"]:
print("--expected-first-session is required when --end or --sessions "
"differ from the published window.", file=sys.stderr)
return 2
expected_first = PUBLISHED["window_first"]
unknown = [v for v in variants if v not in VARIANTS]
if unknown:
print(f"unknown variant(s): {unknown}; known: {sorted(VARIANTS)}", file=sys.stderr)
return 2
missing = [v for v in REQUIRED_VARIANTS if v not in variants]
if missing:
print(f"--methodology must include {list(REQUIRED_VARIANTS)}; missing {missing}. "
"v2_reconstruction carries the published reproduction figures, and "
"v3+v4 are needed for the state_v4 <= state_v3 invariant gate.",
file=sys.stderr)
return 2
if not args.quiet:
print(f"fetching {len(symbols)} symbols from Alpaca...", flush=True)
prices = await _load_prices(symbols, end - timedelta(days=args.history_days), end, cache_dir, args.quiet)
vix = await _load_fred("VIXCLS", end - timedelta(days=args.history_days), end, cache_dir)
oas = await _load_fred("BAMLH0A0HYM2", end - timedelta(days=args.oas_fetch_days), end, cache_dir)
leader = prices.get(LEADER, [])
if not leader:
print("no leader (SMH) price data — cannot replay", file=sys.stderr)
return 2
sessions = [d for d, _ in leader if d <= end][-args.sessions:]
breadth, breadth_counts = breadth_service._breadth_with_counts(
{s: prices[s] for s in _basket(config) if prices.get(s)}, window=200, min_tickers=20
)
# Required glue: _item_asof breaks on the first date > as_of, so unsorted
# input silently returns a wrong value rather than erroring.
breadth_series = rms._mapping_series(breadth)
divergence_by_variant = {
"live": rms._mapping_series(breadth_service.compute_divergence_series(breadth, leader)),
"v2": rms._mapping_series(_v2_divergence_series(breadth, leader)),
}
if args.oas_window_days != rms.HY_OAS_WINDOW_DAYS:
rms.HY_OAS_WINDOW_DAYS = args.oas_window_days
results: dict[str, Any] = {}
rows_by_variant: dict[str, list[dict]] = {}
for variant in variants:
if not args.quiet:
print(f"replaying {variant} over {len(sessions)} sessions...", flush=True)
rows = _replay(variant, prices, vix, oas, config, sessions,
breadth_series, divergence_by_variant, breadth_counts)
rows_by_variant[variant] = rows
results[variant] = _stats(rows, variant)
results[variant]["soft_gates"] = _soft_gates(results[variant], variant) \
if variant in ("v3", "v2_reconstruction") else []
states = [r["state"]["score"] for r in rows if r["state"]["score"] is not None]
if variant.startswith("v4"):
results[variant]["band_grid"] = _band_grid(states)
canonical = rows_by_variant.get("v3") or next(iter(rows_by_variant.values()))
hard_gates = _completeness_gates(prices, symbols, sessions, breadth_counts, len(_basket(config)))
hard_gates += _pipeline_gates(canonical, sessions, expected_first)
if "v3" in rows_by_variant and "v4" in rows_by_variant:
hard_gates.append(_invariant_gate(rows_by_variant["v3"], rows_by_variant["v4"]))
blocked_by = [g["gate"] for g in hard_gates if not g["passed"]]
provisional = any(
not g["passed"] for v in results.values() for g in v.get("soft_gates", [])
)
report = {
"generated_at": datetime.now().isoformat(timespec="seconds"),
"source": _source_state(),
"params": vars(args),
"provenance": {
"symbols": {s: {"bars": len(prices.get(s, [])),
"first": prices[s][0][0].isoformat() if prices.get(s) else None,
"last": prices[s][-1][0].isoformat() if prices.get(s) else None}
for s in symbols},
"vix": {"points": len(vix or []),
"first": vix[0][0].isoformat() if vix else None,
"last": vix[-1][0].isoformat() if vix else None},
"oas": {"points": len(oas or []),
"first": oas[0][0].isoformat() if oas else None,
"last": oas[-1][0].isoformat() if oas else None},
"scored_range": [sessions[0].isoformat(), sessions[-1].isoformat()],
},
"hard_gates": hard_gates,
"blocked_by": blocked_by,
"provisional": provisional,
"variants": results,
"scenarios": {
"note": "pillar arithmetic on explicit sensor scores; State weights unchanged",
"vix_a": _scenarios(P5_VIX_ANCHORS_A, P1_TREND_BREAK_ANCHORS),
},
"diagnostics": {
"vix_top10": sorted(((d.isoformat(), v) for d, v in (vix or [])),
key=lambda x: -x[1])[:10],
"candidate_anchors": {
"P5_VIX_ANCHORS_A": P5_VIX_ANCHORS_A,
"P5_VIX_ANCHORS_B": P5_VIX_ANCHORS_B,
"P1_TREND_BREAK_ANCHORS": P1_TREND_BREAK_ANCHORS,
},
},
# Structurally impossible to read a recommendation out of a run whose
# pipeline did not validate.
"v4_recommendation": None if blocked_by else {
"state_bands_candidate": [20.0, 50.0, 65.0],
"provisional": provisional,
"note": "confirm against band_grid + scenarios before shipping",
},
}
stamp = datetime.now().strftime("%Y%m%d-%H%M%S")
out = Path(args.out) if args.out else ROOT / "reports" / f"regime-monitor-v4-calibration-{stamp}.json"
out.parent.mkdir(parents=True, exist_ok=True)
tmp = out.with_suffix(out.suffix + ".tmp")
tmp.write_text(json.dumps(report, indent=2, default=str), encoding="utf-8")
tmp.replace(out)
md = out.with_suffix(".md")
md_tmp = md.with_suffix(".md.tmp")
md_tmp.write_text(_markdown(report), encoding="utf-8")
md_tmp.replace(md)
if not args.quiet:
print(f"wrote {out}", flush=True)
for gate in hard_gates:
mark = "ok " if gate["passed"] else "FAIL"
print(f" [{mark}] {gate['gate']}: expected {gate['expected']}, got {gate['measured']}")
if blocked_by:
print(f"HARD GATES FAILED: {blocked_by} — no v4 recommendation emitted", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(asyncio.run(_main()))
+110
View File
@@ -0,0 +1,110 @@
"""Admin → Jobs listing: categories, ordering, and next-run coherence.
The panel used to render 19 jobs as one alphabetical list in which a pipeline
step, a cron job and a manual job were indistinguishable, and a triggered job
could advertise a next run ten years out.
"""
from datetime import datetime, timedelta, timezone
import pytest
from app import job_catalog
from app.scheduler import configure_scheduler, scheduler
from app.services.admin_service import _visible_next_run, list_jobs
@pytest.fixture(autouse=True)
def _configured_scheduler():
scheduler.remove_all_jobs()
configure_scheduler()
yield
scheduler.remove_all_jobs()
def _by_name(jobs: list[dict]) -> dict[str, dict]:
return {job["name"]: job for job in jobs}
class TestVisibleNextRun:
def test_parked_backstop_is_not_a_schedule(self):
"""Paused jobs carry a 520-week interval; triggering one re-arms it."""
backstop = datetime.now(timezone.utc) + timedelta(weeks=520)
assert _visible_next_run(backstop) is None
def test_a_real_upcoming_run_passes_through(self):
soon = datetime.now(timezone.utc) + timedelta(hours=6)
assert _visible_next_run(soon) == soon
def test_none_stays_none(self):
assert _visible_next_run(None) is None
class TestListJobs:
async def test_hidden_jobs_are_not_listed_but_stay_valid(self, db_session):
jobs = _by_name(await list_jobs(db_session))
assert "data_backfill" not in jobs
# Still triggerable through the API, and still registered.
assert "data_backfill" in job_catalog.VALID_JOB_NAMES
assert scheduler.get_job("data_backfill") is not None
async def test_every_visible_job_has_a_category(self, db_session):
jobs = await list_jobs(db_session)
assert {j["name"] for j in jobs} == set(
job_catalog.VALID_JOB_NAMES - job_catalog.HIDDEN_JOBS
)
assert all(j["category"] in job_catalog.CATEGORY_ORDER for j in jobs)
async def test_jobs_arrive_grouped_by_category(self, db_session):
"""The frontend renders sections in payload order, so ordering is the
API's job — not something each client re-derives."""
categories = [j["category"] for j in await list_jobs(db_session)]
ranks = [job_catalog.CATEGORY_ORDER.index(c) for c in categories]
assert ranks == sorted(ranks)
async def test_pipeline_steps_defer_their_schedule_to_the_parent(self, db_session):
jobs = _by_name(await list_jobs(db_session))
step = jobs["rr_scanner"]
assert step["category"] == job_catalog.CATEGORY_STEP
assert step["next_run_at"] is None
assert step["next_run_source"] == "via_pipeline"
assert step["pipelines"] == ["near_close_pipeline"]
async def test_step_reports_the_soonest_enabled_parent(self, db_session):
due = datetime.now(timezone.utc) + timedelta(hours=3)
scheduler.modify_job("daily_pipeline", next_run_time=due)
collector = _by_name(await list_jobs(db_session))["data_collector"]
assert collector["via_next_run_job"] == "daily_pipeline"
assert collector["via_next_run_at"] == due.isoformat()
# Runs in all four pipelines — the reason steps are not nested under one.
assert set(collector["pipelines"]) == set(job_catalog.PIPELINE_JOBS)
async def test_manual_jobs_say_so_instead_of_showing_a_date(self, db_session):
study = _by_name(await list_jobs(db_session))["event_study"]
assert study["category"] == job_catalog.CATEGORY_MANUAL
assert study["next_run_source"] == "manual_only"
assert study["next_run_at"] is None
async def test_a_triggered_manual_job_still_shows_no_next_run(self, db_session):
"""Regression: triggering re-armed the 520-week backstop, which the panel
rendered as a real 'next run in ~87600h'."""
scheduler.modify_job("event_study", next_run_time=datetime.now(timezone.utc))
scheduler.modify_job("event_study", next_run_time=None)
study = _by_name(await list_jobs(db_session))["event_study"]
assert study["next_run_at"] is None
async def test_pipelines_report_their_own_schedule_and_steps(self, db_session):
pipeline = _by_name(await list_jobs(db_session))["daily_pipeline"]
assert pipeline["category"] == job_catalog.CATEGORY_PIPELINE
assert pipeline["next_run_source"] == "own_schedule"
assert pipeline["steps"] == [
step for step, _ in job_catalog.PIPELINE_STEPS["daily_pipeline"]
]
async def test_standalone_jobs_keep_their_own_schedule(self, db_session):
backtest = _by_name(await list_jobs(db_session))["backtest"]
assert backtest["category"] == job_catalog.CATEGORY_SCHEDULED
assert backtest["next_run_source"] == "own_schedule"
assert backtest["pipelines"] == []
+306 -8
View File
@@ -2,6 +2,7 @@
from __future__ import annotations
import json
import math
from datetime import date, timedelta
from types import SimpleNamespace
@@ -1272,32 +1273,63 @@ def test_build_recommendation_reads_the_report():
{"min_momentum_percentile": 60.0, "net_avg_r": 0.05, "total": 300},
{"min_momentum_percentile": 0.0, "net_avg_r": -0.12, "total": 1000},
],
# Legacy policy book. Its numbers are deliberately DIFFERENT from the
# production monitor's below, so sourcing the benchmark line from here
# again would fail the assertion rather than pass unnoticed.
"portfolio_sim": {"policies": [
{"policy": "target", "cagr_pct": 23.7, "total_return_pct": 134.8,
"spy_return_pct": 95.9, "max_drawdown_pct": 20.7},
{"policy": "hold", "cagr_pct": 31.9, "total_return_pct": 203.6,
"spy_return_pct": 95.9, "max_drawdown_pct": 21.2},
]},
"portfolio_monitor": {
"production_strategy": "prod",
"runs": [{
"strategy": "prod", "lookback": "all", "lookback_label": "All history",
"cagr_pct": 40.0, "sharpe": 1.72, "max_drawdown_pct": 17.7,
"total_return_pct": 297.8, "spy_return_pct": 101.9,
}, {
# A second window with DIFFERENT numbers. Without it the "all"
# preference is untested and a lookback mix-up cannot fail.
"strategy": "prod", "lookback": "3y", "lookback_label": "3y",
"cagr_pct": 47.9, "sharpe": 1.96, "max_drawdown_pct": 17.3,
"total_return_pct": 220.9, "spy_return_pct": 71.5,
}],
},
}
rec = bt._build_recommendation(report)
by_topic: dict[str, list[str]] = {}
for item in rec["items"]:
by_topic.setdefault(item["topic"], []).append(item["text"])
assert rec["headline"] is not None and "hold 30" in rec["headline"]
assert any("hold 30 trading days" in t for t in by_topic["exit"])
assert rec["headline"] is not None and "Production baseline" in rec["headline"]
# The hold-vs-target comparison is gone: both are exits the production book
# replaced, so a recommendation between them cannot lead to an action.
assert "exit" not in by_topic
# Benchmark must quote the SAME row the page's tiles show, not the policy sim.
assert "+297.8%" in by_topic["benchmark"][0]
assert "203.6" not in by_topic["benchmark"][0]
gate_texts = " | ".join(by_topic["gate"])
assert "confidence floor adds nothing" in gate_texts
assert "keep the R:R floor" in gate_texts
assert "keep the NEUTRAL exclusion" in gate_texts
assert "80" in by_topic["cutoff"][0]
assert "beats" in by_topic["benchmark"][0]
# robustness is judged under the RECOMMENDED exit (the 30d hold), not the
# target model the recommendation advises abandoning
assert any(
"not a handful of outliers" in t and "under the recommended 30d hold" in t
for t in by_topic["robustness"]
)
# Every production figure comes from ONE window, and the report says which,
# so the page can default its selector to the same one.
assert rec["basis_lookback"] == "all"
assert rec["basis_lookback_label"] == "All history"
assert "+40.0%" in by_topic["production"][0]
assert "47.9" not in by_topic["production"][0] # the 3y row must not leak in
assert "220.9" not in by_topic["benchmark"][0]
# Robustness names its real basis. It used to claim "under the recommended
# 30d hold" — nothing recommends that exit; production is the ATR trail.
robustness = by_topic["robustness"][0]
assert "not a handful of outliers" in robustness
assert "gate-level grading" in robustness
assert "recommended" not in robustness
def test_build_recommendation_flags_outlier_dependence():
@@ -1620,3 +1652,269 @@ async def test_run_backtest_smoke(session):
sweep = sorted(report["sweep"], key=lambda r: r["min_momentum_percentile"], reverse=True)
counts = [r["total"] for r in sweep]
assert counts == sorted(counts) # ascending as threshold descends
async def test_run_backtest_rolls_back_a_failed_ticker_fetch(session, monkeypatch):
"""A failed per-ticker read must not leave the session mid-failed-transaction.
Every DB call in the replay loop is best-effort, but swallowing the error
without a rollback leaves asyncpg in "current transaction is aborted": every
later statement fails the same way until the first unguarded one the report
write surfaces it as the job error, long after the real cause.
"""
await _seed_oscillating_ticker(session, "AAA")
await _seed_oscillating_ticker(session, "OSC")
real_fetch = bt._fetch_columns
rolled_back: list[str] = []
async def failing_fetch(db, symbol):
if symbol == "AAA":
raise RuntimeError("simulated OHLCV read failure")
return await real_fetch(db, symbol)
real_rollback = session.rollback
async def tracking_rollback():
rolled_back.append("x")
await real_rollback()
monkeypatch.setattr(bt, "_fetch_columns", failing_fetch)
monkeypatch.setattr(session, "rollback", tracking_rollback)
report = await bt.run_backtest(session)
assert rolled_back, "a failed ticker fetch left the session un-rolled-back"
# the surviving ticker is still replayed after the rollback
assert report["tickers"] == 2
assert report["candidates"] >= 1
async def test_run_backtest_rolls_back_a_failed_portfolio_sim_load(session, monkeypatch):
"""The portfolio-sim block loads the benchmark and the live exit policy from
the same session, well after the replay loop. A failure there poisons the
transaction exactly as one in the loop does, and the report write pays for it.
"""
await _seed_oscillating_ticker(session, "OSC")
rolled_back: list[str] = []
called: list[str] = []
async def failing_exit_policy(db):
called.append("x")
raise RuntimeError("simulated exit-policy read failure")
real_rollback = session.rollback
async def tracking_rollback():
rolled_back.append("x")
await real_rollback()
monkeypatch.setattr(
"app.services.paper_trade_service.get_exit_policy", failing_exit_policy
)
monkeypatch.setattr(session, "rollback", tracking_rollback)
report = await bt.run_backtest(session)
assert called, "the portfolio-sim block never ran; test proves nothing"
assert rolled_back, "a failed portfolio-sim load left the session un-rolled-back"
assert report["tickers"] == 1
class TestPortfolioQualityMetrics:
"""Sortino / Gain-to-Pain / dollar profit factor.
Each derives its expectation from the returned ``equity_curve`` rather than
hand-tracing position sizing, and each also asserts the *wrong* variant is
NOT what came back the denominator and the numerator are exactly where
these ratios are usually got wrong.
"""
ORD = date(2025, 1, 6).toordinal()
@staticmethod
def _daily_returns(sim: dict) -> list[float]:
eq = [row["equity"] for row in sim["equity_curve"]]
return [b / a - 1.0 for a, b in zip(eq, eq[1:]) if a > 0]
@staticmethod
def _monthly_returns(sim: dict) -> list[float]:
monthly: list[float] = []
rows = sim["equity_curve"]
start = last = rows[0]["equity"]
cur = date.fromisoformat(rows[0]["date"]).replace(day=1)
for row in rows:
m = date.fromisoformat(row["date"]).replace(day=1)
if m != cur:
monthly.append(last / start - 1.0)
cur, start = m, last
last = row["equity"]
monthly.append(last / start - 1.0)
return monthly
def _wobbly_sim(self) -> dict:
"""~70 sessions crossing four month boundaries with a real mid drawdown,
so monthly returns include both signs (a short fixture yields one month
and zero pain, which reads as a broken formula)."""
closes = (
[100.0 + i for i in range(20)] # climb
+ [120.0 - 1.5 * i for i in range(20)] # drawdown
+ [90.0 + 1.2 * i for i in range(30)] # recovery
)
prices = {"AAA": _sim_prices(self.ORD, closes)}
cand = _sim_cand("AAA", self.ORD, entry=100.0, stop=80.0, target=400.0)
sim = bt._simulate_portfolio(
[cand], prices, None, "hold", 65, include_curve=True
)
assert sim is not None
return sim
def test_sortino_denominator_is_full_sample_not_downside_count(self):
sim = self._wobbly_sim()
rets = self._daily_returns(sim)
downside = [r for r in rets if r < 0.0]
assert downside, "fixture must produce down days or the test proves nothing"
mean_ret = sum(rets) / len(rets)
correct = mean_ret / math.sqrt(
sum(r * r for r in downside) / len(rets)
) * math.sqrt(252.0)
# The classic error: dividing by the count of down days shrinks the
# denominator and inflates the ratio.
inflated = mean_ret / math.sqrt(
sum(r * r for r in downside) / len(downside)
) * math.sqrt(252.0)
assert sim["sortino"] == pytest.approx(round(correct, 2), abs=0.01)
assert sim["sortino"] != pytest.approx(round(inflated, 2), abs=0.01)
def test_gain_to_pain_is_schwager_on_monthly_returns(self):
sim = self._wobbly_sim()
monthly = self._monthly_returns(sim)
assert len(monthly) >= 3, "fixture must span several months"
pain = -sum(r for r in monthly if r < 0.0)
assert pain > 0, "fixture must have a losing month or pain is zero"
schwager = sum(monthly) / pain
# sum(all) = sum(pos) - |sum(neg)|, so the profit-factor-shaped variant
# sits exactly 1.0 higher for every input.
profit_factor_shaped = sum(r for r in monthly if r > 0.0) / pain
assert profit_factor_shaped == pytest.approx(schwager + 1.0, abs=1e-9)
assert sim["gain_to_pain"] == pytest.approx(round(schwager, 2), abs=0.01)
assert sim["gain_to_pain"] != pytest.approx(
round(profit_factor_shaped, 2), abs=0.01
)
def test_profit_factor_is_dollar_based(self):
"""One winner, one loser, on separate symbols so both fill."""
up = [100.0 + 2.0 * i for i in range(8)]
down = [100.0 - 2.0 * i for i in range(8)]
prices = {
"WIN": _sim_prices(self.ORD, up),
"LOSE": _sim_prices(self.ORD, down),
}
cands = [
_sim_cand("WIN", self.ORD, entry=100.0, stop=90.0, target=400.0, mp=95.0),
_sim_cand("LOSE", self.ORD, entry=100.0, stop=80.0, target=400.0, mp=94.0),
]
sim = bt._simulate_portfolio([cands[0], cands[1]], prices, None, "hold", 5)
assert sim is not None
assert sim["trades"] == 2
# With exactly two trades the reported best/worst ARE the win and the loss.
gross_win = sim["best_trade_pnl"]
gross_loss = -sim["worst_trade_pnl"]
assert gross_win > 0 and gross_loss > 0, "fixture must produce one of each"
assert sim["profit_factor"] == pytest.approx(
round(gross_win / gross_loss, 2), abs=0.01
)
def test_keys_always_present_and_no_downside_is_none(self):
"""Monotonic rise: no down days. Sortino must be None, never inf — and
all three keys must still be emitted, because the UI reads an ABSENT key
as 'report predates these metrics'."""
closes = [100.0, 102.0, 104.0, 106.0, 108.0, 110.0]
prices = {"AAA": _sim_prices(self.ORD, closes)}
cand = _sim_cand("AAA", self.ORD, entry=100.0, stop=95.0, target=130.0)
sim = bt._simulate_portfolio([cand], prices, None, "hold", 3)
assert sim is not None
for key in ("sortino", "gain_to_pain", "profit_factor"):
assert key in sim
assert sim["sortino"] is None
def test_build_recommendation_states_no_baseline_without_a_production_row():
"""A report with no portfolio monitor cannot describe the production book.
It used to fall back to recommending the fixed-hold exit advice for a model
the production book had already replaced."""
report = {
"overall_qualified": {"net_avg_r": 0.13, "net_avg_r_ex_top5": 0.05},
"time_exit_sweep": [{"hold_days": 30, "net_avg_r": 0.50, "net_avg_r_ex_top5": 0.21}],
"portfolio_sim": {"policies": [
{"policy": "hold", "cagr_pct": 31.9, "total_return_pct": 203.6,
"spy_return_pct": 95.9, "max_drawdown_pct": 21.2},
]},
}
rec = bt._build_recommendation(report)
topics = {item["topic"] for item in rec["items"]}
assert rec["headline"] is None
# Nothing may be sourced from the legacy policy book.
assert "benchmark" not in topics
assert "exit" not in topics
async def test_cached_report_recommendation_is_rebuilt_on_read(session):
"""A report cached by an older build carries that build's recommendation.
Served verbatim, the page would show the legacy wording and no
basis_lookback which let the lookback selector default elsewhere, putting
3y tiles beside an all-history recommendation with no warning. This is the
shape of the report sitting in production right now.
"""
from app.services.admin_service import update_setting
stale = {
"generated_at": "2026-08-12T05:00:00+00:00",
"tickers": 512, "candidates": 100, "qualified": 10,
"params": {"horizon_days": 30},
"overall_qualified": {"net_avg_r": 0.13, "net_avg_r_ex_top5": 0.20},
"portfolio_sim": {"policies": [
{"policy": "hold", "cagr_pct": 31.9, "total_return_pct": 175.0,
"spy_return_pct": 101.9, "max_drawdown_pct": 23.7},
]},
"portfolio_monitor": {
"production_strategy": "prod",
"runs": [{
"strategy": "prod", "lookback": "all", "lookback_label": "All history",
"cagr_pct": 40.0, "sharpe": 1.72, "max_drawdown_pct": 17.7,
"total_return_pct": 297.8, "spy_return_pct": 101.9,
}],
},
# What the old build stored: sourced from the policy book, and naming an
# exit the production book replaced.
"recommendation": {
"headline": "Trade the qualified list long-only; hold 30 trading days.",
"items": [
{"topic": "benchmark", "text": "Book vs SPY: beats buy-and-hold by "
"+73.1 points (+175.0% vs +101.9%)."},
{"topic": "robustness", "text": "Robustness: expectancy survives removing "
"the top 5% of winners (+0.20R net/trade "
"under the recommended 30d hold)."},
],
"note": "stale",
},
}
await update_setting(session, bt.KEY_REPORT, json.dumps(stale))
report = await bt.get_backtest_report(session)
assert report is not None
rec = report["recommendation"]
# Rebuilt: the basis is published, so the page cannot default elsewhere.
assert rec["basis_lookback"] == "all"
texts = " | ".join(i["text"] for i in rec["items"])
# ...and it quotes the production book, not the policy sim it used to.
assert "+297.8%" in texts and "175.0" not in texts
assert "recommended 30d hold" not in texts
assert "Production baseline" in (rec["headline"] or "")
+396 -1
View File
@@ -1,17 +1,25 @@
"""Tests for v3 correction events, warning alarm episodes, and report caveats."""
"""Tests for correction events, alarm episodes, the shipped-rule replay, and caveats."""
from __future__ import annotations
from copy import deepcopy
from datetime import date, timedelta
from app.services.breadth_service import _breadth_from_closes, compute_divergence_series
from app.services.event_study_service import (
MIN_EVENTS_FOR_CONFIDENCE,
STRESS_QUADRANT,
WARNING_QUADRANTS,
_era_split,
_null_model,
_percentile,
_reliability,
alarm_episodes,
below_average_series,
detect_events,
entry_alarms,
evaluate_alarms,
replay_quadrant_changes,
)
@@ -19,6 +27,33 @@ def _days(count: int, start: date = date(2021, 1, 1)) -> list[date]:
return [start + timedelta(days=index) for index in range(count)]
def _row(
warning: float,
state: float = 0.0,
*,
warning_coverage: float = 100.0,
state_coverage: float = 100.0,
fresh: bool = True,
) -> dict:
return {
"state": state,
"warning": warning,
"state_coverage": state_coverage,
"warning_coverage": warning_coverage,
"inputs_fresh": fresh,
}
def _rows(
dates: list[date], warnings: list[float], patch: dict[int, dict] | None = None
) -> dict[date, dict]:
"""One publishable row per date, with per-position replacements."""
built = {day: _row(value) for day, value in zip(dates, warnings)}
for index, replacement in (patch or {}).items():
built[dates[index]] = replacement
return built
def test_detect_events_uses_rising_edge_and_cooldown():
closes = [100.0] * 300 + [85.0] * 5 + [100.0] * 50 + [85.0] * 5
events = detect_events(closes, _days(len(closes)), threshold_pct=15.0, cooldown=40)
@@ -88,6 +123,366 @@ def test_evaluate_alarms_counts_episodes_not_alarm_days():
assert result["median_lead_days"] == 17.5
# ---------------------------------------------------------------------------
# The shipped quadrant rule, replayed
# ---------------------------------------------------------------------------
def test_replay_seeds_silently_and_needs_two_sessions():
"""A one-session spike is not an alert; the second session confirms it.
The alarm is therefore dated at the confirmation rather than at the first
crossing, which costs one session of lead. That is what ships.
"""
dates = _days(10)
spike = _rows(dates, [30] * 5 + [70] + [30] * 4)
assert replay_quadrant_changes(spike, dates) == []
held = _rows(dates, [30] * 5 + [70, 70] + [30] * 3)
fires = replay_quadrant_changes(held, dates)
# The rule alerts on quadrant changes in both directions, so the return to
# calm fires too. Only the entry is a warning about anything.
assert [(f["index"], f["from"], f["to"]) for f in fires] == [
(6, "3", "1"),
(9, "1", "3"),
]
assert entry_alarms(fires, WARNING_QUADRANTS) == [6]
def test_confirmation_classifies_the_prior_session_against_the_baseline():
"""Not against its own predecessor -- the distinction changes the answer.
Warning 42 sits inside the hysteresis deadband. Measured from the standing
"3" baseline it is still "3", so it cannot confirm a move to "1". A chain
that classified each session against the one before it would read 42 as "1"
(having just seen 70) and fire a day later, which production does not do.
"""
dates = _days(10)
rows = _rows(dates, [30, 30, 30, 30, 70, 42, 70, 30, 30, 30])
assert replay_quadrant_changes(rows, dates) == []
def test_cooldown_suppresses_and_the_baseline_only_advances_on_a_fire():
dates = _days(10)
rows = _rows(dates, [30, 30, 30, 30, 70, 70, 30, 30, 30, 30])
fires = replay_quadrant_changes(rows, dates)
# Entry confirmed on day 5. The exit confirms on day 7 but lands inside the
# 3-day cooldown, so it is re-evaluated and fires on day 8 instead.
assert [(f["index"], f["from"], f["to"]) for f in fires] == [
(5, "3", "1"),
(8, "1", "3"),
]
assert entry_alarms(fires, WARNING_QUADRANTS) == [5]
def test_low_coverage_sessions_cannot_confirm():
"""The confirmation source has to be a session that published a band."""
dates = _days(10)
warnings = [30, 30, 30, 30, 30, 70, 70, 30, 30, 30]
visible = replay_quadrant_changes(_rows(dates, warnings), dates)
assert entry_alarms(visible, WARNING_QUADRANTS) == [6]
# Day 5 is the only session that could confirm the entry on day 6; below
# MIN_COVERAGE it never published a band, so day 4 is the prior instead.
hidden = _rows(dates, warnings, {5: _row(70, warning_coverage=70.0)})
assert replay_quadrant_changes(hidden, dates) == []
def test_stale_inputs_block_todays_alert_but_not_tomorrows_confirmation():
"""is_fresh gates the live reading only; the prior session comes from history."""
dates = _days(10)
rows = _rows(dates, [30] * 4 + [70, 70, 70] + [30] * 3, {5: _row(70, fresh=False)})
fires = replay_quadrant_changes(rows, dates)
assert entry_alarms(fires, WARNING_QUADRANTS) == [6]
def test_entry_alarms_ignore_movement_inside_the_set():
fires = [
{"index": 3, "from": "3", "to": "1"},
{"index": 9, "from": "1", "to": "2"},
{"index": 20, "from": "2", "to": "4"},
]
assert entry_alarms(fires, WARNING_QUADRANTS) == [3]
assert entry_alarms(fires, STRESS_QUADRANT) == [9]
def test_below_average_series_needs_a_full_window():
series = list(zip(_days(6), [10.0, 10.0, 10.0, 10.0, 4.0, 20.0]))
indicator = below_average_series(series, window=3)
assert _days(6)[1] not in indicator # warm-up
assert indicator[_days(6)[4]] == 100.0 # 4 is under the 3-day mean of 8
assert indicator[_days(6)[5]] == 0.0
def test_null_model_is_seeded_and_drawn_from_evaluable_sessions_only():
dates = _days(300)
events = [100, 180, 260]
first = _null_model(6, events, dates, horizon=20, start_index=50, observed_warned=2, draws=200)
second = _null_model(6, events, dates, horizon=20, start_index=50, observed_warned=2, draws=200)
assert first == second # a re-run must not move the report
assert 0.0 <= first["p_at_least_observed"] <= 1.0
assert first["alarms_per_draw"] == 6
assert first["mean_warned"] <= len(events)
# More alarms than there are sessions to place them on is not a null.
assert _null_model(500, events, dates, 20, 50, 2, draws=10) is None
assert _null_model(6, [], dates, 20, 50, 0, draws=10) is None
def test_era_split_reports_the_two_sensor_eras_separately():
"""The fuller sample is mostly pre-credit, where Warning is W1+W2 only."""
dates = _days(400)
eras = _era_split(
alarms=[80, 300],
event_indices=[90, 310],
dates=dates,
horizon=20,
start_index=10,
credit_from=dates[200],
)
assert eras["pre_credit"]["events"] == 1
assert eras["pre_credit"]["events_warned"] == 1
assert eras["full_coverage"]["events"] == 1
assert eras["full_coverage"]["events_warned"] == 1
assert eras["credit_from"] == dates[200].isoformat()
# No credit series at all means there is no boundary to split on.
assert _era_split([80], [90], dates, 20, 10, None) is None
def _business_days(count: int, end: date = date(2026, 8, 7)) -> list[date]:
out: list[date] = []
cursor = end
while len(out) < count:
if cursor.weekday() < 5:
out.append(cursor)
cursor -= timedelta(days=1)
return list(reversed(out))
def _synthetic_path(sessions: int) -> list[float]:
"""A rising leader with two deep drawdowns, so corrections exist to detect."""
closes: list[float] = []
for index in range(sessions):
if index < 350:
closes.append(100.0 + index * 0.25)
elif index < 400:
closes.append(187.5 - (index - 350) * 0.9)
elif index < 650:
closes.append(142.5 + (index - 400) * 0.4)
elif index < 700:
closes.append(242.5 - (index - 650) * 1.1)
else:
closes.append(187.5 + (index - 700) * 0.3)
return closes
async def test_report_assembles_every_rule_from_synthetic_inputs(monkeypatch):
"""End-to-end: the shipped replay, ablations, baselines and null all score.
Synthetic rather than recorded because the point is the wiring -- that every
rule is measured on the same events over the same sessions and the report
carries what the panel reads. The numbers are meaningless by construction.
"""
import app.services.event_study_service as ess
sessions = 900
dates = _business_days(sessions)
closes = _synthetic_path(sessions)
leader = list(zip(dates, closes))
# SPY grinds up throughout, so the leader's relative strength rolls over
# exactly when it falls.
market = list(zip(dates, [100.0 + index * 0.12 for index in range(sessions)]))
# Breadth deteriorates ~15 sessions ahead of each decline, which is the
# divergence W1 exists to catch.
breadth = {}
for index, day in enumerate(dates):
weak = 335 <= index < 400 or 635 <= index < 700
breadth[day] = 30.0 if weak else 70.0
vix = [(day, 32.0 if (350 <= i < 400 or 650 <= i < 700) else 15.0) for i, day in enumerate(dates)]
# Credit starts late, exactly as ICE's 3-year cap makes it in production.
oas = [(day, 4.2 if (650 <= i < 700) else 3.0) for i, day in enumerate(dates) if i >= 500]
async def fake_config(_db):
return deepcopy(ess.rms.DEFAULT_CONFIG)
async def fake_prices(_config, _start, _end):
return {"SMH": leader, "QQQ": leader, "SPY": market}
async def fake_fred(series_id, _start, _end):
return {"VIXCLS": vix, "BAMLH0A0HYM2": oas}.get(series_id)
async def fake_breadth(_db, _symbols, window=200, min_tickers=20):
return breadth, {day: 30 for day in dates}
async def fake_observations(_db):
return []
monkeypatch.setattr(ess.rms, "get_regime_config", fake_config)
monkeypatch.setattr(ess.rms, "_fetch_prices", fake_prices)
monkeypatch.setattr(ess.rms, "_fetch_fred_series", fake_fred)
monkeypatch.setattr(ess.rms, "get_fundamental_observations", fake_observations)
monkeypatch.setattr(ess.breadth_service, "compute_breadth_details", fake_breadth)
monkeypatch.setattr(ess, "NULL_DRAWS", 100)
report = await ess.run_event_study(None)
assert report["available"] is True
assert report["schema"] == ess.STUDY_SCHEMA
# The shipped rule is measured on the whole sample, not a 30% holdout.
shipped = report["shipped"]
assert shipped["metrics"]["events"] == report["sample"]["events_evaluable"]
assert report["sample"]["events_evaluable"] >= 2
assert shipped["metrics"]["events"] >= report["fitted"]["metrics"]["events"]
assert len(shipped["events"]) == shipped["metrics"]["events"]
assert {row["kind"] for row in report["comparison"]} == {
"ablation", "baseline", "fundamental",
}
# Market rows share the headline's events, or the table lies. Fundamental
# rows deliberately do not: they are coverage-matched to the sessions the
# channel actually existed on, which is a different (here empty) window.
for row in report["comparison"]:
if row["kind"] != "fundamental":
assert row["events"] == shipped["metrics"]["events"]
assert row["false_alarms_per_year"] >= 0
else:
# No eligible sessions means the rate is undefined, not zero. A
# tiny-divisor fallback here printed 5e9 alarms/year.
assert row["false_alarms_per_year"] is None
# The credit sensor starts mid-sample, so the era split must be populated.
eras = shipped["by_era"]
assert eras["credit_from"] == dates[500].isoformat()
assert eras["pre_credit"]["events"] + eras["full_coverage"]["events"] == shipped["metrics"]["events"]
if report["null_model"] is not None:
assert 0.0 <= report["null_model"]["p_at_least_observed"] <= 1.0
assert report["null_model"]["observed_warned"] == shipped["metrics"]["events_warned"]
# With an empty observation series the fundamental rows are *untested*, not
# failed, and the report has to carry that distinction or a 0/10 in the table
# reads as a measured result.
coverage = report["fundamental_coverage"]
assert coverage["observations"] == 0
assert coverage["sessions_eligible"] == 0
assert coverage["events_covered"] == 0
assert coverage["measurable"] is False
fundamental_rows = [r for r in report["comparison"] if r["kind"] == "fundamental"]
assert {r["id"] for r in fundamental_rows} == {
"fundamental_adverse", "confluence", "market_over_covered",
}
assert all(row["measurable"] is False for row in fundamental_rows)
# Coverage-matched denominators: with no exposure these rows must not claim
# to have been scored against the market rows' 10 corrections.
assert all(row["events"] == 0 for row in fundamental_rows)
# Market rows are unaffected: their inputs exist for the whole window.
assert all(
row["measurable"] is True
for row in report["comparison"]
if row["kind"] != "fundamental"
)
def test_fundamental_rows_are_scored_only_on_their_own_exposure():
"""One day of coverage must not render as 0/10.
A fundamental rule scores zero whether it is wrong or merely absent, so
scoring it against corrections it could never have seen manufactures a
failed result out of a thin one the same mistake the `measurable` flag
prevents for an empty table, arriving one observation later.
"""
import app.services.event_study_service as ess
dates = _days(300)
events = [50, 120, 200, 280]
# Context exists for a single stretch, covering only the 120 event's horizon.
rows = {
day: {
"fundamental_state": "adverse",
"fundamental_usable": 105 <= index <= 115,
}
for index, day in enumerate(dates)
}
covered = ess.covered_events(events, rows, dates, horizon=20)
assert covered == [120]
assert ess.eligible_sessions(rows, dates, start_index=0) == 11
# A stale stretch counts for nothing, however adverse it reads.
stale = {
day: {"fundamental_state": "adverse", "fundamental_usable": False}
for day in dates
}
assert ess.covered_events(events, stale, dates, horizon=20) == []
assert ess.eligible_sessions(stale, dates, start_index=0) == 0
assert ess.adverse_episodes(stale, dates, 0) == []
assert ess.confluence_episodes([120], stale, dates) == []
# And neither does a *fresh* observation that determined nothing. Repeated
# extraction failures would otherwise accumulate exposure until the rows
# flipped to a measurable 0/8 for a channel that never knew anything —
# the same tested-versus-unavailable confusion, arriving by a slower route.
empty = {
day: {"fundamental_state": "unknown", "fundamental_usable": False}
for day in dates
}
assert ess.covered_events(events, empty, dates, horizon=20) == []
assert ess.eligible_sessions(empty, dates, start_index=0) == 0
async def test_the_fundamental_channel_never_moves_the_warning_score():
"""The channel is compared, never fused. Warning must be identical either way.
A weighted modifier was built and reverted: with ~10 correction events and
almost no fundamental history any fusion weight is a policy preference
presented as a measurement.
"""
import app.services.event_study_service as ess
end = date(2026, 6, 26)
dates = _business_days(400, end)
rising = [(day, 100.0 + index * 0.2) for index, day in enumerate(dates)]
prices = {"SMH": rising, "QQQ": rising, "SPY": rising}
args = (prices, [(end, 20.0)], [(day, 4.0) for day in dates])
config = deepcopy(ess.rms.DEFAULT_CONFIG)
names = config["tickers"]["hyperscalers"]
tail = (rising, [(day, 20.0) for day in dates], dates, config)
def adverse(effective: date) -> list[dict]:
return [{
"effective_date": effective,
"f1_score": 100.0,
"f3_score": 100.0,
"capex": dict.fromkeys(names, "cutting"),
"good_news_stock_down": "yes",
"fetched_at": "2026-01-01T00:00:00+00:00",
}]
bare = ess._axis_rows(*args, *tail, None)
observed = ess._axis_rows(*args, *tail, adverse(dates[-20]))
latest, early = dates[-1], dates[-90]
assert observed[latest]["warning"] == bare[latest]["warning"]
assert observed[latest]["fundamental_state"] == "adverse"
assert bare[latest]["fundamental_state"] == "unknown"
# Sessions before the effective date stay unknown, so a rebuild cannot stamp
# today's reading onto history.
assert observed[early]["fundamental_state"] == "unknown"
# The confluence rule keeps only crossings the channel agrees with, and the
# fundamental rule fires on the transition into adverse -- both rising-edge,
# so both stay comparable with the market rows.
adverse_alarms = ess.adverse_episodes(observed, dates, 0)
assert [dates[i] for i in adverse_alarms] == [dates[-20]]
assert ess.adverse_episodes(bare, dates, 0) == []
assert ess.confluence_episodes([dates.index(early), dates.index(latest)], observed, dates) == [
dates.index(latest)
]
def test_breadth_from_fixed_closes_and_tapered_divergence():
dates = _days(10)
closes_by_symbol = {
+141 -1
View File
@@ -1,7 +1,7 @@
from __future__ import annotations
import json
from datetime import date, datetime, timezone
from datetime import date, datetime, timedelta, timezone
from app.models.data_import_run import DataImportRun
from app.models.fundamental_snapshot import FundamentalSnapshot
@@ -139,3 +139,143 @@ async def test_ticker_quality_explains_no_xbrl_block(db_session):
assert await fundamentals_quality_service.ticker_is_eligible(
db_session, ticker.id
) is False
def _escalated_gap(cik: str, *, escalated: bool = True) -> SecFilingGap:
first_seen = datetime.now(timezone.utc) - timedelta(days=24)
return SecFilingGap(
cik=cik,
accession=f"{cik}-STALE-Q",
form="10-Q",
index_date=(first_seen.date()),
reason="not_in_companyfacts",
first_seen_at=first_seen,
last_attempted_at=datetime.now(timezone.utc),
escalated_at=(
datetime.now(timezone.utc) - timedelta(days=10) if escalated else None
),
)
def _prior_quarter(cik: str, *, age_days: int) -> FundamentalSnapshot:
"""The issuer's last successfully ingested filing, older than the gap so it
cannot supersede it exactly the production shape of a stale companyfacts
file: Q1 stored, Q2 missing."""
filed = date.today() - timedelta(days=age_days)
return FundamentalSnapshot(
cik=cik,
accession=f"{cik}-PRIOR-Q",
form="10-Q",
filed_date=filed,
accepted_at=datetime.now(timezone.utc) - timedelta(days=age_days),
period_end=filed,
fiscal_year=filed.year,
fiscal_period="Q1",
)
async def test_escalated_gap_stops_blocking_when_fundamentals_are_recent(db_session):
ticker = Ticker(symbol="STALEFACTS", cik="0000000046")
db_session.add(ticker)
db_session.add(_escalated_gap(ticker.cik))
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
ticker.id
}
# The alert has run and the issuer still has last quarter to score on.
db_session.add(_prior_quarter(ticker.cik, age_days=120))
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == set()
# ...but the filing is still queued, so the importer keeps retrying it.
assert len(await fundamentals_quality_service.active_gaps(db_session)) == 1
async def test_escalated_gap_keeps_blocking_when_fundamentals_are_stale(db_session):
ticker = Ticker(symbol="NOTHINGFRESH", cik="0000000047")
db_session.add_all([
ticker,
_escalated_gap(ticker.cik),
_prior_quarter(
ticker.cik,
age_days=fundamentals_quality_service.GAP_GATE_RECENT_FILING_DAYS + 30,
),
])
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
ticker.id
}
async def test_unescalated_gap_still_blocks_alongside_an_escalated_one(db_session):
ticker = Ticker(symbol="TWOGAPS", cik="0000000048")
fresh = datetime.now(timezone.utc)
db_session.add_all([
ticker,
_escalated_gap(ticker.cik),
SecFilingGap(
cik=ticker.cik,
accession="TWOGAPS-FRESH-Q",
form="10-Q",
index_date=date.today(),
reason="not_in_companyfacts",
first_seen_at=fresh,
last_attempted_at=fresh,
),
_prior_quarter(ticker.cik, age_days=120),
])
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
ticker.id
}
async def test_summary_path_does_not_reblock_an_exempt_cik(db_session):
"""The run summary mirrors the same filings as the queue — it must honour the
same hand-off, or the bound is inert in production."""
ticker = Ticker(symbol="MIRRORED", cik="0000000049")
db_session.add(ticker)
db_session.add(_escalated_gap(ticker.cik))
db_session.add(_prior_quarter(ticker.cik, age_days=120))
await db_session.flush()
db_session.add(
DataImportRun(
source="sec_facts",
status="promoted",
validation_json=json.dumps({
"setup_blocked_ciks": [ticker.cik],
"missing_xbrl": [
{"cik": ticker.cik, "accession": f"{ticker.cik}-STALE-Q"}
],
}),
started_at=datetime.now(timezone.utc),
)
)
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == set()
async def test_a_newer_gap_ends_the_exemption(db_session):
"""Production's second exit path: Q3 also fails to ingest, so an un-escalated
gap joins the escalated one and the issuer pauses again immediately."""
ticker = Ticker(symbol="NEWGAP", cik="0000000050")
db_session.add_all([
ticker, _escalated_gap(ticker.cik), _prior_quarter(ticker.cik, age_days=120)
])
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == set()
fresh = datetime.now(timezone.utc)
db_session.add(SecFilingGap(
cik=ticker.cik, accession="NEWGAP-Q3", form="10-Q", index_date=date.today(),
reason="not_in_companyfacts", first_seen_at=fresh, last_attempted_at=fresh,
))
await db_session.flush()
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
ticker.id
}
+271
View File
@@ -0,0 +1,271 @@
"""Durable last-run state.
Job outcomes lived only in an in-memory dict, so every deploy wiped them and
Admin Jobs could only report "Active" with no indication a job had ever run.
"""
from datetime import datetime, timezone
import pytest
from sqlalchemy import select
from app import scheduler as sched
from app.models.job_run_state import JobRunState
from app.services import job_run_store
from tests.conftest import _test_session_factory
@pytest.fixture
def session_factory():
"""A real, independently committing session.
_persist_job_run opens its own session and commits, which is what production
does; the shared db_session fixture holds an outer transaction that a commit
would tear down.
"""
return _test_session_factory
async def _rows(session) -> dict[str, JobRunState]:
result = await session.execute(select(JobRunState))
return {row.job_name: row for row in result.scalars().all()}
async def _committed_rows() -> dict[str, JobRunState]:
async with _test_session_factory() as session:
return await _rows(session)
class TestRecordFinish:
async def test_inserts_then_updates_one_row_per_job(self, db_session):
await job_run_store.record_finish(
db_session,
"rr_scanner",
{"status": "completed", "finished_at": "2026-08-08T10:00:00+00:00", "processed": 5, "total": 5},
)
await db_session.flush()
await job_run_store.record_finish(
db_session,
"rr_scanner",
{"status": "error", "finished_at": "2026-08-08T12:00:00+00:00", "message": "boom"},
)
await db_session.flush()
rows = await _rows(db_session)
assert list(rows) == ["rr_scanner"], "upsert, not append-only history"
assert rows["rr_scanner"].status == "error"
assert rows["rr_scanner"].message == "boom"
async def test_missing_finish_time_falls_back_to_now(self, db_session):
await job_run_store.record_finish(db_session, "alerts", {"status": "completed"})
await db_session.flush()
assert (await _rows(db_session))["alerts"].finished_at is not None
class TestPipelinePersistence:
"""_run_pipeline is the only write path for steps: they are plain coroutine
calls, so they emit no scheduler events for the listener to catch."""
@pytest.fixture(autouse=True)
def _enabled(self, monkeypatch, session_factory):
async def enabled(db, job_name):
return True
monkeypatch.setattr("app.scheduler.async_session_factory", session_factory)
monkeypatch.setattr("app.scheduler._is_job_enabled", enabled)
async def test_persists_both_the_step_and_the_orchestrator(self, monkeypatch):
async def ok_step():
sched._runtime_finish("rr_scanner", "completed", processed=3, total=3)
monkeypatch.setattr(sched, "ok_step", ok_step, raising=False)
await sched._run_pipeline("near_close_pipeline", [("rr_scanner", "ok_step")])
rows = await _committed_rows()
assert rows["rr_scanner"].status == "completed"
assert rows["near_close_pipeline"].status == "completed"
async def test_a_failing_step_still_records_its_error(self, monkeypatch):
"""The persist sits after the except that swallows step errors — inside
it, exactly the runs worth seeing would be skipped."""
async def boom():
sched._runtime_finish("rr_scanner", "error", processed=0, total=1, message="kaboom")
raise RuntimeError("kaboom")
monkeypatch.setattr(sched, "boom", boom, raising=False)
await sched._run_pipeline("near_close_pipeline", [("rr_scanner", "boom")])
rows = await _committed_rows()
assert rows["rr_scanner"].status == "error"
assert rows["rr_scanner"].message == "kaboom"
# The pipeline itself survives a failing step.
assert rows["near_close_pipeline"].status == "completed"
async def test_disabled_pipeline_records_skipped(self, monkeypatch):
async def disabled(db, job_name):
return False
monkeypatch.setattr("app.scheduler._is_job_enabled", disabled)
await sched._run_pipeline("daily_pipeline", [])
assert (await _committed_rows())["daily_pipeline"].status == "skipped"
async def test_persistence_failure_never_breaks_the_pipeline(self, monkeypatch):
calls: list[str] = []
async def exploding_record(db, job_name, runtime):
calls.append(job_name)
raise RuntimeError("db down")
async def ok_step():
sched._runtime_finish("rr_scanner", "completed", processed=1, total=1)
monkeypatch.setattr(sched.job_run_store, "record_finish", exploding_record)
monkeypatch.setattr(sched, "ok_step", ok_step, raising=False)
await sched._run_pipeline("near_close_pipeline", [("rr_scanner", "ok_step")])
assert calls, "persistence was attempted"
assert sched.get_job_runtime_snapshot("near_close_pipeline")["status"] == "completed"
async def test_a_job_that_never_finished_writes_nothing(self):
sched._runtime_start("event_study", total=1)
await sched._persist_job_run("event_study")
assert "event_study" not in await _committed_rows()
class TestListJobsSplitsLiveFromPersisted:
async def test_a_stale_error_does_not_pin_the_status_chip(self, db_session):
"""runtime_* must stay live-only: the chip and the rate-limit banner read
it, so a week-old error there would read as the current state forever."""
from app.scheduler import configure_scheduler, scheduler
from app.services.admin_service import list_jobs
scheduler.remove_all_jobs()
configure_scheduler()
# _job_runtime is module-global and survives across tests; pin the live
# row to idle so the assertion is about the split, not about ordering.
sched._job_runtime["rr_scanner"] = sched._idle_runtime()
await job_run_store.record_finish(
db_session,
"rr_scanner",
{
"status": "error",
"finished_at": datetime(2026, 8, 1, tzinfo=timezone.utc).isoformat(),
"message": "old failure",
},
)
await db_session.flush()
job = {j["name"]: j for j in await list_jobs(db_session)}["rr_scanner"]
assert job["last_run_status"] == "error"
assert job["last_run_message"] == "old failure"
assert job["runtime_status"] == "idle"
assert job["running"] is False
scheduler.remove_all_jobs()
class TestConcurrentWrites:
"""Pipelines are separate scheduler jobs that can overlap, and they share
step ids data_collector belongs to all four."""
async def test_interleaved_first_writes_do_not_collide(self):
"""Both sessions SELECT before either INSERTs: select-then-insert lost
this race with an IntegrityError, and the caller swallows it."""
async with _test_session_factory() as a, _test_session_factory() as b:
await job_run_store.record_finish(
a, "data_collector",
{"status": "completed", "finished_at": "2026-08-08T10:00:00+00:00"},
)
await job_run_store.record_finish(
b, "data_collector",
{"status": "completed", "finished_at": "2026-08-08T10:00:01+00:00"},
)
await a.commit()
await b.commit() # must not raise
rows = await _committed_rows()
assert rows["data_collector"].status == "completed"
async def test_an_older_finish_never_rewinds_the_row(self):
"""A slower pipeline finishing an older run last must not overwrite a
newer outcome with a stale one."""
async with _test_session_factory() as s:
await job_run_store.record_finish(
s, "alerts",
{"status": "completed", "finished_at": "2026-08-08T12:00:00+00:00"},
)
await s.commit()
await job_run_store.record_finish(
s, "alerts",
{"status": "error", "finished_at": "2026-08-08T09:00:00+00:00", "message": "stale"},
)
await s.commit()
row = (await _committed_rows())["alerts"]
assert row.finished_at.isoformat().startswith("2026-08-08T12:00")
assert row.status == "completed"
assert row.message is None
async def test_a_newer_finish_still_wins(self):
async with _test_session_factory() as s:
await job_run_store.record_finish(
s, "rr_scanner",
{"status": "completed", "finished_at": "2026-08-08T09:00:00+00:00"},
)
await s.commit()
await job_run_store.record_finish(
s, "rr_scanner",
{"status": "error", "finished_at": "2026-08-08T12:00:00+00:00", "message": "boom"},
)
await s.commit()
row = (await _committed_rows())["rr_scanner"]
assert row.status == "error"
assert row.message == "boom"
class TestShutdownDrain:
async def test_drains_a_task_queued_after_the_flush_starts(self):
"""scheduler.shutdown(wait=False) returns before APScheduler dispatches
its completion events, so writes can appear mid-drain. Snapshotting the
task set once would miss them and dispose the engine underneath."""
import asyncio
done: list[str] = []
async def slow_first():
await asyncio.sleep(0.02)
done.append("first")
# Queued only once the first write is already finishing.
sched._persist_tasks.add(asyncio.get_running_loop().create_task(late()))
async def late():
await asyncio.sleep(0.02)
done.append("late")
sched._persist_tasks.clear()
sched._persist_tasks.add(asyncio.get_running_loop().create_task(slow_first()))
await sched.flush_job_run_persists(timeout=2.0)
assert done == ["first", "late"]
sched._persist_tasks.clear()
async def test_returns_promptly_when_there_is_nothing_to_drain(self):
sched._persist_tasks.clear()
await sched.flush_job_run_persists(timeout=2.0)
async def test_gives_up_rather_than_hanging_shutdown(self):
import asyncio
async def never():
await asyncio.sleep(30)
sched._persist_tasks.clear()
task = asyncio.get_running_loop().create_task(never())
sched._persist_tasks.add(task)
await sched.flush_job_run_persists(timeout=0.15) # returns, does not hang
task.cancel()
sched._persist_tasks.clear()
@@ -0,0 +1,183 @@
"""The calibration harness's pure helpers. No network, no data.
The harness's whole value is that its numbers can be trusted, so the parts that
decide whether a run is trustworthy the variant patching and the gate
evaluation are worth pinning even though the script is research-only.
"""
import importlib.util
from pathlib import Path
import pytest
from app.services import regime_monitor_service as rms
_SPEC = importlib.util.spec_from_file_location(
"regime_calibration",
Path(__file__).resolve().parents[2] / "scripts" / "run_regime_monitor_calibration.py",
)
calib = importlib.util.module_from_spec(_SPEC)
_SPEC.loader.exec_module(calib)
class TestVariantPatching:
def test_patching_restores_module_state_exactly(self):
"""A leaked patch would silently contaminate every later variant."""
before = {
name: getattr(rms, name)
for name in ("p1_trend_break", "p5_volatility", "p3_drawdown",
"f2_credit_spreads", "HY_OAS_WINDOW_DAYS")
}
with calib.patched(calib.VARIANTS["v2_reconstruction"]):
assert rms.p3_drawdown is not before["p3_drawdown"]
assert rms.HY_OAS_WINDOW_DAYS == 3653
for name, original in before.items():
assert getattr(rms, name) is original or getattr(rms, name) == original
def test_patching_restores_even_when_the_body_raises(self):
original = rms.p5_volatility
with pytest.raises(RuntimeError):
with calib.patched(calib.VARIANTS["v3"]):
raise RuntimeError("boom")
assert rms.p5_volatility is original
def test_v4_variant_patches_nothing(self):
"""The shipped methodology must be exercised as live code, not a copy,
or the harness and the service can drift apart silently."""
assert calib.VARIANTS["v4"] == {}
class TestCandidateFormulas:
def test_graduated_trend_break_never_exceeds_the_binary_one(self):
"""Half of the provable state_v4 <= state_v3 invariant."""
closes = [100.0] * 200
for last in (105.0, 100.0, 99.0, 92.0, 80.0, 50.0):
series = closes[:-1] + [last]
v4 = rms._under_200(series) # shipped
v3 = calib._v3_under_200(series) # retired
assert v4 <= v3, f"close={last}: v4 {v4} > v3 {v3}"
def test_anchored_vix_never_exceeds_the_live_formula(self):
for vix in (10, 15, 17, 20, 25, 30, 40, 55, 82):
assert rms.p5_volatility(vix) <= calib._v3_p5_volatility(vix), f"vix={vix}"
def test_the_vix_table_keeps_resolving_past_thirty(self):
assert rms.p5_volatility(30) < rms.p5_volatility(40) < rms.p5_volatility(50)
assert rms.p5_volatility(55) == 100.0
# the retired formula's defect, kept as the contrast
assert calib._v3_p5_volatility(30) == calib._v3_p5_volatility(82) == 100.0
def test_a_shallow_break_no_longer_pegs(self):
closes = [100.0] * 199 + [98.0] # ~2% below a flat 200-DMA
assert calib._v3_under_200(closes) == 100.0 # retired: pegged
assert rms._under_200(closes) < 40.0 # shipped: graded
def test_candidate_tables_are_well_formed(self):
for table in (calib.P5_VIX_ANCHORS_A, calib.P5_VIX_ANCHORS_B,
calib.P1_TREND_BREAK_ANCHORS):
xs = [x for x, _ in table]
ys = [y for _, y in table]
assert xs == sorted(xs) and len(set(xs)) == len(xs)
assert ys == sorted(ys)
assert 0.0 <= min(ys) and max(ys) <= 100.0
class TestGates:
def _row(self, date_str, state, coverage=100.0, w1=5.0):
return {
"date": date_str,
"state": {"score": state, "coverage": coverage, "pillars": []},
"warning": {"score": 10.0, "coverage": 100.0, "pillars": [
{"id": "breadth_divergence", "sensors": [{"id": "W1", "score": w1}]}
]},
"data_quality": {"stale_inputs": []},
}
def test_the_invariant_gate_catches_a_v4_row_above_its_v3_row(self):
v3 = [self._row("2026-01-02", 40.0), self._row("2026-01-05", 50.0)]
v4 = [self._row("2026-01-02", 38.0), self._row("2026-01-05", 55.0)]
gate = calib._invariant_gate(v3, v4)
assert gate["passed"] is False
assert gate["detail"][0]["date"] == "2026-01-05"
def test_the_invariant_gate_passes_when_v4_is_never_higher(self):
v3 = [self._row("2026-01-02", 40.0)]
v4 = [self._row("2026-01-02", 40.0)]
assert calib._invariant_gate(v3, v4)["passed"] is True
def test_sensor_lookup_searches_both_axes(self):
"""W1 lives under warning; a state-only search silently returns None and
made the W1 census read 0."""
row = self._row("2026-01-02", 40.0, w1=7.5)
assert calib._sensor_score(row, "breadth_divergence", "W1") == 7.5
class TestBandShares:
def test_shares_use_the_candidate_bands_not_the_imported_default(self):
"""band_for binds bands=STATE_BANDS as an import-time default, so the
grid must pass candidates explicitly or every row scores identically."""
scores = [10.0, 30.0, 55.0, 75.0]
loose = calib._band_shares(scores, (20.0, 50.0, 80.0))
tight = calib._band_shares(scores, (20.0, 50.0, 60.0))
assert loose["breaking"] == 0.0
assert tight["breaking"] == 25.0
def test_shares_sum_to_one_hundred(self):
shares = calib._band_shares([5.0, 25.0, 55.0, 85.0], (20.0, 50.0, 80.0))
assert sum(shares.values()) == pytest.approx(100.0)
class TestScenarios:
def test_the_decisive_scenario_is_computed_not_asserted(self):
rows = {s["scenario"].split()[0]: s for s in
calib._scenarios(calib.P5_VIX_ANCHORS_A, calib.P1_TREND_BREAK_ANCHORS)}
# A 2022-style AI/tech drawdown with genuinely calm credit. This is the
# meaning anchor for the breaking threshold, so it is machine-checked.
assert rows["S3a"]["C1"] == 0.0
assert rows["S3a"]["state"] == pytest.approx(70.33, abs=0.01)
assert rows["S3b"]["state"] == pytest.approx(74.00, abs=0.01)
# ...and it must clear the chosen threshold under either P2 assumption.
assert min(rows["S3a"]["state"], rows["S3b"]["state"]) > 65.0
assert rows["S1"]["state"] < 20.0
class TestRefusals:
"""The harness's safety contract: it must decline rather than under-report.
Both paths return before any network call, so these are fast and offline.
"""
def _run(self, argv, monkeypatch):
import asyncio
import sys
monkeypatch.setattr(sys, "argv", ["run_regime_monitor_calibration.py", *argv])
return asyncio.run(calib._main())
def test_refuses_without_every_required_variant(self, monkeypatch):
assert self._run(["--methodology", "v3,v4"], monkeypatch) == 2
assert self._run(["--methodology", "v2_reconstruction,v3"], monkeypatch) == 2
def test_refuses_an_unknown_variant(self, monkeypatch):
assert self._run(["--methodology", "v3,v4,nonsense"], monkeypatch) == 2
def test_refuses_a_custom_window_without_a_calendar_anchor(self, monkeypatch):
"""--sessions alone can only ever be tautological, so the start date must
be supplied explicitly once the published window is left behind."""
assert self._run(
["--methodology", ",".join(calib.REQUIRED_VARIANTS), "--sessions", "100"],
monkeypatch,
) == 2
assert self._run(
["--methodology", ",".join(calib.REQUIRED_VARIANTS), "--end", "2026-01-05"],
monkeypatch,
) == 2
def test_the_default_invocation_satisfies_its_own_requirement(self, monkeypatch):
"""A default that the requirement rejects would make every bare run fail."""
import sys
monkeypatch.setattr(sys, "argv", ["run_regime_monitor_calibration.py"])
default = calib._parse_args().methodology.split(",")
assert set(calib.REQUIRED_VARIANTS) <= set(default)
assert set(calib.REQUIRED_VARIANTS) <= set(calib.VARIANTS)
+396 -41
View File
@@ -1,4 +1,4 @@
"""Pure-function tests for the v3 AI/Tech Risk Monitor contract."""
"""Pure-function tests for the v4 AI/Tech Risk Monitor contract."""
from __future__ import annotations
@@ -28,7 +28,7 @@ from app.services.regime_monitor_service import (
drawdown_pct,
f2_credit_spreads,
current_observation,
fundamental_overlay,
fundamental_context,
p1_trend_break,
p2_death_cross,
p3_drawdown,
@@ -40,6 +40,24 @@ from app.services.regime_monitor_service import (
)
async def _no_observations(_db):
return []
async def _skip_recording(_db, _observation):
return None
class _CommitOnlyDB:
"""Enough session for writers that own their own transaction boundary."""
def __init__(self) -> None:
self.commits = 0
async def commit(self) -> None:
self.commits += 1
def _dated(values: list[float], end: date = date(2026, 6, 26)) -> list[tuple[date, float]]:
return [
(end - timedelta(days=len(values) - 1 - index), value)
@@ -52,6 +70,10 @@ def test_band_for_is_per_axis():
assert band_for(20, STATE_BANDS) == "watch"
assert band_for(50, STATE_BANDS) == "elevated"
assert band_for(80, STATE_BANDS) == "breaking"
# v4 moved the top band 80 -> 65; pin the new boundary from both sides so a
# silent revert cannot pass. band_for is inclusive at the threshold.
assert band_for(64.9, STATE_BANDS) == "elevated"
assert band_for(65, STATE_BANDS) == "breaking"
# Warning's realized range is far narrower, so it gets its own thresholds.
assert band_for(45, STATE_BANDS) == "watch"
assert band_for(45, WARNING_BANDS) == "elevated"
@@ -172,7 +194,7 @@ def test_relative_strength_flat_or_better_is_zero():
def test_volatility_and_breadth_zero_points():
assert p5_volatility(15) == 0
assert p5_volatility(30) == 100
assert p5_volatility(30) == 55
assert breadth_level_score(60) == 0
assert breadth_level_score(20) == 100
assert breadth_level_score(None) is None
@@ -211,7 +233,7 @@ def test_score_pillars_gates_band_below_75_percent_coverage():
assert result["band"] is None
def test_fundamental_overlay_never_replays_before_effective_date_and_expires():
def test_fundamental_context_never_replays_before_effective_date_and_expires():
overrides = {
"f1_score": 0.0,
"f3_score": 100.0,
@@ -222,19 +244,19 @@ def test_fundamental_overlay_never_replays_before_effective_date_and_expires():
}
config = {**DEFAULT_CONFIG, "fundamental_staleness_days": 80}
pending = fundamental_overlay(overrides, config, date(2026, 6, 1))
pending = fundamental_context(overrides, config, date(2026, 6, 1))
assert pending["pending"] is True
assert pending["available"] is False
assert pending["capex"] is None
# The effective date is still reported so a pending refresh is visible.
assert pending["effective_date"] == "2026-06-02"
live = fundamental_overlay(overrides, config, date(2026, 6, 2))
live = fundamental_context(overrides, config, date(2026, 6, 2))
assert live["available"] is True
assert live["good_news_stock_down"] == "yes"
assert live["earnings_stress"] == 100.0
expired = fundamental_overlay(overrides, config, date(2026, 8, 22))
expired = fundamental_context(overrides, config, date(2026, 8, 22))
assert expired["stale"] is True
assert expired["available"] is False
@@ -259,7 +281,7 @@ def test_live_observation_is_visible_before_its_effective_date():
config = {**DEFAULT_CONFIG, "fundamental_staleness_days": 80}
before = date(2026, 6, 1)
record = fundamental_overlay(overrides, config, before)
record = fundamental_context(overrides, config, before)
now = current_observation(overrides, config, before)
# Same day, same observation: the record hides it, the live reading shows it.
@@ -310,35 +332,256 @@ def test_an_uncollected_observation_is_not_reported_as_collected():
assert current_observation(collected, DEFAULT_CONFIG, date(2026, 8, 7))["observed"] is True
def test_fundamentals_do_not_move_the_warning_score():
"""The v3 complaint: a maxed-out LLM read must not silently do nothing.
def test_fundamental_state_never_averages_unknown_into_neutral():
"""Missing evidence must not present as evidence of normality.
It no longer feeds Warning at all, so Warning is identical either way and
the observation is reported beside the score instead of buried in it.
This is the trap that mattered when the channel replaced the weighted
modifier: treating ``unknown`` as a middle value would let two ``cutting``
reads and two ``unknown`` ones land on "neutral". A single adverse read
carries on partial evidence; ``unknown`` survives only when *nothing* was
observed.
"""
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
assert rms._capex_signal(dict.fromkeys(names, "unknown"), names) == "unknown"
assert rms._capex_signal(dict.fromkeys(names, "raising"), names) == "supportive"
assert rms._capex_signal(dict.fromkeys(names, "holding"), names) == "neutral"
half_cut = {names[0]: "cutting", names[1]: "cutting", **dict.fromkeys(names[2:], "unknown")}
assert rms._capex_signal(half_cut, names) == "adverse"
assert rms._reaction_signal("yes") == "adverse"
assert rms._reaction_signal("no") == "supportive"
assert rms._reaction_signal("mixed") == "neutral"
assert rms._reaction_signal(None) == "unknown"
combine = rms.combine_fundamental_signals
assert combine("unknown", "unknown") == "unknown"
assert combine("adverse", "supportive") == "adverse" # one adverse read carries
assert combine("supportive", "unknown") == "supportive"
assert combine("neutral", "unknown") == "neutral"
assert combine("supportive", "neutral") == "neutral"
# Nothing combines *into* unknown -- that would be inventing missing evidence.
assert "unknown" not in {
combine(a, b)
for a in rms.FUNDAMENTAL_STATES
for b in rms.FUNDAMENTAL_STATES
if not (a == "unknown" and b == "unknown")
}
def test_fundamental_context_is_a_channel_not_a_term_in_warning():
"""The read is reported beside the scores and never added into them.
A weighted modifier was built and reverted: with ~10 correction events and
almost no fundamental history, any fusion weight is a policy preference
presented as a measurement, and adding a slow categorical judgement to a fast
continuous score manufactures precision by summing unlike things.
"""
end = date(2026, 6, 26)
rising = [100.0 + index * 0.2 for index in range(700)]
prices = {"SMH": _dated(rising, end), "QQQ": _dated(rising, end), "SPY": _dated(rising, end)}
args = (prices, [(end, 20.0)], [(end - timedelta(days=i), 4.0) for i in reversed(range(100))])
tail = (copy.deepcopy(DEFAULT_CONFIG), end, [(end, 55.0)], [(end, 20.0)], {end: 25})
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
quiet = _compute_index(*args, {"f1_score": None, "f3_score": None}, *tail)
screaming = _compute_index(
*args,
{
"f1_score": 100.0,
"f3_score": 100.0,
"capex": dict.fromkeys(DEFAULT_CONFIG["tickers"]["hyperscalers"], "cutting"),
"good_news_stock_down": "yes",
def observed(capex_state: str, reaction: str) -> dict:
return {
"capex": dict.fromkeys(names, capex_state),
"good_news_stock_down": reaction,
"effective_date": "2026-06-01",
"fetched_at": "2026-06-01T00:00:00+00:00",
"source": "openai",
}
unobserved = _compute_index(*args, {"f1_score": None, "f3_score": None}, *tail)
supportive = _compute_index(*args, observed("raising", "no"), *tail)
adverse = _compute_index(*args, observed("cutting", "yes"), *tail)
# Every Warning is identical: the channel is not a term in the score.
scores = {
snapshot["warning"]["score"]
for snapshot in (unobserved, supportive, adverse)
}
assert len(scores) == 1
assert {p["id"] for p in unobserved["warning"]["pillars"]} == set(WARNING_WEIGHTS)
# And it never touches coverage, so a missing observation cannot suppress a
# band or silently redistribute weight onto the technical sensors.
assert len({s["warning"]["coverage"] for s in (unobserved, supportive, adverse)}) == 1
assert unobserved["fundamental_context"]["state"] == "unknown"
assert unobserved["fundamental_context"]["evidence_quality"] == "unavailable"
assert supportive["fundamental_context"]["state"] == "supportive"
assert adverse["fundamental_context"]["state"] == "adverse"
assert adverse["fundamental_context"]["evidence_quality"] == "complete"
def test_a_fresh_but_empty_observation_is_available_to_show_and_not_usable():
"""Collected-but-determined-nothing must not count as evidence.
`available` is about timing (there is an effective, non-stale record to
display); `usable` is about content. An LLM run that failed to extract
anything produces a perfectly fresh observation that knows nothing and if
that counted, repeated extraction failures would slowly accumulate study
exposure until the fundamental rows reported a measurable 0/8 for a channel
that had never seen a thing.
"""
config = copy.deepcopy(DEFAULT_CONFIG)
names = config["tickers"]["hyperscalers"]
as_of = date(2026, 6, 26)
base = {
"effective_date": "2026-06-01",
"fetched_at": "2026-06-01T00:00:00+00:00",
"source": "openai",
}
empty = fundamental_context(
{**base, "capex": dict.fromkeys(names, "unknown"), "good_news_stock_down": "unknown"},
config, as_of,
)
assert empty["state"] == "unknown"
assert empty["available"] is True # there is a record, and it has a date
assert empty["usable"] is False # but it says nothing
# One real signal is enough to be usable, on partial evidence.
partial = fundamental_context(
{
**base,
"capex": {names[0]: "cutting", **dict.fromkeys(names[1:], "unknown")},
"good_news_stock_down": "unknown",
},
*tail,
config, as_of,
)
assert partial["state"] == "adverse"
assert partial["usable"] is True
assert partial["evidence_quality"] == "partial"
# Stale is neither available nor usable — `available` means effective *and*
# non-stale. What survives is `state`, which the card renders on its own
# (with the stale badge) so the last thing observed stays visible.
stale = fundamental_context(
{
**base,
"effective_date": "2026-01-01",
"capex": dict.fromkeys(names, "cutting"),
"good_news_stock_down": "yes",
},
config, as_of,
)
assert stale["state"] == "adverse"
assert stale["stale"] is True
assert stale["available"] is False
assert stale["usable"] is False
# Nothing collected at all: neither.
absent = fundamental_context({}, config, as_of)
assert (absent["available"], absent["usable"]) == (False, False)
def test_the_live_reading_publishes_the_same_fields_as_the_record():
""""Same shape" has to mean the same fields, not the same ones it needs.
The frontend types both payloads as one interface, so a field present on the
record and missing from the live reading is an undefined at runtime that
TypeScript cannot catch across a trusted server boundary.
"""
config = copy.deepcopy(DEFAULT_CONFIG)
names = config["tickers"]["hyperscalers"]
as_of = date(2026, 6, 26)
observation = {
"effective_date": "2026-06-01",
"fetched_at": "2026-06-01T00:00:00+00:00",
"source": "openai",
"capex": dict.fromkeys(names, "cutting"),
"good_news_stock_down": "yes",
}
record = fundamental_context(observation, config, as_of)
live = current_observation(observation, config, as_of)
assert set(record) <= set(live)
assert (live["state"], live["usable"]) == ("adverse", True)
# A just-collected observation is shown but is not yet in force, so it is
# available to read and not yet usable as evidence.
pending = current_observation(
{**observation, "effective_date": "2026-07-01"}, config, as_of
)
assert (pending["pending"], pending["available"], pending["usable"]) == (True, True, False)
# And an extraction that determined nothing is never usable, however fresh.
empty = current_observation(
{**observation, "capex": dict.fromkeys(names, "unknown"), "good_news_stock_down": "unknown"},
config, as_of,
)
assert (empty["state"], empty["usable"]) == ("unknown", False)
def test_pre_rename_snapshots_keep_their_recorded_fundamental_evidence():
"""The rename shipped without a methodology bump, so those rows were never reseeded.
Reading only the new key would turn real observations into `unknown` and
silently drop historical Path colours and legitimate study exposure.
"""
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
legacy = {
"methodology": rms.METHODOLOGY,
"date": "2026-07-01",
"state": {"score": 10.0, "band": "stable"},
"warning": {"score": 20.0, "band": "stable"},
"fundamental_overlay": {
"available": True,
"pending": False,
"stale": False,
"effective_date": "2026-06-20",
"capex": {names[0]: "cutting", **dict.fromkeys(names[1:], "raising")},
"good_news_stock_down": "yes",
"source": "openai",
"fetched_at": "2026-06-19T00:00:00+00:00",
},
}
parsed = rms._parse_snapshot(json.dumps(legacy))
context = parsed["fundamental_context"]
assert context["state"] == "adverse"
assert context["evidence_quality"] == "complete"
assert context["usable"] is True
assert context["effective_date"] == "2026-06-20"
# A pending legacy overlay carried no facts, so it stays unknown rather than
# inventing an observation for a session nobody had looked at.
blank = json.loads(json.dumps(legacy))
blank["fundamental_overlay"] = {"pending": True, "stale": False, "capex": None}
blank_context = rms._parse_snapshot(json.dumps(blank))["fundamental_context"]
assert blank_context["state"] == "unknown"
assert blank_context["evidence_quality"] == "unavailable"
assert blank_context["usable"] is False
# A row already carrying the new key is left exactly as written.
modern = json.loads(json.dumps(legacy))
modern["fundamental_context"] = {"state": "supportive", "usable": True}
assert rms._parse_snapshot(json.dumps(modern))["fundamental_context"]["state"] == "supportive"
def test_evidence_quality_ranks_what_an_operator_needs_first():
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
config = copy.deepcopy(DEFAULT_CONFIG)
full = dict.fromkeys(names, "raising")
partial = {names[0]: "raising", **dict.fromkeys(names[1:], "unknown")}
def quality(capex, reaction, *, observed=True, stale=False, source="openai"):
return rms._evidence_quality(
capex, reaction, names, observed=observed, stale=stale, source=source
)
assert quiet["warning"]["score"] == screaming["warning"]["score"]
assert {p["id"] for p in quiet["warning"]["pillars"]} == set(WARNING_WEIGHTS)
assert screaming["fundamental_overlay"]["available"] is True
assert screaming["fundamental_overlay"]["capex_stress"] == 100.0
assert quality(full, "no") == "complete"
assert quality(partial, "no") == "partial"
assert quality(full, None) == "partial" # reaction unknown
assert quality(full, "no", source="manual") == "manual"
assert quality(full, "no", stale=True) == "stale"
# Nothing collected outranks every other grade.
assert quality(full, "no", observed=False, stale=True, source="manual") == "unavailable"
assert set(rms.EVIDENCE_QUALITY) >= {quality(full, "no"), quality(partial, "no")}
assert config["tickers"]["hyperscalers"] == names
def test_capex_score_separates_holding_from_raising():
@@ -363,7 +606,7 @@ def test_fundamental_api_rejects_numeric_ordinal_overrides():
@pytest.mark.asyncio
async def test_legacy_numeric_fundamentals_do_not_leak_into_v3(monkeypatch):
async def test_legacy_numeric_fundamentals_do_not_leak_into_v4(monkeypatch):
async def fake_value(_db, _key):
return json.dumps({"f1_score": 75.0, "f3_score": 75.0, "source": "manual"})
@@ -371,24 +614,31 @@ async def test_legacy_numeric_fundamentals_do_not_leak_into_v3(monkeypatch):
result = await rms.get_fundamental_overrides(object())
assert result["methodology"] == "v3"
assert result["methodology"] == rms.METHODOLOGY
assert result["f1_score"] is None
assert result["f3_score"] is None
assert result["good_news_stock_down"] == "mixed"
# Not "mixed": an unreadable blob is an absence of an observation, and
# "mixed" is a genuinely observed mixed reaction.
assert result["good_news_stock_down"] == "unknown"
@pytest.mark.asyncio
async def test_v2_observation_survives_the_methodology_bump(monkeypatch):
@pytest.mark.parametrize("stored_methodology", ["v2", "v3"])
async def test_v2_observation_survives_the_methodology_bump(monkeypatch, stored_methodology):
"""A snapshot reseed must not throw away a hand/LLM-collected observation.
The categorical format is unchanged, so the stored capex map is still valid;
only the capex scale moved, and f1 is recomputed from the categories.
Parametrised over every methodology that could be sitting in the settings row
at cutover time -- "v3" is the one the v4 bump actually meets in production,
and losing it would silently start a paid LLM refresh on every run.
"""
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
async def fake_value(_db, _key):
return json.dumps({
"methodology": "v2",
"methodology": stored_methodology,
"f1_score": 0.0, # stale v2 scale, must be recomputed
"f3_score": 100.0,
"capex": {names[0]: "raising", **dict.fromkeys(names[1:], "holding")},
@@ -396,6 +646,7 @@ async def test_v2_observation_survives_the_methodology_bump(monkeypatch):
"source": "gemini",
"fetched_at": "2026-07-24T14:25:47+00:00",
"effective_date": "2026-07-27",
"locked": True,
})
monkeypatch.setattr(rms.settings_store, "get_value", fake_value)
@@ -405,7 +656,11 @@ async def test_v2_observation_survives_the_methodology_bump(monkeypatch):
assert result["source"] == "gemini"
assert result["good_news_stock_down"] == "yes"
assert result["effective_date"] == "2026-07-27"
assert result["f1_score"] == 37.5 # recomputed on the v3 scale, not the stored 0.0
assert result["f1_score"] == 37.5 # recomputed on the current scale, not the stored 0.0
assert result["fetched_at"] == "2026-07-24T14:25:47+00:00" # or a refresh loop starts
# locked is the operator saying "do not overwrite this". Losing it is half the
# failure mode: update_regime_monitor only auto-refreshes when locked is false.
assert result["locked"] is True
@pytest.mark.asyncio
@@ -428,11 +683,12 @@ async def test_unlock_does_not_redate_a_fundamental_observation(monkeypatch):
async def fake_update(_db, _key, value):
saved.update(json.loads(value))
return None
monkeypatch.setattr(rms, "get_fundamental_overrides", fake_get)
monkeypatch.setattr(rms, "update_setting", fake_update)
monkeypatch.setattr(rms.settings_store, "upsert_setting", fake_update)
result = await rms.set_fundamental_overrides(object(), locked=False)
result = await rms.set_fundamental_overrides(_CommitOnlyDB(), locked=False)
assert result["locked"] is False
assert result["fetched_at"] == stored["fetched_at"]
@@ -462,14 +718,22 @@ async def test_manual_fundamentals_are_categorical_and_derived(monkeypatch):
async def fake_update(_db, _key, value):
saved.update(json.loads(value))
return None
monkeypatch.setattr(rms, "get_fundamental_overrides", fake_get)
monkeypatch.setattr(rms, "update_setting", fake_update)
monkeypatch.setattr(rms.settings_store, "upsert_setting", fake_update)
# A manual save now also appends to the point-in-time series.
monkeypatch.setattr(rms, "record_fundamental_observation", _skip_recording)
capex = {names[0]: "cutting", **dict.fromkeys(names[1:], "holding")}
db = _CommitOnlyDB()
result = await rms.set_fundamental_overrides(
object(), capex=capex, good_news_stock_down="mixed"
db, capex=capex, good_news_stock_down="mixed"
)
# The series row is a second write after update_setting's own commit, so the
# writer has to take one -- record_fundamental_observation deliberately does
# not, or it would steal update_regime_monitor's transaction boundary.
assert db.commits == 1
assert result["f1_score"] == 62.5 # one cutting (100) + three holding (50)
assert result["f3_score"] is None
@@ -484,7 +748,9 @@ async def test_manual_fundamentals_are_categorical_and_derived(monkeypatch):
async def test_prior_snapshot_is_immutable_without_explicit_rebuild(db_session):
snapshot_date = date(2026, 6, 26)
first = {
"methodology": "v3",
# Must be the *current* methodology: a foreign row does not parse, so it
# reads as absent and the rewrite guard never comes into play.
"methodology": rms.METHODOLOGY,
"date": snapshot_date.isoformat(),
"state": {"score": 10.0, "band": "stable"},
"warning": {"score": 20.0, "band": "stable"},
@@ -540,7 +806,7 @@ async def test_routine_can_refresh_latest_trading_session_after_civil_day_rolls(
return {}, {}
async def fake_latest(_db):
return object(), {"methodology": "v3", "sensor_revision": rms.SENSOR_REVISION}
return object(), {"methodology": "v4", "sensor_revision": rms.SENSOR_REVISION}
async def fake_upsert(_db, result, *, rewrite_existing):
rewrites.append(rewrite_existing)
@@ -557,6 +823,8 @@ async def test_routine_can_refresh_latest_trading_session_after_civil_day_rolls(
monkeypatch.setattr(rms.breadth_service, "compute_breadth_details", fake_breadth)
monkeypatch.setattr(rms, "_latest_snapshot_row", fake_latest)
monkeypatch.setattr(rms, "_upsert_snapshot", fake_upsert)
monkeypatch.setattr(rms, "get_fundamental_observations", _no_observations)
monkeypatch.setattr(rms, "record_fundamental_observation", _skip_recording)
result = await rms.update_regime_monitor(FakeDB())
@@ -568,9 +836,9 @@ async def test_routine_can_refresh_latest_trading_session_after_civil_day_rolls(
@pytest.mark.parametrize(
("stored", "expect_reseed"),
[
({"methodology": "v3"}, True), # written before the marker existed
({"methodology": "v3", "sensor_revision": 1}, True),
({"methodology": "v3", "sensor_revision": rms.SENSOR_REVISION}, False),
({"methodology": "v4"}, True), # written before the marker existed
({"methodology": "v4", "sensor_revision": 1}, True),
({"methodology": "v4", "sensor_revision": rms.SENSOR_REVISION}, False),
],
)
async def test_a_stale_sensor_revision_reseeds_stored_history(
@@ -622,6 +890,8 @@ async def test_a_stale_sensor_revision_reseeds_stored_history(
("_fetch_fred_series", fake_fred),
("_latest_snapshot_row", fake_latest),
("_upsert_snapshot", fake_upsert),
("get_fundamental_observations", _no_observations),
("record_fundamental_observation", _skip_recording),
):
monkeypatch.setattr(rms, name, value)
monkeypatch.setattr(rms.breadth_service, "compute_breadth_details", fake_breadth)
@@ -708,6 +978,91 @@ def test_compute_index_uses_one_max_price_vote_and_has_no_combined_score():
price = next(p for p in result["state"]["pillars"] if p["id"] == "price")
sensor_scores = [sensor["score"] for sensor in price["sensors"] if sensor["score"] is not None]
assert price["score"] == max(sensor_scores)
assert result["methodology"] == "v3"
assert result["methodology"] == rms.METHODOLOGY
assert "combined" not in result
assert result["basket"]["members_available"] == 25
def test_v4_carries_categorical_fundamental_observations():
"""The costliest failure mode in the v3 -> v4 cut.
CATEGORICAL_FUNDAMENTAL_METHODOLOGIES is checked against the *stored* blob.
Omit the current methodology and the first write discards the observation;
the default that replaces it has fetched_at None and locked False, so
_fundamentals_stale is true and update_regime_monitor fires a paid LLM
refresh on every run, forever, with the operator's locked read gone.
"""
assert rms.METHODOLOGY in rms.CATEGORICAL_FUNDAMENTAL_METHODOLOGIES
# Older categorical blobs must still carry forward across the bump.
assert {"v2", "v3"} <= rms.CATEGORICAL_FUNDAMENTAL_METHODOLOGIES
def test_quadrant_dividers_match_the_band_boundaries():
"""The doc asserts dividers sit at each axis's watch/elevated boundary.
Nothing enforced it, and alert_service keeps its own fallback copies -- so a
band move could silently leave the alert path classifying on the old grid.
"""
from app.services import alert_service
assert rms.QUADRANT_STATE_DIVIDER == STATE_BANDS[1]
assert rms.QUADRANT_WARNING_DIVIDER == WARNING_BANDS[1]
assert alert_service.QUAD_X_DIV == rms.QUADRANT_STATE_DIVIDER
assert alert_service.QUAD_Y_DIV == rms.QUADRANT_WARNING_DIVIDER
def test_the_vix_sensor_keeps_headroom_past_a_thirty_print():
"""v3 read VIX 30, 50 and 82 as an identical 100 -- the same saturation v3
itself had just removed from P3."""
assert p5_volatility(30) < p5_volatility(40) < p5_volatility(50)
assert p5_volatility(55) == 100.0
assert p5_volatility(82) == 100.0
assert p5_volatility(15) == 0.0
assert p5_volatility(10) == 0.0
def test_a_shallow_trend_break_does_not_peg_the_price_pillar():
"""v3's binary _under_200 printed 100 the moment price crossed, pinning the
pillar's max() and stopping P3's ladder resolving for the whole selloff."""
end = date(2026, 6, 26)
# ~2% below a flat 200-DMA, with a shallow drawdown to match.
flat = [100.0] * 260
shallow = flat[:-1] + [98.0]
prices = {
"SMH": _dated(shallow, end),
"QQQ": _dated(shallow, end),
"SPY": _dated(flat, end),
}
result = _compute_index(
prices, [(end, 16.0)], [(end, 2.8)], {},
copy.deepcopy(DEFAULT_CONFIG), end, [(end, 55.0)], [(end, 0.0)], {end: 30},
)
price = next(p for p in result["state"]["pillars"] if p["id"] == "price")
assert price["score"] < 40.0, "a 2% break must not read as maximum stress"
p1 = next(s for s in price["sensors"] if s["id"] == "P1")
assert 0.0 < p1["score"] < 40.0
def test_anchor_tables_are_well_formed():
"""Cheap guard against a fat-fingered edit to any interpolation table."""
tables = {
"P3_DRAWDOWN_ANCHORS": rms.P3_DRAWDOWN_ANCHORS,
"P1_TREND_BREAK_ANCHORS": rms.P1_TREND_BREAK_ANCHORS,
"P5_VIX_ANCHORS": rms.P5_VIX_ANCHORS,
}
for name, table in tables.items():
xs = [x for x, _ in table]
ys = [y for _, y in table]
assert xs == sorted(xs) and len(set(xs)) == len(xs), f"{name}: x not increasing"
assert ys == sorted(ys), f"{name}: y not non-decreasing"
assert 0.0 <= min(ys) and max(ys) <= 100.0, f"{name}: out of [0,100]"
# Slopes ease off only on the two v4 tables. P3 is deliberately gentle at the
# onset then steepens (2.5, 3.75, 3.125, 2.33, 1.83), so it is excluded.
for name in ("P1_TREND_BREAK_ANCHORS", "P5_VIX_ANCHORS"):
table = tables[name]
slopes = [
(table[i + 1][1] - table[i][1]) / (table[i + 1][0] - table[i][0])
for i in range(len(table) - 1)
]
assert all(a >= b for a, b in zip(slopes, slopes[1:])), f"{name}: {slopes}"
+143
View File
@@ -5,10 +5,16 @@ different realized ranges -- Warning never exceeded 64.9 in the 408 calibration
sessions, so a shared 60 left the whole upper half of that axis unreachable.
"""
import pytest
from app.services import alert_service
from app.services.alert_service import (
CONFLUENCE_TYPE,
FUND_TYPE,
QUAD_X_DIV,
QUAD_Y_DIV,
_classify_quadrant,
_collect_regime_fundamental,
_parse_quadrant_log_key,
_quadrant_log_key,
)
@@ -48,3 +54,140 @@ def test_quadrant_key_carries_basket_hash_and_parses_legacy_keys():
assert _parse_quadrant_log_key(key) == ("abc123", "3", 32.4, 54.6)
assert _parse_quadrant_log_key("3:32.4:54.6") == (None, "3", 32.4, 54.6)
assert _parse_quadrant_log_key("3") == (None, "3", None, None)
# ---------------------------------------------------------------------------
# Fundamental-context and confluence alerts
# ---------------------------------------------------------------------------
def _monitor(
warning_score: float, state: str, *, coverage: float = 100.0, usable: bool = True
) -> dict:
return {
"available": True,
"warning": {"score": warning_score, "coverage": coverage},
"fundamental_context": {
"state": state,
"evidence_quality": "complete" if usable else "stale",
# The state survives going stale so the card can still show it, and
# a failed extraction is fresh but knows nothing; `usable` is what
# says whether it may still confirm anything.
"available": usable,
"usable": usable,
},
"data_quality": {"is_fresh": True},
"quadrant_config": {"warning_divider": QUAD_Y_DIV},
}
class _LogSpyDB:
"""Records what would be logged; returns a canned "last logged key"."""
def __init__(self, last: dict[str, str | None]) -> None:
self.last = last
self.logged: list[tuple[str, str]] = []
@pytest.fixture
def patched(monkeypatch):
def apply(data: dict, last: dict[str, str | None]):
db = _LogSpyDB(last)
async def fake_monitor(_db):
return data
async def fake_last(_db, alert_type):
return db.last.get(alert_type)
def fake_log(_db, alert_type, key, value=None):
db.logged.append((alert_type, key))
import app.services.regime_monitor_service as rms
monkeypatch.setattr(rms, "get_regime_monitor", fake_monitor)
monkeypatch.setattr(alert_service, "_last_logged_key", fake_last)
monkeypatch.setattr(alert_service, "_log_alert", fake_log)
return db
return apply
@pytest.mark.asyncio
async def test_first_run_seeds_both_channels_without_alerting(patched):
db = patched(_monitor(60.0, "adverse"), {FUND_TYPE: None, CONFLUENCE_TYPE: None})
assert await _collect_regime_fundamental(db) == []
assert dict(db.logged) == {FUND_TYPE: "adverse", CONFLUENCE_TYPE: "yes"}
@pytest.mark.asyncio
async def test_fundamental_change_and_confluence_are_separate_messages(patched):
db = patched(_monitor(60.0, "adverse"), {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "no"})
out = await _collect_regime_fundamental(db)
assert [alert_type for alert_type, _, _ in out] == [FUND_TYPE, CONFLUENCE_TYPE]
assert "neutral → adverse" in out[0][2]
assert "Confluence" in out[1][2]
# Neither message reports a fused score; they name which channel moved.
assert "not a score" in out[0][2]
@pytest.mark.asyncio
async def test_unknown_never_alerts(patched):
"""Absence of evidence is not a change in the evidence."""
db = patched(_monitor(60.0, "unknown"), {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "no"})
assert await _collect_regime_fundamental(db) == []
@pytest.mark.asyncio
async def test_adverse_alone_is_not_confluence(patched):
"""A calm tape with adverse fundamentals is a context change, not confluence."""
db = patched(_monitor(10.0, "adverse"), {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "no"})
out = await _collect_regime_fundamental(db)
assert [alert_type for alert_type, _, _ in out] == [FUND_TYPE]
@pytest.mark.asyncio
async def test_leaving_confluence_rebaselines_quietly(patched):
db = patched(_monitor(10.0, "neutral"), {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "yes"})
assert await _collect_regime_fundamental(db) == []
assert (CONFLUENCE_TYPE, "no") in db.logged
@pytest.mark.asyncio
async def test_low_coverage_or_stale_inputs_stay_quiet(patched):
thin = _monitor(60.0, "adverse", coverage=50.0)
assert await _collect_regime_fundamental(
patched(thin, {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "no"})
) == []
stale = _monitor(60.0, "adverse")
stale["data_quality"]["is_fresh"] = False
assert await _collect_regime_fundamental(
patched(stale, {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "no"})
) == []
@pytest.mark.asyncio
async def test_a_stale_observation_cannot_confirm_a_new_crossing(patched):
"""The state is kept for display, but it stops being evidence.
Without this, one adverse read corroborates every Warning crossing for the
rest of time the strongest claim the channel makes, from the data with the
least right to make it.
"""
stale = _monitor(60.0, "adverse", usable=False)
db = patched(stale, {FUND_TYPE: "adverse", CONFLUENCE_TYPE: "no"})
assert await _collect_regime_fundamental(db) == []
# It also rebaselines to "no", so recollecting the observation re-arms it.
assert (CONFLUENCE_TYPE, "no") not in db.logged # already "no"; nothing to log
fresh = _monitor(60.0, "adverse", usable=True)
db2 = patched(fresh, {FUND_TYPE: "adverse", CONFLUENCE_TYPE: "no"})
out = await _collect_regime_fundamental(db2)
assert [alert_type for alert_type, _, _ in out] == [CONFLUENCE_TYPE]
@pytest.mark.asyncio
async def test_a_stale_state_change_does_not_alert(patched):
db = patched(_monitor(10.0, "adverse", usable=False), {FUND_TYPE: "neutral", CONFLUENCE_TYPE: "no"})
assert await _collect_regime_fundamental(db) == []
+105 -43
View File
@@ -1,16 +1,19 @@
"""Unit tests for app.scheduler module."""
import asyncio
from datetime import datetime, timezone
from types import SimpleNamespace
import pytest
from app import job_catalog
from app.scheduler import (
_DAILY_PIPELINE_STEPS,
_NEAR_CLOSE_PIPELINE_STEPS,
_consume_backtest_options,
_consume_backtest_target_model,
_parse_frequency,
_repause_after_manual_run,
_resume_tickers,
_last_successful,
_run_source_import,
@@ -112,60 +115,119 @@ class TestResumeTickers:
class TestConfigureScheduler:
def test_configure_adds_all_jobs(self):
# Remove any existing jobs first
# Derived from the catalog, not a fourth hand-maintained copy of the
# job list: a job added to the catalog but never registered now fails
# here instead of silently rendering "Not registered" in the admin UI.
scheduler.remove_all_jobs()
configure_scheduler()
jobs = scheduler.get_jobs()
job_ids = {j.id for j in jobs}
assert job_ids == {
"data_collector",
"data_backfill",
"benchmark_collector",
"sentiment_collector",
"dolt_earnings_import",
"sec_fundamentals_import",
"rr_scanner",
"shadow_book",
"ticker_universe_sync",
"outcome_evaluator",
"alerts",
"market_regime",
"regime_monitor",
"event_study",
"backtest",
"daily_pipeline",
"near_close_pipeline",
"after_close_pipeline",
"intraday_pipeline",
}
assert {j.id for j in scheduler.get_jobs()} == set(job_catalog.VALID_JOB_NAMES)
def test_configure_is_idempotent(self):
scheduler.remove_all_jobs()
configure_scheduler()
configure_scheduler() # Should replace, not duplicate
job_ids = [j.id for j in scheduler.get_jobs()]
# Each ID should appear exactly once
assert sorted(job_ids) == sorted([
"after_close_pipeline",
"alerts",
"backtest",
"benchmark_collector",
"daily_pipeline",
"intraday_pipeline",
assert sorted(job_ids) == sorted(job_catalog.VALID_JOB_NAMES)
def test_independent_jobs_use_cron_not_interval(self):
"""Interval countdowns restart on every deploy, so a weekly interval on a
frequently-redeployed box can defer forever. Both standalone jobs were
migrated to cron; this pins them there."""
scheduler.remove_all_jobs()
configure_scheduler()
for job_id in ("backtest", "ticker_universe_sync"):
trigger = type(scheduler.get_job(job_id).trigger).__name__
assert trigger == "CronTrigger", f"{job_id} regressed to {trigger}"
class TestJobCatalog:
def test_pipeline_members_are_derived_from_step_lists(self):
derived = {
step
for steps in job_catalog.PIPELINE_STEPS.values()
for step, _ in steps
}
assert job_catalog.PIPELINE_MEMBERS == derived
# ...and reproduces the set that used to be maintained by hand, so the
# derivation is behaviour-preserving rather than merely self-consistent.
assert job_catalog.PIPELINE_MEMBERS == {
"data_collector",
"data_backfill",
"dolt_earnings_import",
"sec_fundamentals_import",
"market_regime",
"near_close_pipeline",
"regime_monitor",
"event_study",
"outcome_evaluator",
"rr_scanner",
"benchmark_collector",
"sentiment_collector",
"rr_scanner",
"shadow_book",
"ticker_universe_sync",
])
"outcome_evaluator",
"alerts",
"market_regime",
"regime_monitor",
}
def test_categories_partition_every_job_exactly_once(self):
buckets = [
job_catalog.PIPELINE_JOBS,
job_catalog.PIPELINE_STEP_JOBS,
job_catalog.SCHEDULED_JOBS,
job_catalog.MANUAL_JOBS,
]
flat = [name for bucket in buckets for name in bucket]
assert len(flat) == len(set(flat)), "a job is in two categories"
assert set(flat) == set(job_catalog.VALID_JOB_NAMES)
assert all(name in job_catalog.JOB_CATEGORY for name in flat)
def test_every_job_has_a_label_and_a_unique_sort_order(self):
names = job_catalog.VALID_JOB_NAMES
assert set(job_catalog.JOB_LABELS) == set(names)
assert len({job_catalog.sort_order(n) for n in names}) == len(names)
def test_multi_pipeline_members_report_every_parent(self):
"""Membership is many-to-many — the reason the UI groups into sections
rather than nesting steps under one parent."""
by_member = job_catalog.PIPELINES_BY_MEMBER
assert set(by_member["data_collector"]) == set(job_catalog.PIPELINE_JOBS)
assert set(by_member["alerts"]) == {"daily_pipeline", "near_close_pipeline"}
assert set(by_member["outcome_evaluator"]) == {
"intraday_pipeline",
"after_close_pipeline",
}
assert "backtest" not in by_member
def test_every_job_has_a_runtime_row_before_it_first_runs(self):
"""The old private _JOB_NAMES list held 16 of 19, so three jobs showed no
last-run line until their first run in a given process."""
assert set(get_job_runtime_snapshot()) == set(job_catalog.VALID_JOB_NAMES)
class TestRepauseListener:
def _configured(self):
scheduler.remove_all_jobs()
configure_scheduler()
def test_manual_job_is_repaused_after_running(self):
"""Triggering a paused job re-arms its 520-week backstop, which used to
surface as a "next run in ~87600h"."""
self._configured()
scheduler.modify_job("event_study", next_run_time=datetime.now(timezone.utc))
_repause_after_manual_run(SimpleNamespace(job_id="event_study"))
assert scheduler.get_job("event_study").next_run_time is None
def test_pipeline_step_is_repaused_after_running(self):
self._configured()
scheduler.modify_job("rr_scanner", next_run_time=datetime.now(timezone.utc))
_repause_after_manual_run(SimpleNamespace(job_id="rr_scanner"))
assert scheduler.get_job("rr_scanner").next_run_time is None
def test_cron_jobs_are_left_alone(self):
# Set an explicit next run first: an unstarted scheduler leaves the
# attribute unset, so comparing None to None would prove nothing.
self._configured()
due = datetime.now(timezone.utc)
scheduler.modify_job("daily_pipeline", next_run_time=due)
_repause_after_manual_run(SimpleNamespace(job_id="daily_pipeline"))
assert scheduler.get_job("daily_pipeline").next_run_time == due
def test_unknown_job_is_ignored(self):
self._configured()
_repause_after_manual_run(SimpleNamespace(job_id="not_a_job"))
class _SessionContext:
+85
View File
@@ -538,3 +538,88 @@ def test_a_wrong_declared_year_end_no_longer_collides_two_periods():
keys = {(r.fiscal_year, r.fiscal_period) for r in res.rows}
assert len(keys) == 2, f"periods collided on one key: {keys}"
assert keys == {(2026, "Q1"), (2026, "Q2")}
# --- debt composition across the tagging styles large filers actually use ----
# Values are the real shapes measured 2026-08; before this composition, seven of
# nineteen sampled large caps carried a materially wrong or absent total_debt.
_RD = date(2026, 3, 28)
def _f(concept, val):
return Fact("us-gaap", concept, "USD", None, _RD, val, 2026, "Q2")
def test_debt_from_a_noncurrent_lease_aggregate_adds_its_current_side():
"""KO/HD/T/XOM/CVX tag LongTermDebtAndCapitalLeaseObligations, which nothing
read before AT&T reported no debt at all against 134bn tagged."""
facts = [_f("LongTermDebtAndCapitalLeaseObligations", 134_630), _f("DebtCurrent", 9_320)]
assert _compose_debt(facts, _RD) == 143_950
def test_debt_current_is_the_whole_current_side_not_an_addition():
"""DebtCurrent already spans short-term borrowing AND current maturities, so
adding commercial paper on top would count it twice."""
facts = [
_f("LongTermDebtNoncurrent", 22_840),
_f("DebtCurrent", 11_300),
_f("LongTermDebtCurrent", 6_460),
_f("CommercialPaper", 4_840),
]
assert _compose_debt(facts, _RD) == 34_140
def test_debt_falls_back_to_the_split_current_parts():
facts = [
_f("LongTermDebtNoncurrent", 36_890),
_f("LongTermDebtCurrent", 3_900),
_f("ShortTermBorrowings", 10_670),
]
assert _compose_debt(facts, _RD) == 51_460
def test_notes_payable_is_the_unsecured_side_when_nothing_names_it():
"""Realty Income and VMRK tag a secured and an unsecured side, no aggregate."""
facts = [_f("NotesPayable", 25_090), _f("SecuredDebt", 40), _f("CommercialPaper", 1_400)]
assert _compose_debt(facts, _RD) == 26_530
def test_an_explicit_unsecured_side_wins_over_notes_payable():
"""MAA tags NotesPayable 5.66bn = UnsecuredDebt 5.30bn + SecuredDebt 0.36bn, so
NotesPayable is the total there and adding SecuredDebt to it double-counts.
Preferring the explicit unsecured side reproduces the total either way."""
facts = [_f("NotesPayable", 5_660), _f("UnsecuredDebt", 5_300), _f("SecuredDebt", 360)]
assert _compose_debt(facts, _RD) == 5_660
def test_one_side_of_a_reits_debt_is_not_a_total():
"""Boston Properties tags SecuredDebt 4.28bn and commercial paper against ~15bn
of real debt; Ventas the same shape. Composing from one side invents a total."""
assert _compose_debt([_f("SecuredDebt", 4_280), _f("CommercialPaper", 750)], _RD) is None
assert _compose_debt([_f("UnsecuredDebt", 5_300)], _RD) is None
def test_current_maturities_alone_are_not_a_total():
"""LongTermDebtCurrent used to stand in for the whole long-term side, which
reports the slice due within a year as if it were the debt."""
assert _compose_debt([_f("LongTermDebtCurrent", 6_460)], _RD) is None
def test_an_aggregate_beats_the_reit_parts():
"""AvalonBay tags all three; summing the parts would understate the total."""
facts = [_f("LongTermDebt", 9_020), _f("SecuredDebt", 700), _f("UnsecuredDebt", 7_410),
_f("CommercialPaper", 920)]
assert _compose_debt(facts, _RD) == 9_940
def test_a_short_term_only_filing_reports_no_total_at_all():
"""Chevron tags its full debt only in the 10-K, so a 10-Q carries 0.40bn of
short-term borrowing alone reporting that as *total* debt reads as a
near-unlevered issuer carrying 50bn. None costs a leverage read; the partial
value produces a confidently wrong one."""
assert _compose_debt([_f("ShortTermBorrowings", 401)], _RD) is None
def test_no_debt_facts_at_all_is_still_none():
assert _compose_debt([_f("CashAndCashEquivalentsAtCarryingValue", 100)], _RD) is None
+337 -1
View File
@@ -6,7 +6,7 @@ from __future__ import annotations
import json
import os
import tempfile
from datetime import date, datetime, timezone
from datetime import date, datetime, timedelta, timezone
import pytest
from sqlalchemy import func, select
@@ -950,6 +950,9 @@ async def test_discrepancy_in_shares_is_detected_and_reported(engine):
assert k.shares_outstanding == 999.0 and k.import_run_id == 1 # immutable — not overwritten
events = (await s.execute(select(SystemEvent).where(SystemEvent.code == "snapshot_discrepancy"))).scalars().all()
assert len(events) == 1 and events[0].severity == "warning"
# The alert has to say WHICH column moved: a differing cik is a co-registrant
# attribution, a differing revenue is our numbers changing.
assert "K (shares_outstanding, shares_outstanding_date)" in events[0].message
# --- reparse: rewriting rows a fixed parser reconstructs differently --------
@@ -1128,3 +1131,336 @@ async def test_malformed_cik_override_is_ignored_not_fatal(engine):
resolved = await resolve_ciks(db, client)
assert resolved.symbol_to_cik["AAPL"] == 320193 # fell back to company_tickers
# --- aggregate deferral ceiling -------------------------------------------
def _blocking_staged(today: date) -> StagedFundamentals:
"""One filing still inside the per-filing retry window."""
return StagedFundamentals(
resolved=ResolvedUniverse(),
missing_xbrl=[{
"cik": "0000320193",
"accession": "YOUNG-1",
"form": "10-Q",
"index_date": today,
"age_days": 0,
"reason": "not_in_companyfacts",
}],
)
async def _add_run(factory, *, status: str, started_at: datetime) -> None:
from app.models.data_import_run import DataImportRun
async with factory() as db:
db.add(DataImportRun(
source="sec_facts", status=status, started_at=started_at,
))
await db.commit()
async def _validate_with_history(engine, *, promoted_days_ago: int | None):
factory = _factory(engine)
today = date(2026, 5, 20)
now = datetime(2026, 5, 20, 12, 0, tzinfo=timezone.utc)
if promoted_days_ago is not None:
await _add_run(
factory,
status=STATUS_PROMOTED,
started_at=now - timedelta(days=promoted_days_ago),
)
importer = SecFundamentalsImporter(today=today)
importer._latest_index_date = date(2026, 5, 19)
staged = _blocking_staged(today)
async with factory() as db:
return await importer.validate(db, staged), staged, importer
async def test_ceiling_forces_a_promotion_once_deferral_outlasts_it(engine):
"""The per-filing window bounds one filing; this bounds the whole import."""
result, staged, importer = await _validate_with_history(engine, promoted_days_ago=10)
assert result.ok is True
assert result.summary["missing_xbrl_blocking"] == 0
assert result.summary["promotion_ceiling_tripped"] == {"forced": 1, "unresolved": 1}
# Aged in place, so promote() re-derives the same verdict and queues it.
assert staged.missing_xbrl[0]["age_days"] > 3
# The symbol stays barred from setups — promoting is not trusting the data.
assert result.summary["setup_blocked_ciks"] == ["0000320193"]
async def test_a_recent_promotion_keeps_the_normal_block(engine):
result, staged, _ = await _validate_with_history(engine, promoted_days_ago=1)
assert result.ok is False
assert result.summary["missing_xbrl_blocking"] == 1
assert result.summary["promotion_ceiling_tripped"] is None
assert result.retryable is True
assert staged.missing_xbrl[0]["age_days"] == 0
async def test_ceiling_never_fires_before_a_first_promotion(engine):
"""No baseline means initial setup, not a wedge — forcing it through would
mask a misconfiguration instead of recovering from an SEC gap."""
result, _, _ = await _validate_with_history(engine, promoted_days_ago=None)
assert result.ok is False
assert result.summary["promotion_ceiling_tripped"] is None
async def test_ceiling_promotes_queues_and_alerts_end_to_end(engine, monkeypatch):
"""The self-heal claim, end to end: a filing that would block forever gets
promoted through, queued for retry, and announced."""
from app.models.data_import_run import DataImportRun
from app.services.sec_facts_parser import ParseResult
from sqlalchemy import update as sa_update
factory = _factory(engine)
await _seed(factory, ["AAPL"])
backfill = FakeSecClient(
tickers={"AAPL": 320193},
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
submissions={320193: _submissions(SUB_FILINGS)},
latest_index=date(2026, 1, 31),
)
assert (await run_import(_importer(backfill), engine=engine)).status == STATUS_PROMOTED
# Age the only promotion past the ceiling: this is the wedge the ceiling exists
# for — the filing below stays young, so nothing else would ever release it.
async with factory() as db:
await db.execute(
sa_update(DataImportRun)
.where(DataImportRun.source == "sec_facts")
.values(started_at=datetime(2026, 4, 20, tzinfo=timezone.utc))
)
await db.commit()
# Present in Company Facts but unparseable — one missing_xbrl entry, not the
# two an absent-from-facts accession would also raise.
stuck_fact = _rev("2025-09-28", "2026-03-28", 254940, 2026, "Q2", "STUCK")
stuck_share = _shares("2026-04-17", 14687, "STUCK", 2026, "Q2")
client = FakeSecClient(
tickers={"AAPL": 320193},
companyfacts={
320193: _companyfacts(
[CF_K, CF_Q1, stuck_fact], [SH_K, SH_Q1, stuck_share]
)
},
submissions={320193: _submissions(SUB_FILINGS + [
_filing("STUCK", "10-Q", "2026-03-28", "2026-05-01",
"2026-05-01T10:01:00.000Z"),
])},
latest_index=date(2026, 5, 2),
daily={date(2026, 5, 1): [
{"form": "10-Q", "cik": 320193, "accession": "STUCK"},
]},
)
monkeypatch.setattr(
"app.services.sec_facts_parser.parse_snapshots",
lambda *a, **k: ParseResult(
skipped_filings=[{"accession": "STUCK", "reason": "unparseable"}]
),
)
# index_date 2026-05-01 vs today 2026-05-03 => 2 days old, still inside the
# per-filing window, so only the aggregate ceiling can let this through.
run = await run_import(_importer(client, today=date(2026, 5, 3)), engine=engine)
assert run.status == STATUS_PROMOTED
summary = json.loads(run.validation_json)
assert summary["promotion_ceiling_tripped"] == {"forced": 1, "unresolved": 1}
async with factory() as db:
gap = (await db.execute(select(SecFilingGap))).scalar_one()
events = (
await db.execute(
select(SystemEvent).where(
SystemEvent.code == "promotion_ceiling_forced"
)
)
).scalars().all()
# Queued, so later runs retry it without it ever blocking again...
assert gap.accession == "STUCK"
# ...and the safety valve firing is visible, not silent.
assert len(events) == 1
assert events[0].severity == "warning"
assert "7 days" in events[0].message
# --- attribution collisions: two tracked CIKs claiming one filing ----------
# A REIT and its operating partnership co-file one 10-K, and SEC's
# company_tickers.json points the old symbol at the partnership (EQR ->
# ERP Operating LP) while the issuer itself trades under a new one (VMRK).
_COMBINED = [_filing("COMBINED-K", "10-K", "2025-12-31", "2026-02-13",
"2026-02-13T21:00:00.000Z")]
_CF_COMBINED = _rev("2025-01-01", "2025-12-31", 2900000, 2025, "FY", "COMBINED-K")
_SH_COMBINED = _shares("2026-02-01", 380000, "COMBINED-K", 2025, "FY")
def _reit_submissions(cik, tickers):
return {"cik": cik, "sic": "6798", "sic_description": "REIT",
"fiscal_year_end": "1231", "tickers": tickers, "filings": _COMBINED}
def _reit_client(tickers):
return FakeSecClient(
tickers=tickers,
companyfacts={
cik: _companyfacts([_CF_COMBINED], [_SH_COMBINED], cik=cik)
for cik in tickers.values()
},
submissions={
cik: _reit_submissions(cik, [sym]) for sym, cik in tickers.items()
},
latest_index=date(2026, 3, 1),
)
async def test_cik_collision_is_reported_as_attribution_not_discrepancy(engine):
"""Only `cik` differs, so nothing was re-parsed differently — the universe
resolves a co-registrant it should not track, and the alert must say that."""
factory = _factory(engine)
await _seed(factory, ["VMRK"])
run = await run_import(
_importer(_reit_client({"VMRK": 906107}), today=date(2026, 3, 2)), engine=engine
)
assert run.status == STATUS_PROMOTED
# The stale symbol is added, resolving to the partnership's CIK.
await _seed(factory, ["EQR"])
run = await run_import(
_importer(_reit_client({"VMRK": 906107, "EQR": 931182}), today=date(2026, 3, 2)),
engine=engine,
)
assert run.status == STATUS_PROMOTED
# Production's shape: a run-level incremental in which the untracked-until-now
# CIK is individually backfilled (run 63 recorded exactly this).
assert '"backfill": false' in (run.validation_json or "")
async with factory() as s:
rows = (await s.execute(select(FundamentalSnapshot))).scalars().all()
events = (await s.execute(select(SystemEvent))).scalars().all()
# The filing stays with the issuer that filed it, stored once.
assert [(r.accession, r.cik) for r in rows] == [("COMBINED-K", "0000906107")]
codes = {e.code for e in events}
assert "accession_cik_collision" in codes
assert "snapshot_discrepancy" not in codes # not a reconstruction change
collision = next(e for e in events if e.code == "accession_cik_collision")
assert "stored 0000906107, parsed 0000931182" in collision.message
assert "sec_cik_overrides" in collision.message # names the actual fix
async def test_reparse_never_restamps_a_collision_onto_the_co_registrant(engine):
"""A reparse rewrites rows a fixed parser reconstructs differently. A cik-only
difference is not that: rewriting would hand the filing to the co-registrant."""
from app.services.sec_facts_parser import SnapshotRow
factory = _factory(engine)
await _seed(factory, ["VMRK"])
assert (await run_import(
_importer(_reit_client({"VMRK": 906107}), today=date(2026, 3, 2)), engine=engine
)).status == STATUS_PROMOTED
importer = _importer(_reit_client({"VMRK": 906107}), today=date(2026, 3, 2))
importer.reparse = True
staged = StagedFundamentals(
resolved=ResolvedUniverse(),
rows=[SnapshotRow(
cik="0000931182", accession="COMBINED-K", form="10-K",
filed_date=date(2026, 2, 13),
accepted_at=datetime(2026, 2, 13, 21, tzinfo=timezone.utc),
period_end=date(2025, 12, 31), fiscal_year=2025, fiscal_period="FY",
)],
existing_accessions={"COMBINED-K"},
discrepancies=[{
"accession": "COMBINED-K", "fields": ["cik"],
"cik": "0000931182", "stored_cik": "0000906107",
}],
)
async with _factory(engine)() as db:
counts = await importer.promote(db, staged, run_id=999)
await db.commit()
assert counts["updated"] == 0
async with factory() as s:
row = (await s.execute(select(FundamentalSnapshot))).scalar_one()
assert row.cik == "0000906107" # still the issuer that filed it
# --- the reprieve ending: an exemption that lapses must not do so silently ---
def _stale_gap(cik, *, exempted: bool):
now = datetime.now(timezone.utc)
return SecFilingGap(
cik=cik, accession=f"{cik}-AGED-Q", form="10-Q",
index_date=(now - timedelta(days=30)).date(), reason="not_in_companyfacts",
first_seen_at=now - timedelta(days=30), last_attempted_at=now,
escalated_at=now - timedelta(days=16),
exempted_at=(now - timedelta(days=16)) if exempted else None,
)
def _snapshot(cik, *, age_days):
filed = date.today() - timedelta(days=age_days)
return FundamentalSnapshot(
cik=cik, accession=f"{cik}-PRIOR", form="10-Q", filed_date=filed,
accepted_at=datetime.now(timezone.utc) - timedelta(days=age_days),
period_end=filed, fiscal_year=filed.year, fiscal_period="Q1",
)
async def _promote_only(engine, seed):
"""Run promote() alone against seeded gap/snapshot state."""
factory = _factory(engine)
async with factory() as s:
for obj in seed:
s.add(obj)
await s.commit()
importer = _importer(FakeSecClient(
tickers={}, companyfacts={}, submissions={}, latest_index=date(2026, 3, 1)
))
async with factory() as db:
await importer.promote(db, StagedFundamentals(resolved=ResolvedUniverse()), run_id=77)
await db.commit()
async with factory() as s:
gaps = (await s.execute(select(SecFilingGap))).scalars().all()
events = (await s.execute(select(SystemEvent))).scalars().all()
return gaps, events
async def test_a_lapsed_exemption_raises_its_own_alert(engine):
"""filing_gap_aged fires once and never again, so nothing else would say the
pause came back when the issuer's own fundamentals aged out."""
cik = "0000000060"
gaps, events = await _promote_only(
engine, [_stale_gap(cik, exempted=True), _snapshot(cik, age_days=400)]
)
repaused = [e for e in events if e.code == "filing_gap_repaused"]
assert len(repaused) == 1
assert f"{cik}/{cik}-AGED-Q" in repaused[0].message
# Cleared, so a later recovery can re-arm and lapse again.
assert gaps[0].exempted_at is None
async def test_an_exemption_taking_effect_is_stamped_silently(engine):
"""Setups resuming is what filing_gap_aged already described — stamping the
state must not raise a second alert for it."""
cik = "0000000061"
gaps, events = await _promote_only(
engine, [_stale_gap(cik, exempted=False), _snapshot(cik, age_days=120)]
)
assert [e.code for e in events if e.code.startswith("filing_gap")] == []
assert gaps[0].exempted_at is not None
async def test_a_still_paused_gap_is_not_reported_as_lapsing(engine):
"""It never became exempt, so there is no transition to report."""
cik = "0000000062"
gaps, events = await _promote_only(
engine, [_stale_gap(cik, exempted=False), _snapshot(cik, age_days=400)]
)
assert [e for e in events if e.code == "filing_gap_repaused"] == []
assert gaps[0].exempted_at is None
+416
View File
@@ -0,0 +1,416 @@
"""Delisting lifecycle: marking, the active_only filter, and SEC confirmation.
The behaviour under test is that a delisted symbol leaves the *live* path while
its rows stay put deleting it instead is what makes the backtest universe
survivorship-biased, so retention is the point, not a side effect.
"""
from __future__ import annotations
import json
from collections.abc import AsyncGenerator
from datetime import date
import httpx
import pytest
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine
from app.database import Base
from app.models.ticker import Ticker
from app.services import ticker_service
from app.services.sec_client import SecClient
_engine = create_async_engine("sqlite+aiosqlite://", echo=False)
_session_factory = async_sessionmaker(_engine, class_=AsyncSession, expire_on_commit=False)
@pytest.fixture(autouse=True)
async def _setup_tables() -> AsyncGenerator[None, None]:
async with _engine.begin() as conn:
await conn.run_sync(Base.metadata.create_all)
yield
async with _engine.begin() as conn:
await conn.run_sync(Base.metadata.drop_all)
@pytest.fixture
async def session() -> AsyncGenerator[AsyncSession, None]:
async with _session_factory() as s:
yield s
def _submissions(forms: list[str], dates: list[str]) -> dict:
return {
"cik": 712515,
"name": "ELECTRONIC ARTS INC.",
"filings": {
"recent": {
"form": forms,
"filingDate": dates,
"accessionNumber": [f"0001354457-26-{i:06d}" for i in range(len(forms))],
"primaryDocument": ["xslF25X02/primary_doc.xml"] * len(forms),
}
},
}
def _sec_client(payload: dict, security: str | None = "Common Stock") -> SecClient:
"""Mock submissions + the Form 25 primary document the class check reads."""
def handler(request: httpx.Request) -> httpx.Response:
if request.url.path.endswith("primary_doc.xml"):
if security is None:
return httpx.Response(404)
body = (
"<?xml version='1.0'?><notificationOfRemoval>"
f"<descriptionClassSecurity>{security}</descriptionClassSecurity>"
"</notificationOfRemoval>"
)
return httpx.Response(200, content=body.encode())
return httpx.Response(200, content=json.dumps(payload).encode())
return SecClient(transport=httpx.MockTransport(handler), spacing_seconds=0)
async def test_mark_delisted_is_idempotent(session: AsyncSession):
session.add(Ticker(symbol="EA"))
await session.commit()
assert await ticker_service.mark_delisted(
session, "EA", delisted_on=date(2026, 8, 4)
) is True
# A second call must not churn the row — the staleness path retries daily.
assert await ticker_service.mark_delisted(
session, "EA", delisted_on=date(2026, 9, 1)
) is False
row = (await session.execute(select(Ticker).where(Ticker.symbol == "EA"))).scalar_one()
assert row.delisted_on == date(2026, 8, 4) # first date wins, not the retry
assert row.delisted_reason == ticker_service.REASON_MANUAL
async def test_clear_delisted_restores_the_symbol(session: AsyncSession):
session.add(Ticker(symbol="EA"))
await session.commit()
await ticker_service.mark_delisted(session, "EA", delisted_on=date(2026, 8, 4))
assert await ticker_service.clear_delisted(session, "EA") is True
assert await ticker_service.clear_delisted(session, "EA") is False
row = (await session.execute(select(Ticker).where(Ticker.symbol == "EA"))).scalar_one()
assert row.delisted_on is None and row.delisted_reason is None
async def test_active_only_filters_but_the_row_survives(session: AsyncSession):
session.add_all([Ticker(symbol="AAPL"), Ticker(symbol="EA")])
await session.commit()
await ticker_service.mark_delisted(session, "EA", delisted_on=date(2026, 8, 4))
active = (
await session.execute(ticker_service.active_only(select(Ticker.symbol)))
).scalars().all()
assert list(active) == ["AAPL"]
# The whole point: the row — and everything cascading off it — is still there.
everything = [t.symbol for t in await ticker_service.list_tickers(session)]
assert everything == ["AAPL", "EA"]
async def test_confirm_delisting_marks_on_a_form_25(session: AsyncSession, monkeypatch):
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["8-K", "25-NSE"], ["2026-07-01", "2026-08-04"])),
raising=False,
)
marked = await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
)
# Removal is effective ten days after the 2026-08-04 filing, not on it.
assert marked == date(2026, 8, 14)
row = (await session.execute(select(Ticker).where(Ticker.symbol == "EA"))).scalar_one()
assert row.delisted_reason == ticker_service.REASON_FORM_25
async def test_confirm_delisting_leaves_a_halt_alone(session: AsyncSession, monkeypatch):
"""A halt or a rename files no Form 25 — those must keep warning, not retire."""
session.add(Ticker(symbol="SATS", cik="0000012345"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["8-K", "10-Q"], ["2026-07-01", "2026-08-04"])),
raising=False,
)
assert await ticker_service.confirm_delisting(
session, "SATS", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) is None
row = (await session.execute(select(Ticker).where(Ticker.symbol == "SATS"))).scalar_one()
assert row.delisted_on is None
async def test_confirm_delisting_skips_a_symbol_without_a_cik(session: AsyncSession):
"""No CIK, no SEC lookup — must not raise, and must not mark."""
session.add(Ticker(symbol="ADRX"))
await session.commit()
assert await ticker_service.confirm_delisting(
session, "ADRX", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) is None
async def test_delisting_filing_picks_the_newest_match():
client = _sec_client(
_submissions(
["25", "8-K", "25-NSE", "15-12B"],
["2024-01-02", "2026-08-01", "2026-08-04", "2025-05-05"],
)
)
async with client as c:
found = await c.delisting_filing("0000712515")
assert found["form"] == "25-NSE"
assert found["filing_date"] == date(2026, 8, 4)
async def test_delisting_filing_returns_none_without_one():
client = _sec_client(_submissions(["10-K", "8-K"], ["2026-01-02", "2026-08-01"]))
async with client as c:
assert await c.delisting_filing("0000320193") is None
async def test_confirm_delisting_waits_before_spending_a_request(session: AsyncSession, monkeypatch):
"""A one-day gap is a weekend or a hiccup. Probing every stale symbol during a
market-data outage would be one SEC request per symbol per run."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
def _explode():
raise AssertionError("must not reach SEC before the stale threshold")
monkeypatch.setattr(ticker_service, "_sec_client_factory", _explode, raising=False)
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 10), today=date(2026, 8, 11)
) is None
# ...and no bars at all is an ingestion problem, not a delisting.
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=None, today=date(2026, 8, 11)
) is None
async def test_sec_confirmation_upgrades_a_manual_mark(session: AsyncSession, monkeypatch):
"""An operator's estimated date is a guess; Form 25 carries the real one."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
await ticker_service.mark_delisted(session, "EA", delisted_on=date(2026, 8, 11))
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["25-NSE"], ["2026-08-04"])),
raising=False,
)
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) == date(2026, 8, 14)
row = (await session.execute(select(Ticker).where(Ticker.symbol == "EA"))).scalar_one()
assert row.delisted_on == date(2026, 8, 14)
assert row.delisted_reason == ticker_service.REASON_FORM_25
async def test_a_confirmed_row_is_never_reprobed(session: AsyncSession, monkeypatch):
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
await ticker_service.mark_delisted(
session, "EA", delisted_on=date(2026, 8, 4),
reason=ticker_service.REASON_FORM_25,
)
def _explode():
raise AssertionError("a SEC-confirmed row must not cost another request")
monkeypatch.setattr(ticker_service, "_sec_client_factory", _explode, raising=False)
# Costs no request, and still reports the date so the caller knows this gap
# is explained and must not warn about it again.
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 9, 1)
) == date(2026, 8, 4)
async def test_ohlcv_priority_ordering_skips_delisted(session: AsyncSession):
"""Covers the one statement where active_only wraps a compound select."""
from app.scheduler import _get_ohlcv_priority_tickers
session.add_all([Ticker(symbol="AAPL"), Ticker(symbol="EA"), Ticker(symbol="MSFT")])
await session.commit()
await ticker_service.mark_delisted(session, "EA", delisted_on=date(2026, 8, 4))
symbols = await _get_ohlcv_priority_tickers(session)
assert "EA" not in symbols
assert sorted(symbols) == ["AAPL", "MSFT"]
async def test_a_form_25_for_another_security_class_is_ignored(session: AsyncSession, monkeypatch):
"""Form 25 is per security class. An issuer delisting its notes, preferred or
warrants files one while the common keeps trading retiring the ticker on
that would remove an actively traded symbol from every signal."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(
_submissions(["25-NSE"], ["2026-08-04"]),
security="6.25% Notes due 2030",
),
raising=False,
)
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) is None
row = (await session.execute(select(Ticker).where(Ticker.symbol == "EA"))).scalar_one()
assert row.delisted_on is None
async def test_a_stale_historical_form_25_cannot_retire_a_symbol(session: AsyncSession, monkeypatch):
"""A 2019 filing for a long-gone class must not retire a symbol whose bars
ran until 2026 and must certainly not stamp 2019 as the date."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["25"], ["2019-03-01"])),
raising=False,
)
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) is None
async def test_form_15_alone_never_retires_a_symbol(session: AsyncSession, monkeypatch):
"""Form 15 ends a reporting obligation; it is not evidence trading stopped."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["15-12B", "15-12G"], ["2026-08-04", "2026-08-05"])),
raising=False,
)
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) is None
async def test_an_unreadable_form_25_fails_closed(session: AsyncSession, monkeypatch):
"""Pre-2009 filings have no primary_doc.xml. Unknown class must read as no."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["25"], ["2026-08-04"]), security=None),
raising=False,
)
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
) is None
@pytest.mark.parametrize(
"description,expected",
[
("Common Stock", True),
("Class A Common Stock, $0.01 par value", True),
("Common Shares, no par value", True),
("6.25% Notes due 2030", False),
("7.5% Series B Cumulative Preferred Stock", False),
("Warrants to purchase Common Stock", False),
("Depositary Shares each representing 1/1000th interest", False),
("", False),
],
)
def test_common_stock_classification(description: str, expected: bool):
from app.services.sec_client import _is_common_stock
assert _is_common_stock(description) is expected
async def test_prune_keeps_delisted_rows(session: AsyncSession, monkeypatch):
"""A prune must not destroy rows the delisting flow deliberately retained —
their price history is the whole reason those rows still exist."""
from app.services import ticker_universe_service as tus
session.add_all([Ticker(symbol="AAPL"), Ticker(symbol="EA"), Ticker(symbol="GONE")])
await session.commit()
await ticker_service.mark_delisted(session, "EA", delisted_on=date(2026, 8, 14))
async def fake_fetch(db, universe):
return ["AAPL"], "test"
monkeypatch.setattr(tus, "fetch_universe_symbols", fake_fetch)
summary = await tus.bootstrap_universe(session, "sp500", prune_missing=True)
remaining = sorted(t.symbol for t in await ticker_service.list_tickers(session))
assert remaining == ["AAPL", "EA"] # GONE pruned, EA protected
assert summary["deleted"] == 1
assert summary["kept_delisted"] == ["EA"]
async def test_a_future_effective_date_keeps_the_symbol_live(session: AsyncSession):
"""Form 25 is known ten days before removal takes effect. The symbol is still
trading in that window and must keep being scanned and ingested."""
session.add_all([Ticker(symbol="AAPL"), Ticker(symbol="EA")])
await session.commit()
await ticker_service.mark_delisted(session, "EA", delisted_on=date(2026, 8, 14))
def active(as_of: date) -> list[str]:
return ticker_service.active_only(select(Ticker.symbol), as_of=as_of)
before = (await session.execute(active(date(2026, 8, 11)))).scalars().all()
on_the_day = (await session.execute(active(date(2026, 8, 14)))).scalars().all()
after = (await session.execute(active(date(2026, 8, 15)))).scalars().all()
assert sorted(before) == ["AAPL", "EA"] # still trading
assert sorted(on_the_day) == ["AAPL"] # removal effective
assert sorted(after) == ["AAPL"]
async def test_the_pending_window_does_not_re_warn(session: AsyncSession, monkeypatch):
"""Between filing and effect the symbol is active but produces no bars. That
must not resurrect the daily staleness warning this flow exists to end."""
session.add(Ticker(symbol="EA", cik="0000712515"))
await session.commit()
monkeypatch.setattr(
ticker_service,
"_sec_client_factory",
lambda: _sec_client(_submissions(["25-NSE"], ["2026-08-04"])),
raising=False,
)
first = await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=date(2026, 8, 11)
)
assert first == date(2026, 8, 14)
def _explode():
raise AssertionError("must not re-probe a confirmed row")
monkeypatch.setattr(ticker_service, "_sec_client_factory", _explode, raising=False)
# Every later run inside the window still reports the delisting, so the
# caller keeps emitting "delisted" rather than "no new bars".
for day in (date(2026, 8, 12), date(2026, 8, 13)):
assert await ticker_service.confirm_delisting(
session, "EA", last_bar=date(2026, 8, 4), today=day
) == date(2026, 8, 14)