Compare commits
50
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6f1ee450f1 | ||
|
|
aa6cd5cac4 | ||
|
|
24482c62fe | ||
|
|
6fc82ae857 | ||
|
|
23fe39fd78 | ||
|
|
477aa4b2da | ||
|
|
e58d2bb2cf | ||
|
|
1ace6688dd | ||
|
|
07d864cf64 | ||
|
|
2435abacaf | ||
|
|
d29d603158 | ||
|
|
70157ccfc2 | ||
|
|
f49b422095 | ||
|
|
59ac108c90 | ||
|
|
e0f3d43efb | ||
|
|
4c0c0579f5 | ||
|
|
d1caac86b5 | ||
|
|
d431ee283d | ||
|
|
3a6900d45a | ||
|
|
7d703ea524 | ||
|
|
7bcdf77ef9 | ||
|
|
c8c660e63d | ||
|
|
f58f8b0818 | ||
|
|
862d1d536b | ||
|
|
5b4fdab85c | ||
|
|
a7aefa6fe7 | ||
|
|
d2a27d4a78 | ||
|
|
83c0555e52 | ||
|
|
019ca1342a | ||
|
|
49bf3b140e | ||
|
|
b0537ebe9a | ||
|
|
3df36a9bfb | ||
|
|
98793925ef | ||
|
|
2e083d3bb9 | ||
|
|
4135946cb3 | ||
|
|
3d42ca7241 | ||
|
|
77fa8b8c65 | ||
|
|
0e556d8a43 | ||
|
|
fae621475b | ||
|
|
e54f03cba6 | ||
|
|
921f3d06fb | ||
|
|
259001e419 | ||
|
|
ce8c60d957 | ||
|
|
361cfd7883 | ||
|
|
dba7ea739b | ||
|
|
34d6dda1ab | ||
|
|
7f944d718f | ||
|
|
eae4d34c06 | ||
|
|
ddc88b130b | ||
|
|
8be7635506 |
@@ -52,6 +52,10 @@ SEC_REQUEST_SPACING_SECONDS=0.2
|
||||
SEC_MAX_RETRIES=4
|
||||
SEC_REQUEST_TIMEOUT_SECONDS=30.0
|
||||
|
||||
# A5 read-only parity report archive. In production keep this outside the
|
||||
# rsync deployment tree, e.g. /var/lib/signal-platform/reports/fundamentals-parity.
|
||||
FUNDAMENTALS_PARITY_REPORT_DIR=reports/fundamentals-parity
|
||||
|
||||
# Regime Monitor — FRED (VIX + HY credit spreads). Free key: https://fred.stlouisfed.org/docs/api/api_key.html
|
||||
# Optional: without it the volatility (V1) and credit (C1) pillars show as n/a.
|
||||
FRED_API_KEY=
|
||||
|
||||
@@ -51,3 +51,6 @@ backtest_snapshots/
|
||||
reports/*.pkl
|
||||
reports/*.pk1
|
||||
reports/.cache/
|
||||
# Runtime A5 parity bundles are generated on the production server. Research
|
||||
# conclusions belong in docs/research, not as an ever-growing artifact archive.
|
||||
reports/fundamentals-parity/
|
||||
|
||||
@@ -133,7 +133,7 @@ indicators.
|
||||
|
||||
1. **OHLCV** — latest daily bars (Alpaca); new tickers backfill ~5 years.
|
||||
2. **Sentiment** — stale names that matter (top-pick feeders, watchlist, open paper, discovery net). Display context only; the activation gate is price-only.
|
||||
3. **Market Regime** + **Regime Monitor** — breadth/trend and the v2 risk thermometer; feed no trades.
|
||||
3. **Market Regime** + **Regime Monitor** — breadth/trend and the v3 risk thermometer; feed no trades.
|
||||
4. **Telegram alerts** — change-driven (regime-quadrant etc.); quiet days stay quiet. Setup alerts still fire on the near-close pipeline after the scan.
|
||||
|
||||
**Near-close** (~15:30 ET Mon–Fri) — the only full-universe qualifying observation:
|
||||
@@ -255,11 +255,18 @@ A systematic single-variable sweep (offline prod snapshot, production gate/rank/
|
||||
| ATR trail multiple {1.5–4.0} | **Keep 3.0** | Return+Sharpe peak; ≤2.0 whipsaws out the momentum right tail; ≥2.5 is a plateau |
|
||||
| SPY 200d-MA regime overlay (block entries / go flat) | **Reject** | Halves return (315%→138%) with zero drawdown benefit — the ATR trail already manages downside, and the filter blocks the recovery-phase entries that make the money |
|
||||
| Momentum lookback: 6-1, 3-1, 12-7 (Novy-Marx), composites | **Keep residual 12-1** | 6-1/3-1 rank-IC ≈ 0; 12-7 IC 0.045 / t 1.58 — weaker than residual 12-1 (0.055 / t 1.98) |
|
||||
| Selection cutoff {70, 75, 85, 90} × book size {10, 15, 20} | **Keep 80 × 10** | Monotonically worse in both directions from 80; the 10-slot cap never binds (<10 concurrent) |
|
||||
| Selection cutoff {70, 75, 85, 90} × book size {10, 15, 20} | **Keep cutoff 80; capacity reopened** | The older weekly replay favored 80 × 10, but its no-cap-pressure conclusion is superseded by 519 book-full rejections versus 472 trades under the current daily gate-reset control |
|
||||
| Position sizing: equal-weight, inverse-vol, risk-% sweep | **Keep 1% fixed-fractional** | See the inverse-vol warning below |
|
||||
| Post-stop re-entry: immediate, fixed 2–5 sessions, gate resets, confirmation filters | **Keep normal gate reset for the 10-position production book** | Sharpe 1.77 vs 1.67 immediate and 1.47 cooldown 5; rerun before changing portfolio capacity |
|
||||
| FIP path-smoothness as an in-book tie-breaker/filter | **Reject** (but see the lead below) | Non-monotonic across FIP quintiles within the qualified set; either half of a median split underperforms the full book — thinning the entry stream costs more compounding than the tilt returns |
|
||||
|
||||
> **Capacity correction (2026-08-05):** the table's older weekly conclusion
|
||||
> that the ten-slot cap never binds is superseded. Under the current daily
|
||||
> gate-reset Phase A control, 472 trades were admitted and 519 qualified entries
|
||||
> were rejected because the book was full (52.4% of admitted+blocked
|
||||
> opportunities). Cutoff 80 remains the signal setting; portfolio capacity is
|
||||
> reopened in the focused capacity-bracket study.
|
||||
|
||||
Two findings future sessions must not re-litigate:
|
||||
|
||||
- **The "inverse-vol sizing win" (July 2026) was mis-attributed — do not resurrect.** The diagnostic sized `notional = equity × 1% / vol_6m`, and the 20% notional cap bound on 95% of entries, so it actually measured "~5 positions × 20% notional each" — a concentration/risk-appetite bump economically equivalent to raising risk to 1.5%, not vol-managed sizing. Genuine inverse-vol sizing (risk budget × median-vol/vol) cuts max drawdown to −18.2% but costs ~58pp total return at flat Sharpe: a risk-preference trade, not edge.
|
||||
@@ -319,7 +326,7 @@ Corollaries: never let an unvalidated score gate setups; the outcome evaluator m
|
||||
- Activation gate — qualifies setups on a residual-momentum percentile floor (the actual selection), a headline gate-target R:R floor (prod: 2.0) and a 20% primary-target reach-probability floor (validated long-only edge)
|
||||
- Recommendation layer — directional confidence, conflict detection, per-target reach-probability
|
||||
- Paper trading — take a setup, mark-to-market vs. latest close, auto-close per the exit policy (default: 3x ATR trail with a 30-trading-day max hold; time / percent-trailing / target-stop selectable), realized track record + outcome evaluation
|
||||
- Market-regime guard + observational State/Warning monitor (fixed-basket breadth, VIX, credit, PIT fundamentals) with a manual chronological correction study
|
||||
- Market-regime guard + observational State/Warning monitor (fixed-basket breadth, VIX, credit level + impulse) with a manual chronological correction study
|
||||
- Telegram alerts (e.g. regime-quadrant changes)
|
||||
- User-curated watchlist (cap: 20), enriched with composite score, R:R and S/R summary
|
||||
- JWT auth with admin role, configurable registration, user access control
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
"""fundamental_snapshots.weighted_avg_diluted_shares — market-cap fallback
|
||||
|
||||
Revision ID: 027
|
||||
Revises: 026
|
||||
Create Date: 2026-07-24 00:00:00.000000
|
||||
|
||||
Multi-class issuers report the cover-page share count per share class. That is a
|
||||
dimensional fact and Company Facts is non-dimensional, so it is absent entirely:
|
||||
META has never tagged it, CMCSA stops in 2009, BRK-B in 2011, CHTR in 2016 (when
|
||||
the Time Warner Cable deal made it multi-class). `shares_outstanding` is
|
||||
therefore null for a large slice of the mega-cap universe, which silently removes
|
||||
both `market_cap_est` and `fcf_yield`.
|
||||
|
||||
The weighted-average diluted count is always present (EPS requires it) and is
|
||||
consolidated across classes. Measured against issuers where the true
|
||||
point-in-time count IS available, it lands within ~0.6%: GOOGL 0.9936, MRNA
|
||||
1.0045, AAPL 0.9974, MSFT 0.9978.
|
||||
|
||||
Stored as its own column rather than backfilled into `shares_outstanding`, so the
|
||||
point-in-time column keeps its strict meaning and the fallback stays an explicit,
|
||||
labelled read-time decision. Existing rows are null until a reparse.
|
||||
"""
|
||||
from typing import Sequence, Union
|
||||
|
||||
from alembic import op
|
||||
import sqlalchemy as sa
|
||||
|
||||
|
||||
revision: str = "027"
|
||||
down_revision: Union[str, None] = "026"
|
||||
branch_labels: Union[str, Sequence[str], None] = None
|
||||
depends_on: Union[str, Sequence[str], None] = None
|
||||
|
||||
|
||||
def upgrade() -> None:
|
||||
op.add_column(
|
||||
"fundamental_snapshots",
|
||||
sa.Column("weighted_avg_diluted_shares", sa.Float(), nullable=True),
|
||||
)
|
||||
|
||||
|
||||
def downgrade() -> None:
|
||||
op.drop_column("fundamental_snapshots", "weighted_avg_diluted_shares")
|
||||
@@ -0,0 +1,147 @@
|
||||
"""SEC filing retry queue and setup-quality gate
|
||||
|
||||
Revision ID: 028
|
||||
Revises: 027
|
||||
Create Date: 2026-08-03 00:00:00.000000
|
||||
"""
|
||||
from datetime import date, datetime, timezone
|
||||
import json
|
||||
from typing import Sequence, Union
|
||||
|
||||
from alembic import op
|
||||
import sqlalchemy as sa
|
||||
|
||||
|
||||
revision: str = "028"
|
||||
down_revision: Union[str, None] = "027"
|
||||
branch_labels: Union[str, Sequence[str], None] = None
|
||||
depends_on: Union[str, Sequence[str], None] = None
|
||||
|
||||
|
||||
def upgrade() -> None:
|
||||
op.create_table(
|
||||
"sec_filing_gaps",
|
||||
sa.Column("id", sa.Integer(), primary_key=True),
|
||||
sa.Column("cik", sa.String(length=10), nullable=False),
|
||||
sa.Column("accession", sa.String(length=25), nullable=False),
|
||||
sa.Column("form", sa.String(length=12), nullable=True),
|
||||
sa.Column("index_date", sa.Date(), nullable=True),
|
||||
sa.Column("reason", sa.String(length=64), nullable=False),
|
||||
sa.Column("coregistrant_ciks_json", sa.Text(), nullable=True),
|
||||
sa.Column("first_seen_at", sa.DateTime(timezone=True), nullable=False),
|
||||
sa.Column("last_attempted_at", sa.DateTime(timezone=True), nullable=False),
|
||||
sa.Column("escalated_at", sa.DateTime(timezone=True), nullable=True),
|
||||
sa.UniqueConstraint("accession", name="uq_sec_filing_gaps_accession"),
|
||||
)
|
||||
op.create_index("ix_sec_filing_gaps_cik", "sec_filing_gaps", ["cik"])
|
||||
_backfill_retry_queue()
|
||||
|
||||
|
||||
def downgrade() -> None:
|
||||
op.drop_index("ix_sec_filing_gaps_cik", table_name="sec_filing_gaps")
|
||||
op.drop_table("sec_filing_gaps")
|
||||
|
||||
|
||||
def _as_date(value) -> date | None:
|
||||
if isinstance(value, datetime):
|
||||
return value.date()
|
||||
if isinstance(value, date):
|
||||
return value
|
||||
if isinstance(value, str):
|
||||
try:
|
||||
return date.fromisoformat(value)
|
||||
except ValueError:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def _backfill_retry_queue() -> None:
|
||||
"""Materialize pre-queue promoted gaps once; runtime never scans history."""
|
||||
bind = op.get_bind()
|
||||
runs = sa.table(
|
||||
"data_import_runs",
|
||||
sa.column("source", sa.String()),
|
||||
sa.column("status", sa.String()),
|
||||
sa.column("validation_json", sa.Text()),
|
||||
sa.column("source_max_date", sa.Date()),
|
||||
sa.column("started_at", sa.DateTime(timezone=True)),
|
||||
)
|
||||
snapshots = sa.table(
|
||||
"fundamental_snapshots",
|
||||
sa.column("cik", sa.String()),
|
||||
sa.column("accession", sa.String()),
|
||||
sa.column("filed_date", sa.Date()),
|
||||
)
|
||||
gaps = sa.table(
|
||||
"sec_filing_gaps",
|
||||
sa.column("cik", sa.String()),
|
||||
sa.column("accession", sa.String()),
|
||||
sa.column("form", sa.String()),
|
||||
sa.column("index_date", sa.Date()),
|
||||
sa.column("reason", sa.String()),
|
||||
sa.column("coregistrant_ciks_json", sa.Text()),
|
||||
sa.column("first_seen_at", sa.DateTime(timezone=True)),
|
||||
sa.column("last_attempted_at", sa.DateTime(timezone=True)),
|
||||
sa.column("escalated_at", sa.DateTime(timezone=True)),
|
||||
)
|
||||
|
||||
snapshot_rows = bind.execute(
|
||||
sa.select(snapshots.c.cik, snapshots.c.accession, snapshots.c.filed_date)
|
||||
).all()
|
||||
resolved_accessions = {row.accession for row in snapshot_rows}
|
||||
latest_filed_by_cik: dict[str, date] = {}
|
||||
for row in snapshot_rows:
|
||||
if row.filed_date is not None:
|
||||
current = latest_filed_by_cik.get(row.cik)
|
||||
if current is None or row.filed_date > current:
|
||||
latest_filed_by_cik[row.cik] = row.filed_date
|
||||
|
||||
audit_rows = bind.execute(
|
||||
sa.select(
|
||||
runs.c.validation_json,
|
||||
runs.c.source_max_date,
|
||||
runs.c.started_at,
|
||||
).where(
|
||||
runs.c.source == "sec_facts",
|
||||
runs.c.status == "promoted",
|
||||
runs.c.validation_json.is_not(None),
|
||||
)
|
||||
).all()
|
||||
now = datetime.now(timezone.utc)
|
||||
candidates: dict[str, dict] = {}
|
||||
for audit in audit_rows:
|
||||
try:
|
||||
summary = json.loads(audit.validation_json)
|
||||
except (TypeError, ValueError):
|
||||
continue
|
||||
if not isinstance(summary, dict):
|
||||
continue
|
||||
for item in summary.get("missing_xbrl") or []:
|
||||
accession = item.get("accession")
|
||||
raw_cik = item.get("cik")
|
||||
if not accession or raw_cik is None or accession in resolved_accessions:
|
||||
continue
|
||||
cik = str(raw_cik).zfill(10)
|
||||
index_date = _as_date(item.get("index_date")) or _as_date(
|
||||
audit.source_max_date
|
||||
)
|
||||
later_filed = latest_filed_by_cik.get(cik)
|
||||
if index_date is not None and later_filed is not None and later_filed > index_date:
|
||||
continue
|
||||
first_seen = audit.started_at or now
|
||||
existing = candidates.get(accession)
|
||||
if existing is not None and existing["first_seen_at"] <= first_seen:
|
||||
continue
|
||||
candidates[accession] = {
|
||||
"cik": cik,
|
||||
"accession": accession,
|
||||
"form": item.get("form"),
|
||||
"index_date": index_date,
|
||||
"reason": item.get("reason") or "not_in_companyfacts",
|
||||
"coregistrant_ciks_json": json.dumps(item.get("coregistrants") or []),
|
||||
"first_seen_at": first_seen,
|
||||
"last_attempted_at": first_seen,
|
||||
"escalated_at": None,
|
||||
}
|
||||
if candidates:
|
||||
op.bulk_insert(gaps, list(candidates.values()))
|
||||
@@ -61,6 +61,10 @@ class Settings(BaseSettings):
|
||||
sec_max_retries: int = 4
|
||||
sec_request_timeout_seconds: float = 30.0
|
||||
|
||||
# A5 read-only comparison artifacts. Production must keep this outside the
|
||||
# rsync deployment tree so the 5-7 day review window survives deploys.
|
||||
fundamentals_parity_report_dir: str = "reports/fundamentals-parity"
|
||||
|
||||
# Regime Monitor — FRED (VIX level + HY credit spreads). Optional: without it
|
||||
# the volatility (P5) and credit-spread (F2) signals are reported as n/a.
|
||||
fred_api_key: str = ""
|
||||
|
||||
@@ -17,6 +17,7 @@ from app.models.regime_snapshot import RegimeSnapshot
|
||||
from app.models.benchmark_price import BenchmarkPrice
|
||||
from app.models.signal_context_snapshot import SignalContextSnapshot
|
||||
from app.models.system_event import SystemEvent
|
||||
from app.models.sec_filing_gap import SecFilingGap
|
||||
|
||||
__all__ = [
|
||||
"Ticker",
|
||||
@@ -40,4 +41,5 @@ __all__ = [
|
||||
"BenchmarkPrice",
|
||||
"SignalContextSnapshot",
|
||||
"SystemEvent",
|
||||
"SecFilingGap",
|
||||
]
|
||||
|
||||
@@ -10,7 +10,8 @@ class DataImportRun(Base):
|
||||
"""One row per bulk-import attempt (SEC facts / Dolt earnings / Dolt stocks).
|
||||
|
||||
Lean audit record for the batch import framework: every attempt is logged,
|
||||
whether it promoted, was a ``no_op`` (unchanged revision), or ``failed``.
|
||||
whether it promoted, was a ``no_op`` (unchanged revision), was ``deferred``
|
||||
for an expected retry, or ``failed``.
|
||||
``row_counts`` and ``validation`` hold JSON strings (repo convention — see
|
||||
``fundamental_data.unavailable_fields_json``), not JSONB; the validation
|
||||
blob carries reconciliation/discrepancy summaries so no separate conflicts
|
||||
@@ -28,7 +29,7 @@ class DataImportRun(Base):
|
||||
source: Mapped[str] = mapped_column(String(32), nullable=False)
|
||||
# Dolt commit hash, or SEC archive SHA-256. Null until known.
|
||||
revision: Mapped[str | None] = mapped_column(String(64), nullable=True)
|
||||
# running | validated | promoted | no_op | failed
|
||||
# running | validated | promoted | no_op | deferred | failed
|
||||
status: Mapped[str] = mapped_column(String(16), nullable=False)
|
||||
source_max_date: Mapped[date | None] = mapped_column(Date, nullable=True)
|
||||
row_counts_json: Mapped[str | None] = mapped_column(Text, nullable=True)
|
||||
@@ -39,4 +40,5 @@ class DataImportRun(Base):
|
||||
completed_at: Mapped[datetime | None] = mapped_column(
|
||||
DateTime(timezone=True), nullable=True
|
||||
)
|
||||
# Failure detail, or the non-error reason when status is deferred.
|
||||
error_details: Mapped[str | None] = mapped_column(Text, nullable=True)
|
||||
|
||||
@@ -12,8 +12,10 @@ class FundamentalSnapshot(Base):
|
||||
Keyed by issuer (CIK), not ticker — multi-class issuers (GOOG/GOOGL) share
|
||||
one CIK and one set of fundamentals; the ``tickers.cik`` column is the only
|
||||
join point. Amendments are retained: every accession is a distinct immutable
|
||||
row, and readers pick the newest valid ``accepted_at`` per
|
||||
(cik, fiscal_year, fiscal_period) at read time — no flags, no mutation.
|
||||
row, and readers resolve (cik, fiscal_year, fiscal_period) at read time by
|
||||
taking the newest ``accepted_at`` **per field**, falling back to the newest
|
||||
accession that actually reports one — a partial amendment (a 10-K/A adding
|
||||
Part III reports no financial facts) must not blank the period — no flags, no mutation.
|
||||
|
||||
**Facts are stored as the filing reports them, never as derived quarters.**
|
||||
Duration facts (revenue, net_income, operating_income, diluted_eps, cfo,
|
||||
@@ -70,6 +72,12 @@ class FundamentalSnapshot(Base):
|
||||
# reported "as of" its own date, which can differ from period_end — store it
|
||||
# so market cap uses the right point-in-time count.
|
||||
shares_outstanding_date: Mapped[date | None] = mapped_column(Date, nullable=True)
|
||||
# Weighted-average diluted count for the filing's most recent quarter — the
|
||||
# market-cap fallback when the cover-page count is absent, which it always is
|
||||
# for multi-class issuers (per-class facts are dimensional, and companyfacts
|
||||
# is not). An average is not cumulative, so unlike the duration facts above
|
||||
# this is NOT a YTD value: it is the shortest-span fact ending at period_end.
|
||||
weighted_avg_diluted_shares: Mapped[float | None] = mapped_column(Float, nullable=True)
|
||||
|
||||
import_run_id: Mapped[int | None] = mapped_column(
|
||||
ForeignKey("data_import_runs.id", ondelete="SET NULL"), nullable=True
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
from datetime import date, datetime
|
||||
|
||||
from sqlalchemy import Date, DateTime, Index, String, Text, UniqueConstraint
|
||||
from sqlalchemy.orm import Mapped, mapped_column
|
||||
|
||||
from app.database import Base
|
||||
|
||||
|
||||
class SecFilingGap(Base):
|
||||
"""Active SEC filing that could not yet be reconstructed.
|
||||
|
||||
Rows form a small retry queue. Successful snapshot ingestion deletes the
|
||||
matching row; a later valid filing supersedes it. While a current row remains,
|
||||
tickers mapped to its CIK are not eligible for actionable trade setups.
|
||||
"""
|
||||
|
||||
__tablename__ = "sec_filing_gaps"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("accession", name="uq_sec_filing_gaps_accession"),
|
||||
Index("ix_sec_filing_gaps_cik", "cik"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
cik: Mapped[str] = mapped_column(String(10), nullable=False)
|
||||
accession: Mapped[str] = mapped_column(String(25), nullable=False)
|
||||
form: Mapped[str | None] = mapped_column(String(12), nullable=True)
|
||||
index_date: Mapped[date | None] = mapped_column(Date, nullable=True)
|
||||
reason: Mapped[str] = mapped_column(String(64), nullable=False)
|
||||
coregistrant_ciks_json: Mapped[str | None] = mapped_column(Text, nullable=True)
|
||||
first_seen_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
|
||||
last_attempted_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
|
||||
escalated_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
|
||||
@@ -13,6 +13,7 @@ from app.schemas.admin import (
|
||||
AlertConfigUpdate,
|
||||
CreateUserRequest,
|
||||
DataCleanupRequest,
|
||||
FundamentalsCutoverConfigUpdate,
|
||||
JobTriggerRequest,
|
||||
JobToggle,
|
||||
RecommendationConfigUpdate,
|
||||
@@ -137,6 +138,27 @@ async def list_settings(
|
||||
)
|
||||
|
||||
|
||||
@router.get("/admin/settings/fundamentals-cutover", response_model=APIEnvelope)
|
||||
async def get_fundamentals_cutover_settings(
|
||||
_admin: User = Depends(require_admin),
|
||||
db: AsyncSession = Depends(get_db),
|
||||
):
|
||||
config = await admin_service.get_fundamentals_cutover_config(db)
|
||||
return APIEnvelope(status="success", data=config)
|
||||
|
||||
|
||||
@router.put("/admin/settings/fundamentals-cutover", response_model=APIEnvelope)
|
||||
async def update_fundamentals_cutover_settings(
|
||||
body: FundamentalsCutoverConfigUpdate,
|
||||
_admin: User = Depends(require_admin),
|
||||
db: AsyncSession = Depends(get_db),
|
||||
):
|
||||
config = await admin_service.update_fundamentals_cutover_config(
|
||||
db, body.enabled
|
||||
)
|
||||
return APIEnvelope(status="success", data=config)
|
||||
|
||||
|
||||
@router.get("/admin/settings/recommendations", response_model=APIEnvelope)
|
||||
async def get_recommendation_settings(
|
||||
_admin: User = Depends(require_admin),
|
||||
@@ -453,6 +475,36 @@ async def toggle_job(
|
||||
)
|
||||
|
||||
|
||||
@router.get("/admin/fundamentals-parity", response_model=APIEnvelope)
|
||||
async def get_fundamentals_parity_report(
|
||||
_admin: User = Depends(require_admin),
|
||||
):
|
||||
"""Latest read-only A5 source/score comparison, or null before first run."""
|
||||
return APIEnvelope(
|
||||
status="success", data=admin_service.get_fundamentals_parity_report()
|
||||
)
|
||||
|
||||
|
||||
@router.get("/admin/fundamentals-parity/csv", response_model=APIEnvelope)
|
||||
async def get_fundamentals_parity_csv(
|
||||
_admin: User = Depends(require_admin),
|
||||
):
|
||||
"""Latest flattened A5 report for an authenticated browser download."""
|
||||
artifact = admin_service.get_fundamentals_parity_csv()
|
||||
data = None if artifact is None else {"filename": artifact[0], "content": artifact[1]}
|
||||
return APIEnvelope(status="success", data=data)
|
||||
|
||||
|
||||
@router.get("/admin/fundamentals-parity/json", response_model=APIEnvelope)
|
||||
async def get_fundamentals_parity_json(
|
||||
_admin: User = Depends(require_admin),
|
||||
):
|
||||
"""Canonical A5 JSON artifact for an authenticated browser download."""
|
||||
artifact = admin_service.get_fundamentals_parity_json()
|
||||
data = None if artifact is None else {"filename": artifact[0], "content": artifact[1]}
|
||||
return APIEnvelope(status="success", data=data)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# System events (operational warnings / errors)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@@ -10,6 +10,7 @@ from app.schemas.common import APIEnvelope
|
||||
from app.schemas.fundamental import FundamentalResponse
|
||||
from app.services.fundamental_service import get_fundamental
|
||||
from app.services.fundamentals_api_service import build_fundamentals_v1
|
||||
from app.services import fundamentals_quality_service
|
||||
|
||||
router = APIRouter(tags=["fundamentals"])
|
||||
|
||||
@@ -34,6 +35,7 @@ async def read_fundamentals(
|
||||
"""Get latest fundamental data for a symbol (legacy fields + additive v1)."""
|
||||
record = await get_fundamental(db, symbol)
|
||||
v1 = await build_fundamentals_v1(db, symbol)
|
||||
quality = await fundamentals_quality_service.ticker_quality(db, symbol)
|
||||
|
||||
legacy: dict = {}
|
||||
if record is not None:
|
||||
@@ -47,5 +49,12 @@ async def read_fundamentals(
|
||||
unavailable_fields=_parse_unavailable_fields(record.unavailable_fields_json),
|
||||
)
|
||||
|
||||
data = FundamentalResponse(symbol=symbol.strip().upper(), **legacy, **v1)
|
||||
data = FundamentalResponse(
|
||||
symbol=symbol.strip().upper(),
|
||||
setup_eligible=quality.eligible,
|
||||
setup_block_code=quality.code,
|
||||
setup_block_reason=quality.message,
|
||||
**legacy,
|
||||
**v1,
|
||||
)
|
||||
return APIEnvelope(status="success", data=data.model_dump())
|
||||
|
||||
+176
-14
@@ -40,8 +40,15 @@ from app.services import (
|
||||
sentiment_service,
|
||||
settings_store,
|
||||
shadow_book_service,
|
||||
fundamentals_parity_service,
|
||||
fundamental_data_refresh_service,
|
||||
)
|
||||
from app.services.data_import import (
|
||||
STATUS_DEFERRED,
|
||||
STATUS_FAILED,
|
||||
SourceImporter,
|
||||
run_import,
|
||||
)
|
||||
from app.services.data_import import STATUS_FAILED, SourceImporter, run_import
|
||||
from app.services.dolt_earnings_importer import DoltEarningsImporter
|
||||
from app.services.sec_fundamentals_importer import SecFundamentalsImporter
|
||||
from app.services.alert_service import dispatch_alerts
|
||||
@@ -98,6 +105,7 @@ _JOB_NAMES = [
|
||||
"fundamental_collector",
|
||||
"dolt_earnings_import",
|
||||
"sec_fundamentals_import",
|
||||
"fundamentals_parity_report",
|
||||
"rr_scanner",
|
||||
"ticker_universe_sync",
|
||||
"alerts",
|
||||
@@ -504,13 +512,14 @@ async def collect_ohlcv(
|
||||
job_name: str = "data_collector",
|
||||
*,
|
||||
refetch_days: int = 0,
|
||||
refresh_sr: bool = True,
|
||||
) -> None:
|
||||
"""Fetch latest daily OHLCV for all tracked tickers.
|
||||
|
||||
Uses AlpacaOHLCVProvider. Processes each ticker independently.
|
||||
On rate limit, records last successful ticker for resume.
|
||||
Start date is resolved by ingestion progress:
|
||||
- existing ticker: resume from last_ingested_date + 1
|
||||
- existing ticker: overlap last_ingested_date so partial bars refresh
|
||||
- new ticker: backfill the configured history window
|
||||
|
||||
``full_backfill`` forces every ticker to re-fetch the full
|
||||
@@ -572,6 +581,7 @@ async def collect_ohlcv(
|
||||
try:
|
||||
result = await ingestion_service.fetch_and_ingest(
|
||||
db, provider, symbol, start_date=backfill_start, end_date=end_date,
|
||||
refresh_sr=refresh_sr,
|
||||
)
|
||||
_last_successful[job_name] = symbol
|
||||
processed += 1
|
||||
@@ -611,6 +621,11 @@ async def collect_ohlcv(
|
||||
_runtime_finish(job_name, "error", processed=processed, total=total, message=str(exc))
|
||||
|
||||
|
||||
async def collect_ohlcv_for_scan() -> None:
|
||||
"""Near-close fetch; the scanner immediately rebuilds S/R per ticker."""
|
||||
await collect_ohlcv(refresh_sr=False)
|
||||
|
||||
|
||||
async def backfill_ohlcv() -> None:
|
||||
"""Deep historical backfill: re-fetch the full ``settings.ohlcv_history_days``
|
||||
window for every ticker, ignoring incremental resume.
|
||||
@@ -650,7 +665,7 @@ async def run_shadow_book() -> None:
|
||||
if not await _is_job_enabled(db, job_name):
|
||||
_log_event(logging.INFO, "job_skipped", job=job_name, reason="disabled")
|
||||
_runtime_finish(job_name, "skipped", processed=0, total=1, message="Disabled")
|
||||
return
|
||||
return False
|
||||
if not await shadow_book_service.is_enabled(db):
|
||||
_log_event(logging.INFO, "job_skipped", job=job_name, reason="not enabled in settings")
|
||||
_runtime_finish(job_name, "skipped", processed=0, total=1, message="Not enabled")
|
||||
@@ -823,6 +838,22 @@ async def collect_fundamentals() -> None:
|
||||
_log_event(logging.INFO, "job_skipped", job=job_name, reason="disabled")
|
||||
_runtime_finish(job_name, "skipped", processed=0, total=0, message="Disabled")
|
||||
return
|
||||
if await fundamental_data_refresh_service.is_enabled(db):
|
||||
message = "SEC + Dolt fundamentals cutover is active"
|
||||
_log_event(
|
||||
logging.INFO,
|
||||
"job_skipped",
|
||||
job=job_name,
|
||||
reason="sec_dolt_cutover_active",
|
||||
)
|
||||
_runtime_finish(
|
||||
job_name,
|
||||
"skipped",
|
||||
processed=0,
|
||||
total=0,
|
||||
message=message,
|
||||
)
|
||||
return
|
||||
|
||||
symbols = await _get_fundamental_priority_tickers(db)
|
||||
if not symbols:
|
||||
@@ -922,8 +953,13 @@ async def collect_fundamentals() -> None:
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
async def _run_shadow_import(job_name: str, importer: SourceImporter) -> None:
|
||||
"""Run one source importer and surface its audit result in Admin → Jobs."""
|
||||
async def _run_shadow_import(job_name: str, importer: SourceImporter) -> bool:
|
||||
"""Run an importer and return whether its scheduled job was enabled.
|
||||
|
||||
The SEC wrapper uses the return value to run its activated local cache step
|
||||
after deferred, failed, no-op, promoted, or source-locked attempts while honoring
|
||||
the job-level disable switch.
|
||||
"""
|
||||
_log_event(logging.INFO, "job_start", job=job_name)
|
||||
_runtime_start(job_name, total=1)
|
||||
|
||||
@@ -939,15 +975,20 @@ async def _run_shadow_import(job_name: str, importer: SourceImporter) -> None:
|
||||
message = "Another import for this source is already running"
|
||||
_log_event(logging.INFO, "job_skipped", job=job_name, reason="source_locked")
|
||||
_runtime_finish(job_name, "skipped", processed=0, total=1, message=message)
|
||||
return
|
||||
return True
|
||||
|
||||
revision = f" · {run.revision[:12]}" if run.revision else ""
|
||||
message = f"{run.status}{revision}"
|
||||
if run.status == STATUS_DEFERRED:
|
||||
message = run.error_details or message
|
||||
_log_event(logging.INFO, "job_deferred", job=job_name, message=message)
|
||||
_runtime_finish(job_name, "deferred", processed=0, total=1, message=message)
|
||||
return True
|
||||
if run.status == STATUS_FAILED:
|
||||
message = run.error_details or message
|
||||
_log_event(logging.ERROR, "job_error", job=job_name, message=message)
|
||||
_runtime_finish(job_name, "error", processed=0, total=1, message=message)
|
||||
return
|
||||
return True
|
||||
|
||||
_log_event(
|
||||
logging.INFO,
|
||||
@@ -957,6 +998,7 @@ async def _run_shadow_import(job_name: str, importer: SourceImporter) -> None:
|
||||
revision=run.revision,
|
||||
)
|
||||
_runtime_finish(job_name, "completed", processed=1, total=1, message=message)
|
||||
return True
|
||||
except asyncio.CancelledError:
|
||||
_runtime_finish(job_name, "error", processed=0, total=1, message="Cancelled")
|
||||
raise
|
||||
@@ -969,6 +1011,7 @@ async def _run_shadow_import(job_name: str, importer: SourceImporter) -> None:
|
||||
message=str(exc),
|
||||
)
|
||||
_runtime_finish(job_name, "error", processed=0, total=1, message=str(exc))
|
||||
return True
|
||||
|
||||
|
||||
async def run_dolt_earnings_import() -> None:
|
||||
@@ -977,8 +1020,110 @@ async def run_dolt_earnings_import() -> None:
|
||||
|
||||
|
||||
async def run_sec_fundamentals_import() -> None:
|
||||
"""Import tracked-universe SEC facts in shadow."""
|
||||
await _run_shadow_import("sec_fundamentals_import", SecFundamentalsImporter())
|
||||
"""Import SEC facts, then run the activated local compat-cache refresh.
|
||||
|
||||
The refresh is deliberately separate from the network import result. Once
|
||||
activated it therefore still runs from stored snapshots/earnings/prices when
|
||||
SEC is unavailable, unchanged, or another SEC import owns the source lock.
|
||||
"""
|
||||
job_name = "sec_fundamentals_import"
|
||||
job_enabled = await _run_shadow_import(job_name, SecFundamentalsImporter())
|
||||
if not job_enabled:
|
||||
return
|
||||
|
||||
try:
|
||||
async with async_session_factory() as db:
|
||||
summary = await fundamental_data_refresh_service.refresh_if_enabled(db)
|
||||
except asyncio.CancelledError:
|
||||
_runtime_finish(
|
||||
job_name, "error", processed=0, total=1, message="Cancelled"
|
||||
)
|
||||
raise
|
||||
except Exception as exc:
|
||||
message = f"Local fundamental_data refresh failed: {exc}"
|
||||
_log_event(
|
||||
logging.ERROR,
|
||||
"fundamental_data_refresh_error",
|
||||
job=job_name,
|
||||
error_type=type(exc).__name__,
|
||||
message=str(exc),
|
||||
)
|
||||
_runtime_finish(job_name, "error", processed=0, total=1, message=message)
|
||||
return
|
||||
|
||||
if not summary["enabled"]:
|
||||
_log_event(
|
||||
logging.INFO,
|
||||
"fundamental_data_refresh_skipped",
|
||||
job=job_name,
|
||||
reason="cutover_disabled",
|
||||
setting=fundamental_data_refresh_service.ACTIVATION_KEY,
|
||||
)
|
||||
return
|
||||
|
||||
_log_event(
|
||||
logging.INFO,
|
||||
"fundamental_data_refresh_complete",
|
||||
job=job_name,
|
||||
**summary,
|
||||
)
|
||||
runtime = get_job_runtime_snapshot(job_name)
|
||||
if runtime.get("status") == "completed":
|
||||
import_message = runtime.get("message") or "import completed"
|
||||
cache_message = (
|
||||
f"cache {summary['refreshed']} · "
|
||||
f"{summary['score_inputs_changed']} score inputs changed"
|
||||
)
|
||||
_runtime_finish(
|
||||
job_name,
|
||||
"completed",
|
||||
processed=1,
|
||||
total=1,
|
||||
message=f"{import_message} · {cache_message}",
|
||||
)
|
||||
|
||||
|
||||
async def run_fundamentals_parity_report() -> None:
|
||||
"""Generate the A5 comparison bundle without mutating live fundamentals/scores."""
|
||||
job_name = "fundamentals_parity_report"
|
||||
_log_event(logging.INFO, "job_start", job=job_name)
|
||||
_runtime_start(job_name, total=1)
|
||||
try:
|
||||
async with async_session_factory() as db:
|
||||
if not await _is_job_enabled(db, job_name):
|
||||
_runtime_finish(
|
||||
job_name, "skipped", processed=0, total=1, message="Disabled"
|
||||
)
|
||||
return
|
||||
report, artifacts = await fundamentals_parity_service.generate_and_store(
|
||||
db, settings.fundamentals_parity_report_dir
|
||||
)
|
||||
summary = report["summary"]
|
||||
message = (
|
||||
f"{summary['universe_count']} tickers · "
|
||||
f"{summary['fundamental_score_material_changes']} material score changes"
|
||||
)
|
||||
_runtime_finish(job_name, "completed", processed=1, total=1, message=message)
|
||||
_log_event(
|
||||
logging.INFO,
|
||||
"job_complete",
|
||||
job=job_name,
|
||||
generated_at=report["generated_at"],
|
||||
json_path=artifacts["json"],
|
||||
csv_path=artifacts["csv"],
|
||||
)
|
||||
except asyncio.CancelledError:
|
||||
_runtime_finish(job_name, "error", processed=0, total=1, message="Cancelled")
|
||||
raise
|
||||
except Exception as exc:
|
||||
_runtime_finish(job_name, "error", processed=0, total=1, message=str(exc))
|
||||
_log_event(
|
||||
logging.ERROR,
|
||||
"job_error",
|
||||
job=job_name,
|
||||
error_type=type(exc).__name__,
|
||||
message=str(exc),
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -1358,8 +1503,8 @@ _DAILY_PIPELINE_STEPS = [
|
||||
("alerts", "dispatch_alerts_job"),
|
||||
]
|
||||
|
||||
# Near-close (~15:30 ET Mon–Fri): refresh in-progress day-t bars (already how
|
||||
# the intraday pipeline keeps the dashboard live), then the only daily
|
||||
# Near-close (~15:30 ET Mon–Fri): refresh in-progress day-t bars (incremental
|
||||
# ingestion overlaps the latest stored session), then the only daily
|
||||
# qualifying R:R scan, then Telegram immediately so manual fills can still hit
|
||||
# MOC cutoffs (~15:50/15:55). Under a 15-minute delayed SIP feed a 15:30 scan
|
||||
# may see ~15:15 prices — immaterial for a 12-1 momentum signal.
|
||||
@@ -1370,7 +1515,7 @@ _DAILY_PIPELINE_STEPS = [
|
||||
_NEAR_CLOSE_PIPELINE_STEPS = [
|
||||
# Must land today's in-progress bar (~20 min behind live), or the scan falls
|
||||
# back to the previous close and execution degrades to the stale_close floor.
|
||||
("data_collector", "collect_ohlcv"),
|
||||
("data_collector", "collect_ohlcv_for_scan"),
|
||||
("rr_scanner", "scan_rr"),
|
||||
# Straight after the scan so shadow entries mark at the same near-close
|
||||
# prices the discretionary book is looking at.
|
||||
@@ -1521,9 +1666,11 @@ SCHEDULE_DEFAULTS: dict[str, str] = {
|
||||
"schedule_timezone": "America/New_York",
|
||||
# Morning data/display refresh (no qualifying R:R scan).
|
||||
"schedule_daily_pipeline_cron": "0 2 * * *",
|
||||
# Shadow source imports. They never write legacy fundamental_data before A5.
|
||||
# Bulk source imports. The SEC job writes the legacy compat cache only after
|
||||
# the explicit, default-off A5 cutover setting is enabled.
|
||||
"schedule_dolt_earnings_cron": "30 2 * * *",
|
||||
"schedule_sec_fundamentals_cron": "0 4 * * *",
|
||||
"schedule_fundamentals_parity_cron": "30 5 * * *",
|
||||
# Fetch in-progress bars → scan → Telegram (manual MOC window).
|
||||
"schedule_near_close_pipeline_cron": "30 15 * * mon-fri",
|
||||
# Fetch final bars → outcome eval (must not run on the partial near-close bar).
|
||||
@@ -1539,6 +1686,7 @@ _CRON_JOBS: dict[str, str] = {
|
||||
"daily_pipeline": "schedule_daily_pipeline_cron",
|
||||
"dolt_earnings_import": "schedule_dolt_earnings_cron",
|
||||
"sec_fundamentals_import": "schedule_sec_fundamentals_cron",
|
||||
"fundamentals_parity_report": "schedule_fundamentals_parity_cron",
|
||||
"near_close_pipeline": "schedule_near_close_pipeline_cron",
|
||||
"after_close_pipeline": "schedule_after_close_pipeline_cron",
|
||||
"intraday_pipeline": "schedule_intraday_pipeline_cron",
|
||||
@@ -1642,7 +1790,18 @@ def configure_scheduler(schedule_config: dict[str, str] | None = None) -> None:
|
||||
"schedule_sec_fundamentals_cron",
|
||||
),
|
||||
id="sec_fundamentals_import",
|
||||
name="SEC Fundamentals Import (shadow)",
|
||||
name="SEC Fundamentals Import",
|
||||
replace_existing=True,
|
||||
)
|
||||
scheduler.add_job(
|
||||
run_fundamentals_parity_report,
|
||||
_cron_trigger(
|
||||
cfg["schedule_fundamentals_parity_cron"],
|
||||
tz,
|
||||
"schedule_fundamentals_parity_cron",
|
||||
),
|
||||
id="fundamentals_parity_report",
|
||||
name="Fundamentals Parity Report (read-only)",
|
||||
replace_existing=True,
|
||||
)
|
||||
scheduler.add_job(
|
||||
@@ -1720,6 +1879,9 @@ def configure_scheduler(schedule_config: dict[str, str] | None = None) -> None:
|
||||
},
|
||||
dolt_earnings_import={"cron": cfg["schedule_dolt_earnings_cron"]},
|
||||
sec_fundamentals_import={"cron": cfg["schedule_sec_fundamentals_cron"]},
|
||||
fundamentals_parity_report={
|
||||
"cron": cfg["schedule_fundamentals_parity_cron"]
|
||||
},
|
||||
near_close_pipeline={
|
||||
"cron": cfg["schedule_near_close_pipeline_cron"],
|
||||
"steps": [name for name, _ in _NEAR_CLOSE_PIPELINE_STEPS],
|
||||
|
||||
@@ -73,11 +73,19 @@ class ActivationConfigUpdate(BaseModel):
|
||||
exclude_neutral: bool | None = None
|
||||
|
||||
|
||||
class FundamentalsCutoverConfigUpdate(BaseModel):
|
||||
"""Switch the legacy fundamentals cache from quota APIs to SEC/Dolt."""
|
||||
enabled: bool
|
||||
|
||||
|
||||
class ScheduleConfigUpdate(BaseModel):
|
||||
"""Cron schedule for the pipelines + fundamentals. Crons are 5-field
|
||||
(min hour dom month dow); timezone is an IANA name (e.g. America/New_York)."""
|
||||
schedule_timezone: str | None = Field(default=None, max_length=64)
|
||||
schedule_daily_pipeline_cron: str | None = Field(default=None, max_length=120)
|
||||
schedule_dolt_earnings_cron: str | None = Field(default=None, max_length=120)
|
||||
schedule_sec_fundamentals_cron: str | None = Field(default=None, max_length=120)
|
||||
schedule_fundamentals_parity_cron: str | None = Field(default=None, max_length=120)
|
||||
schedule_near_close_pipeline_cron: str | None = Field(default=None, max_length=120)
|
||||
schedule_after_close_pipeline_cron: str | None = Field(default=None, max_length=120)
|
||||
schedule_intraday_pipeline_cron: str | None = Field(default=None, max_length=120)
|
||||
|
||||
@@ -26,6 +26,7 @@ class MetricItem(BaseModel):
|
||||
industry: MetricIndustry | None = None
|
||||
period_end: str | None = None
|
||||
filed_date: str | None = None
|
||||
caveat: str | None = None
|
||||
source: str = "sec"
|
||||
|
||||
|
||||
@@ -90,3 +91,6 @@ class FundamentalResponse(BaseModel):
|
||||
metrics: list[MetricItem] | None = None
|
||||
valuation: Valuation | None = None
|
||||
reads: FundamentalsReads | None = None
|
||||
setup_eligible: bool = True
|
||||
setup_block_code: str | None = None
|
||||
setup_block_reason: str | None = None
|
||||
|
||||
@@ -53,3 +53,7 @@ class PaperTradeResponse(BaseModel):
|
||||
# when the trailing exit policy is active.
|
||||
trailing_stop: float | None = None
|
||||
trailing_distance_pct: float | None = None
|
||||
# Trading sessions represented by post-entry OHLCV bars. These are populated
|
||||
# only while the active exit policy has a max-hold rule.
|
||||
sessions_held: int | None = None
|
||||
sessions_remaining: int | None = None
|
||||
|
||||
@@ -17,7 +17,7 @@ from app.models.settings import SystemSetting
|
||||
from app.models.ticker import Ticker
|
||||
from app.models.trade_setup import TradeSetup
|
||||
from app.models.user import User
|
||||
from app.services import settings_store
|
||||
from app.services import fundamental_data_refresh_service, settings_store
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -159,6 +159,28 @@ async def update_setting(db: AsyncSession, key: str, value: str) -> SystemSettin
|
||||
return setting
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Fundamentals source cutover
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
async def get_fundamentals_cutover_config(db: AsyncSession) -> dict[str, bool]:
|
||||
"""Return the explicit A5 cache-cutover switch (default off)."""
|
||||
return {"enabled": await fundamental_data_refresh_service.is_enabled(db)}
|
||||
|
||||
|
||||
async def update_fundamentals_cutover_config(
|
||||
db: AsyncSession, enabled: bool
|
||||
) -> dict[str, bool]:
|
||||
"""Activate or pause SEC/Dolt writes to the legacy fundamentals cache."""
|
||||
await settings_store.upsert_setting(
|
||||
db,
|
||||
fundamental_data_refresh_service.ACTIVATION_KEY,
|
||||
"true" if enabled else "false",
|
||||
)
|
||||
await db.commit()
|
||||
return await get_fundamentals_cutover_config(db)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Activation thresholds
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -614,6 +636,7 @@ VALID_JOB_NAMES = {
|
||||
"fundamental_collector",
|
||||
"dolt_earnings_import",
|
||||
"sec_fundamentals_import",
|
||||
"fundamentals_parity_report",
|
||||
"rr_scanner",
|
||||
"ticker_universe_sync",
|
||||
"outcome_evaluator",
|
||||
@@ -636,7 +659,8 @@ JOB_LABELS = {
|
||||
"sentiment_collector": "Sentiment Collector",
|
||||
"fundamental_collector": "Fundamental Collector",
|
||||
"dolt_earnings_import": "Dolt Earnings Import (shadow)",
|
||||
"sec_fundamentals_import": "SEC Fundamentals Import (shadow)",
|
||||
"sec_fundamentals_import": "SEC Fundamentals Import",
|
||||
"fundamentals_parity_report": "Fundamentals Parity Report (read-only)",
|
||||
"rr_scanner": "R:R Scanner",
|
||||
"ticker_universe_sync": "Ticker Universe Sync",
|
||||
"outcome_evaluator": "Outcome Evaluator",
|
||||
@@ -775,3 +799,30 @@ async def toggle_job(db: AsyncSession, job_name: str, enabled: bool) -> SystemSe
|
||||
|
||||
key = f"job_{job_name}_enabled"
|
||||
return await update_setting(db, key, str(enabled).lower())
|
||||
|
||||
|
||||
def get_fundamentals_parity_report() -> dict | None:
|
||||
"""Return the latest compact A5 summary, if the job has run."""
|
||||
from app.config import settings
|
||||
from app.services.fundamentals_parity_service import load_latest
|
||||
|
||||
report = load_latest(settings.fundamentals_parity_report_dir)
|
||||
if report is not None:
|
||||
report.pop("rows", None) # full per-ticker data is download-only
|
||||
return report
|
||||
|
||||
|
||||
def get_fundamentals_parity_csv() -> tuple[str, str] | None:
|
||||
"""Return the latest A5 CSV filename and content for authenticated download."""
|
||||
from app.config import settings
|
||||
from app.services.fundamentals_parity_service import load_latest_csv
|
||||
|
||||
return load_latest_csv(settings.fundamentals_parity_report_dir)
|
||||
|
||||
|
||||
def get_fundamentals_parity_json() -> tuple[str, str] | None:
|
||||
"""Return the canonical A5 JSON artifact for authenticated download."""
|
||||
from app.config import settings
|
||||
from app.services.fundamentals_parity_service import load_latest_json
|
||||
|
||||
return load_latest_json(settings.fundamentals_parity_report_dir)
|
||||
|
||||
@@ -97,8 +97,8 @@ SIGNAL_BUNDLE_MAX_CHARS = 3900 # Telegram limit is 4096; keep room for HTML par
|
||||
# Hysteresis (a deadband around each divider) stops a point sitting on a boundary
|
||||
# from flip-flopping; the cooldown caps how often a genuine change can re-alert.
|
||||
QUAD_TYPE = "regime_quadrant"
|
||||
QUAD_X_DIV = 60.0 # v2 State divider (backend response is authoritative)
|
||||
QUAD_Y_DIV = 60.0 # v2 Warning divider
|
||||
QUAD_X_DIV = 50.0 # v3 State divider (backend response is authoritative)
|
||||
QUAD_Y_DIV = 40.0 # v3 Warning divider; the axes have different ranges
|
||||
QUAD_MARGIN = 5.0 # half-width of the hysteresis deadband around each divider
|
||||
QUAD_COOLDOWN_DAYS = 3 # min days between quadrant-change alerts
|
||||
QUAD_LABELS = {
|
||||
|
||||
@@ -1320,6 +1320,7 @@ def _replay_candidates_for_period(
|
||||
cadence: str = DEFAULT_BACKTEST_CADENCE,
|
||||
include_short_candidates: bool = False,
|
||||
include_universe_rank_observations: bool = False,
|
||||
outcome_horizon_sessions: int = HORIZON,
|
||||
) -> list[dict]:
|
||||
"""Slim picklable replay used by local event studies.
|
||||
|
||||
@@ -1343,10 +1344,13 @@ def _replay_candidates_for_period(
|
||||
)
|
||||
]
|
||||
cadence = validate_backtest_cadence(cadence)
|
||||
replay_horizon = int(outcome_horizon_sessions)
|
||||
if replay_horizon < 0:
|
||||
raise ValueError('outcome_horizon_sessions must be non-negative')
|
||||
candidates: list[dict] = []
|
||||
for i in range(
|
||||
MIN_LOOKBACK - 1,
|
||||
len(bars) - HORIZON,
|
||||
len(bars) - replay_horizon,
|
||||
backtest_step_sessions(cadence),
|
||||
):
|
||||
if bars[i].date < start_date:
|
||||
@@ -1942,6 +1946,7 @@ def _make_gate_reset_reentry_fn(
|
||||
cadence: str,
|
||||
qualified_fn: Callable[[dict], bool] | None = None,
|
||||
ranking_key: str = PRODUCTION_PERCENTILE_KEY,
|
||||
evaluation_horizon_sessions: int = HORIZON,
|
||||
) -> Callable[[str, int, dict, Any], dict | None]:
|
||||
"""Build the production post-stop gate-reset callback.
|
||||
|
||||
@@ -1959,11 +1964,18 @@ def _make_gate_reset_reentry_fn(
|
||||
|
||||
evaluation_ords: dict[str, set[int]] = {}
|
||||
step_sessions = backtest_step_sessions(cadence)
|
||||
evaluation_horizon = int(evaluation_horizon_sessions)
|
||||
if evaluation_horizon < 0:
|
||||
raise ValueError('evaluation_horizon_sessions must be non-negative')
|
||||
for symbol, columns in prices.items():
|
||||
ordinals = columns[0]
|
||||
evaluation_ords[symbol] = {
|
||||
int(ordinals[index])
|
||||
for index in range(MIN_LOOKBACK - 1, len(ordinals) - HORIZON, step_sessions)
|
||||
for index in range(
|
||||
MIN_LOOKBACK - 1,
|
||||
len(ordinals) - evaluation_horizon,
|
||||
step_sessions,
|
||||
)
|
||||
}
|
||||
|
||||
qualified_by_symbol_date: dict[tuple[str, int], dict] = {}
|
||||
@@ -2010,7 +2022,7 @@ def _simulate_portfolio(
|
||||
*,
|
||||
qualified_fn: Callable[[dict], bool] | None = None,
|
||||
ranking_key: str = PRODUCTION_PERCENTILE_KEY,
|
||||
max_positions: int = SIM_MAX_POSITIONS,
|
||||
max_positions: int | None = SIM_MAX_POSITIONS,
|
||||
risk_per_trade: float = SIM_RISK_PER_TRADE,
|
||||
atr_trail_multiplier: float = ATR_TRAIL_MULTIPLIER,
|
||||
cost_per_side: float = COST_PER_SIDE,
|
||||
@@ -2034,6 +2046,12 @@ def _simulate_portfolio(
|
||||
corr_lookback: int = 120,
|
||||
corr_action: str = "skip",
|
||||
corr_min_overlap: int = 60,
|
||||
min_initial_risk_fraction: float | None = None,
|
||||
weekly_top_n_rebalance: bool = False,
|
||||
daily_rank_map: dict[tuple[str, str], dict[str, float | None]] | None = None,
|
||||
measurement_start_date: date | None = None,
|
||||
hard_end_date: date | None = None,
|
||||
include_capacity_diagnostics: bool = False,
|
||||
) -> dict | None:
|
||||
"""Replay the qualified setups as ONE capital-constrained book and report
|
||||
portfolio economics from the daily equity curve (return, CAGR, drawdown,
|
||||
@@ -2083,6 +2101,20 @@ def _simulate_portfolio(
|
||||
raise ValueError("corr_action must be 'skip' or 'half_size'")
|
||||
if vol_target is not None and vol_target <= 0:
|
||||
raise ValueError("vol_target must be positive when set")
|
||||
if max_positions is not None and int(max_positions) <= 0:
|
||||
raise ValueError("max_positions must be positive or None")
|
||||
if min_initial_risk_fraction is not None and not (
|
||||
0.0 < float(min_initial_risk_fraction) < 1.0
|
||||
):
|
||||
raise ValueError("min_initial_risk_fraction must be between 0 and 1")
|
||||
if weekly_top_n_rebalance and (
|
||||
max_positions is None or daily_rank_map is None
|
||||
):
|
||||
raise ValueError(
|
||||
"weekly_top_n_rebalance requires max_positions and daily_rank_map"
|
||||
)
|
||||
if weekly_top_n_rebalance and fill_mode != FILL_MODE_CLOSE:
|
||||
raise ValueError("weekly_top_n_rebalance requires fill_mode=close")
|
||||
clamp_lo, clamp_hi = float(vol_clamp[0]), float(vol_clamp[1])
|
||||
if clamp_lo <= 0 or clamp_hi < clamp_lo:
|
||||
raise ValueError("vol_clamp must satisfy 0 < lo <= hi")
|
||||
@@ -2094,8 +2126,26 @@ def _simulate_portfolio(
|
||||
|
||||
entries_by_ord: dict[int, list[dict]] = defaultdict(list)
|
||||
start_ord = start_date.toordinal() if start_date is not None else None
|
||||
measurement_start_ord = (
|
||||
measurement_start_date.toordinal()
|
||||
if measurement_start_date is not None
|
||||
else start_ord
|
||||
)
|
||||
hard_end_ord = hard_end_date.toordinal() if hard_end_date is not None else None
|
||||
# Explicit simulator/holdout end dates are exclusive split boundaries.
|
||||
end_ord = end_date.toordinal() if end_date is not None else None
|
||||
if (
|
||||
start_ord is not None
|
||||
and measurement_start_ord is not None
|
||||
and measurement_start_ord < start_ord
|
||||
):
|
||||
raise ValueError("measurement_start_date cannot precede start_date")
|
||||
if (
|
||||
hard_end_ord is not None
|
||||
and measurement_start_ord is not None
|
||||
and hard_end_ord <= measurement_start_ord
|
||||
):
|
||||
raise ValueError("hard_end_date must follow measurement_start_date")
|
||||
for c in candidates:
|
||||
if not qualified_fn(c) or c.get("direction") != "long":
|
||||
continue
|
||||
@@ -2104,6 +2154,8 @@ def _simulate_portfolio(
|
||||
continue
|
||||
if end_ord is not None and entry_ord >= end_ord:
|
||||
continue # holdout/validation: entries strictly before the split
|
||||
if hard_end_ord is not None and entry_ord >= hard_end_ord:
|
||||
continue
|
||||
if not c.get("entry") or not c.get("stop"):
|
||||
continue
|
||||
entries_by_ord[entry_ord].append(c)
|
||||
@@ -2116,7 +2168,12 @@ def _simulate_portfolio(
|
||||
}
|
||||
|
||||
first_ord = start_ord if start_ord is not None else min(entries_by_ord)
|
||||
calendar = sorted({o for cols in prices.values() for o in cols[0] if o >= first_ord})
|
||||
full_calendar = sorted({o for cols in prices.values() for o in cols[0]})
|
||||
calendar = [
|
||||
o
|
||||
for o in full_calendar
|
||||
if o >= first_ord and (hard_end_ord is None or o < hard_end_ord)
|
||||
]
|
||||
if not calendar:
|
||||
return None
|
||||
|
||||
@@ -2124,20 +2181,39 @@ def _simulate_portfolio(
|
||||
# fill lag). Prevents trailing flat-cash after the last resolvable entry —
|
||||
# the clear-air train-window bug — for train, validation, and full-period
|
||||
# books alike (including max-hold sweeps out to 90 days).
|
||||
last_signal_ord = max(entries_by_ord)
|
||||
resolve_pad = hold_days + (1 if fill_mode in DELAYED_FILL_MODES else 0)
|
||||
cut = bisect.bisect_left(calendar, last_signal_ord) + resolve_pad + 1
|
||||
calendar = calendar[:cut]
|
||||
if hard_end_ord is None:
|
||||
last_signal_ord = max(entries_by_ord)
|
||||
resolve_pad = hold_days + (1 if fill_mode in DELAYED_FILL_MODES else 0)
|
||||
cut = bisect.bisect_left(calendar, last_signal_ord) + resolve_pad + 1
|
||||
calendar = calendar[:cut]
|
||||
if not calendar:
|
||||
return None
|
||||
|
||||
weekly_rebalance_ords: set[int] = set()
|
||||
for index, session_ord in enumerate(full_calendar):
|
||||
session_date = date.fromordinal(session_ord)
|
||||
iso = session_date.isocalendar()
|
||||
if index + 1 < len(full_calendar):
|
||||
next_iso = date.fromordinal(full_calendar[index + 1]).isocalendar()
|
||||
if (iso.year, iso.week) != (next_iso.year, next_iso.week):
|
||||
weekly_rebalance_ords.add(session_ord)
|
||||
elif session_date.weekday() == 4:
|
||||
weekly_rebalance_ords.add(session_ord)
|
||||
|
||||
cash = SIM_STARTING_CAPITAL
|
||||
positions: dict[str, dict] = {}
|
||||
curve: list[tuple[int, float]] = []
|
||||
trades: list[dict] = []
|
||||
skipped_full = 0
|
||||
measurement_skipped_full = 0
|
||||
skipped_cooldown = 0
|
||||
skipped_corr = 0
|
||||
skipped_min_initial_risk = 0
|
||||
measurement_skipped_min_initial_risk = 0
|
||||
opened_positions = 0
|
||||
measurement_opened_positions = 0
|
||||
weekly_rank_rejected_entries = 0
|
||||
measurement_weekly_rank_rejected_entries = 0
|
||||
skipped_missing_fill = 0
|
||||
skipped_gap_cap = 0
|
||||
cooldown_until_index: dict[str, int] = {}
|
||||
@@ -2152,6 +2228,12 @@ def _simulate_portfolio(
|
||||
vol_scalars: list[float] = []
|
||||
overnight_slippage_pct: list[float] = []
|
||||
pending_delayed: list[dict] = []
|
||||
measurement_start_equity: float | None = None
|
||||
measurement_start_position_count: int | None = None
|
||||
capacity_samples: list[dict[str, float | int]] = []
|
||||
weekly_rebalance_events: list[dict] = []
|
||||
rebalance_exit_index: dict[str, tuple[int, int]] = {}
|
||||
rebalance_reentry_events: list[dict] = []
|
||||
|
||||
def _bar(sym: str, o: int):
|
||||
idx = index_of.get(sym, {}).get(o)
|
||||
@@ -2221,6 +2303,13 @@ def _simulate_portfolio(
|
||||
cost = proceeds * cost_rate
|
||||
cash += proceeds - cost
|
||||
risk = pos["entry"] - pos["initial_stop"]
|
||||
initial_risk_dollars = pos["shares"] * risk
|
||||
net_pnl = (
|
||||
proceeds
|
||||
- pos["shares"] * pos["entry"]
|
||||
- cost
|
||||
- pos["entry_cost"]
|
||||
)
|
||||
trades.append({
|
||||
"symbol": sym,
|
||||
"entry_ord": pos["entry_ord"],
|
||||
@@ -2229,8 +2318,13 @@ def _simulate_portfolio(
|
||||
"initial_stop": pos["initial_stop"],
|
||||
"active_stop": pos["stop"],
|
||||
"fill": fill,
|
||||
"pnl": proceeds - pos["shares"] * pos["entry"] - cost - pos["entry_cost"],
|
||||
"shares": pos["shares"],
|
||||
"initial_risk_dollars": initial_risk_dollars,
|
||||
"pnl": net_pnl,
|
||||
"r": (fill - pos["entry"]) / risk if risk > 0 else 0.0,
|
||||
"net_r": net_pnl / initial_risk_dollars
|
||||
if initial_risk_dollars > 0
|
||||
else 0.0,
|
||||
"hold": pos["bars_held"],
|
||||
"reason": reason,
|
||||
"stop_refreshes": pos["stop_refreshes"],
|
||||
@@ -2245,6 +2339,13 @@ def _simulate_portfolio(
|
||||
|
||||
cooldown_sessions = max(0, int(reentry_cooldown_sessions))
|
||||
for calendar_index, o in enumerate(calendar):
|
||||
in_measurement = (
|
||||
measurement_start_ord is None or o >= measurement_start_ord
|
||||
)
|
||||
if in_measurement and measurement_start_equity is None:
|
||||
measurement_start_equity = _marked_equity()
|
||||
measurement_start_position_count = len(positions)
|
||||
|
||||
# 1) exits on today's bars (stop intraday, target intraday, time at close)
|
||||
for sym in list(positions):
|
||||
pos = positions[sym]
|
||||
@@ -2358,6 +2459,82 @@ def _simulate_portfolio(
|
||||
reverse=True,
|
||||
)
|
||||
|
||||
weekly_selected_entries: list[dict] | None = None
|
||||
if weekly_top_n_rebalance and o in weekly_rebalance_ords:
|
||||
assert max_positions is not None
|
||||
assert daily_rank_map is not None
|
||||
asof = date.fromordinal(o).isoformat()
|
||||
protected: set[str] = set()
|
||||
ranked_pool: list[tuple[float, int, str, dict | None]] = []
|
||||
for sym in positions:
|
||||
rank_row = daily_rank_map.get((sym, asof))
|
||||
current_rank = (
|
||||
rank_row.get("strategy_rank") if rank_row is not None else None
|
||||
)
|
||||
if current_rank is None or _bar(sym, o) is None:
|
||||
protected.add(sym)
|
||||
continue
|
||||
ranked_pool.append((float(current_rank), 0, sym, None))
|
||||
|
||||
entrants_by_symbol: dict[str, dict] = {}
|
||||
for candidate in signal_todays:
|
||||
sym = str(candidate["symbol"])
|
||||
if sym in positions or sym in entrants_by_symbol:
|
||||
continue
|
||||
entrants_by_symbol[sym] = candidate
|
||||
eligible_entrants = 0
|
||||
for sym, candidate in entrants_by_symbol.items():
|
||||
rank_row = daily_rank_map.get((sym, asof))
|
||||
current_rank = (
|
||||
rank_row.get("strategy_rank") if rank_row is not None else None
|
||||
)
|
||||
if current_rank is None:
|
||||
continue
|
||||
eligible_entrants += 1
|
||||
ranked_pool.append((float(current_rank), 1, sym, candidate))
|
||||
|
||||
available_slots = max(0, int(max_positions) - len(protected))
|
||||
ranked_pool.sort(key=lambda row: (-row[0], row[1], row[2]))
|
||||
selected = ranked_pool[:available_slots]
|
||||
selected_holding_symbols = {
|
||||
sym for _rank, kind, sym, _candidate in selected if kind == 0
|
||||
}
|
||||
weekly_selected_entries = [
|
||||
candidate
|
||||
for _rank, kind, _sym, candidate in selected
|
||||
if kind == 1 and candidate is not None
|
||||
]
|
||||
selected_entrant_symbols = {
|
||||
str(candidate["symbol"]) for candidate in weekly_selected_entries
|
||||
}
|
||||
rejected_now = max(0, eligible_entrants - len(selected_entrant_symbols))
|
||||
weekly_rank_rejected_entries += rejected_now
|
||||
if in_measurement:
|
||||
measurement_weekly_rank_rejected_entries += rejected_now
|
||||
|
||||
exited_symbols: list[str] = []
|
||||
for sym in list(positions):
|
||||
if sym in protected or sym in selected_holding_symbols:
|
||||
continue
|
||||
bar = _bar(sym, o)
|
||||
if bar is None:
|
||||
continue
|
||||
_close_trade(sym, float(bar.close), "weekly_rebalance")
|
||||
rebalance_exit_index[sym] = (calendar_index, o)
|
||||
exited_symbols.append(sym)
|
||||
|
||||
weekly_rebalance_events.append({
|
||||
"ord": o,
|
||||
"fresh_entrant_pool": len(entrants_by_symbol),
|
||||
"rank_eligible_entrant_pool": eligible_entrants,
|
||||
"selected_entrants": len(selected_entrant_symbols),
|
||||
"replacements": len(exited_symbols),
|
||||
"exited_symbols": sorted(exited_symbols),
|
||||
"selected_entrant_symbols": sorted(selected_entrant_symbols),
|
||||
"measurement": in_measurement,
|
||||
})
|
||||
equity = _marked_equity()
|
||||
|
||||
if fill_mode in DELAYED_FILL_MODES:
|
||||
fill_candidates = sorted(
|
||||
pending_delayed,
|
||||
@@ -2366,7 +2543,11 @@ def _simulate_portfolio(
|
||||
)
|
||||
pending_delayed = []
|
||||
else:
|
||||
fill_candidates = signal_todays
|
||||
fill_candidates = (
|
||||
weekly_selected_entries
|
||||
if weekly_selected_entries is not None
|
||||
else signal_todays
|
||||
)
|
||||
|
||||
def _corr_scale_for(sym: str, asof_idx: int) -> float | None:
|
||||
"""1.0 ok, 0.5 half-size, None = skip. Missing history → uncorrelated."""
|
||||
@@ -2411,15 +2592,21 @@ def _simulate_portfolio(
|
||||
corr_scale: float,
|
||||
fill_bar: Any | None,
|
||||
) -> None:
|
||||
nonlocal cash, equity, skipped_full, skipped_cooldown, post_stop_events
|
||||
nonlocal cash, equity, skipped_full, measurement_skipped_full
|
||||
nonlocal skipped_cooldown, post_stop_events
|
||||
nonlocal skipped_min_initial_risk
|
||||
nonlocal measurement_skipped_min_initial_risk
|
||||
nonlocal opened_positions, measurement_opened_positions
|
||||
sym = c["symbol"]
|
||||
if sym in positions:
|
||||
return
|
||||
if calendar_index < cooldown_until_index.get(sym, -1):
|
||||
skipped_cooldown += 1
|
||||
return
|
||||
if len(positions) >= max_positions:
|
||||
if max_positions is not None and len(positions) >= max_positions:
|
||||
skipped_full += 1
|
||||
if in_measurement:
|
||||
measurement_skipped_full += 1
|
||||
return
|
||||
risk_ps = entry - stop
|
||||
if risk_ps <= 0 or entry <= 0:
|
||||
@@ -2436,6 +2623,16 @@ def _simulate_portfolio(
|
||||
(equity * SIM_NOTIONAL_CAP) / entry,
|
||||
max(cash, 0.0) / (entry * (1.0 + cost_rate)),
|
||||
)
|
||||
initial_risk_dollars = shares * risk_ps
|
||||
if (
|
||||
min_initial_risk_fraction is not None
|
||||
and initial_risk_dollars
|
||||
< equity * float(min_initial_risk_fraction)
|
||||
):
|
||||
skipped_min_initial_risk += 1
|
||||
if in_measurement:
|
||||
measurement_skipped_min_initial_risk += 1
|
||||
return
|
||||
if shares * entry < 1.0:
|
||||
return
|
||||
entry_cost = shares * entry * cost_rate
|
||||
@@ -2475,6 +2672,21 @@ def _simulate_portfolio(
|
||||
"vol_scalar": scalar,
|
||||
"corr_scale": corr_scale,
|
||||
}
|
||||
opened_positions += 1
|
||||
if in_measurement:
|
||||
measurement_opened_positions += 1
|
||||
prior_rebalance_exit = rebalance_exit_index.pop(sym, None)
|
||||
if prior_rebalance_exit is not None:
|
||||
prior_exit_index, prior_exit_ord = prior_rebalance_exit
|
||||
rebalance_reentry_events.append({
|
||||
"symbol": sym,
|
||||
"exit_ord": prior_exit_ord,
|
||||
"exit_calendar_index": prior_exit_index,
|
||||
"reentry_calendar_index": calendar_index,
|
||||
"wait_sessions": calendar_index - prior_exit_index,
|
||||
"reentry_ord": entry_ord,
|
||||
"measurement": in_measurement,
|
||||
})
|
||||
# next_open only: fill is at the open, so the rest of the bar can stop out.
|
||||
# stale_close fills at the close — same-day stop after entry does not apply.
|
||||
# bars_held stays 0 on the fill day (matches close-fill cadence).
|
||||
@@ -2576,7 +2788,25 @@ def _simulate_portfolio(
|
||||
# Queue today's signals for the next session's fill.
|
||||
pending_delayed.extend(signal_todays)
|
||||
|
||||
curve.append((o, _marked_equity()))
|
||||
marked_equity = _marked_equity()
|
||||
if in_measurement and include_capacity_diagnostics:
|
||||
gross_notional = sum(
|
||||
pos["shares"] * pos["last_close"] for pos in positions.values()
|
||||
)
|
||||
capacity_samples.append({
|
||||
"positions": len(positions),
|
||||
"cash_pct": cash / marked_equity * 100.0
|
||||
if marked_equity > 0
|
||||
else 0.0,
|
||||
"gross_exposure_pct": gross_notional / marked_equity * 100.0
|
||||
if marked_equity > 0
|
||||
else 0.0,
|
||||
"at_capacity": int(
|
||||
max_positions is not None
|
||||
and len(positions) >= max_positions
|
||||
),
|
||||
})
|
||||
curve.append((o, marked_equity))
|
||||
|
||||
# Close whatever is still open at its last mark so final equity is realized.
|
||||
for sym in list(positions):
|
||||
@@ -2584,32 +2814,57 @@ def _simulate_portfolio(
|
||||
final_equity = cash
|
||||
curve[-1] = (calendar[-1], final_equity)
|
||||
|
||||
total_return_pct = (final_equity / SIM_STARTING_CAPITAL - 1.0) * 100.0
|
||||
years = (calendar[-1] - calendar[0]) / 365.25
|
||||
metric_start_ord = (
|
||||
measurement_start_ord if measurement_start_ord is not None else calendar[0]
|
||||
)
|
||||
metric_curve = [(day_ord, eq) for day_ord, eq in curve if day_ord >= metric_start_ord]
|
||||
if not metric_curve:
|
||||
return None
|
||||
metric_base_equity = (
|
||||
measurement_start_equity
|
||||
if measurement_start_date is not None and measurement_start_equity is not None
|
||||
else SIM_STARTING_CAPITAL
|
||||
)
|
||||
total_return_pct = (final_equity / metric_base_equity - 1.0) * 100.0
|
||||
years = (calendar[-1] - metric_start_ord) / 365.25
|
||||
cagr_pct = (
|
||||
((final_equity / SIM_STARTING_CAPITAL) ** (1.0 / years) - 1.0) * 100.0
|
||||
((final_equity / metric_base_equity) ** (1.0 / years) - 1.0) * 100.0
|
||||
if years > 0.25 and final_equity > 0
|
||||
else None
|
||||
)
|
||||
|
||||
peak = float("-inf")
|
||||
max_dd = 0.0
|
||||
for _, eq in curve:
|
||||
drawdown_equities = (
|
||||
[metric_base_equity, *(eq for _, eq in metric_curve)]
|
||||
if measurement_start_date is not None
|
||||
else [eq for _, eq in metric_curve]
|
||||
)
|
||||
for eq in drawdown_equities:
|
||||
peak = max(peak, eq)
|
||||
if peak > 0:
|
||||
max_dd = max(max_dd, (peak - eq) / peak)
|
||||
|
||||
rets = [b / a - 1.0 for (_, a), (_, b) in zip(curve, curve[1:]) if a > 0]
|
||||
return_equities = (
|
||||
[metric_base_equity, *(eq for _, eq in metric_curve)]
|
||||
if measurement_start_date is not None
|
||||
else [eq for _, eq in metric_curve]
|
||||
)
|
||||
rets = [
|
||||
b / a - 1.0
|
||||
for a, b in zip(return_equities, return_equities[1:])
|
||||
if a > 0
|
||||
]
|
||||
diag = sharpe_diagnostics(rets)
|
||||
sharpe = diag["sharpe"]
|
||||
|
||||
# Per-calendar-year returns off the equity curve — shows whether every year
|
||||
# contributed or one exceptional stretch carried the result.
|
||||
yearly: list[dict] = []
|
||||
year_start_eq = curve[0][1]
|
||||
cur_year = date.fromordinal(curve[0][0]).year
|
||||
last_eq = curve[0][1]
|
||||
for o, eq in curve:
|
||||
year_start_eq = metric_base_equity
|
||||
cur_year = date.fromordinal(metric_start_ord).year
|
||||
last_eq = metric_base_equity
|
||||
for o, eq in metric_curve:
|
||||
y = date.fromordinal(o).year
|
||||
if y != cur_year:
|
||||
yearly.append({
|
||||
@@ -2628,24 +2883,29 @@ def _simulate_portfolio(
|
||||
),
|
||||
})
|
||||
|
||||
pnls = [t["pnl"] for t in trades]
|
||||
metric_trades = [
|
||||
trade for trade in trades if trade["entry_ord"] >= metric_start_ord
|
||||
]
|
||||
pnls = [t["pnl"] for t in metric_trades]
|
||||
wins = sum(1 for p in pnls if p > 0)
|
||||
reason_counts = {
|
||||
reason: sum(1 for t in trades if t["reason"] == reason)
|
||||
for reason in sorted({t["reason"] for t in trades})
|
||||
reason: sum(1 for t in metric_trades if t["reason"] == reason)
|
||||
for reason in sorted({t["reason"] for t in metric_trades})
|
||||
}
|
||||
spy_pct = None
|
||||
if spy_closes:
|
||||
from app.services.benchmark_service import benchmark_return_pct
|
||||
|
||||
spy_pct = benchmark_return_pct(
|
||||
spy_closes, date.fromordinal(calendar[0]), date.fromordinal(calendar[-1])
|
||||
spy_closes,
|
||||
date.fromordinal(metric_start_ord),
|
||||
date.fromordinal(calendar[-1]),
|
||||
)
|
||||
|
||||
curve_payload: list[dict] | None = None
|
||||
benchmark_payload: list[dict] | None = None
|
||||
if include_curve:
|
||||
curve_base = curve[0][1] if curve else SIM_STARTING_CAPITAL
|
||||
curve_base = metric_base_equity
|
||||
curve_payload = [
|
||||
{
|
||||
"date": date.fromordinal(o).isoformat(),
|
||||
@@ -2654,12 +2914,12 @@ def _simulate_portfolio(
|
||||
if curve_base > 0
|
||||
else None,
|
||||
}
|
||||
for o, eq in curve
|
||||
for o, eq in metric_curve
|
||||
]
|
||||
if spy_closes:
|
||||
benchmark_payload = []
|
||||
base_spy = None
|
||||
for o, _ in curve:
|
||||
for o, _ in metric_curve:
|
||||
d = date.fromordinal(o)
|
||||
close = spy_closes.get(d)
|
||||
if close is None or close <= 0:
|
||||
@@ -2678,6 +2938,8 @@ def _simulate_portfolio(
|
||||
calmar = float(cagr_pct) / max_dd_pct
|
||||
result = {
|
||||
"starting_capital": SIM_STARTING_CAPITAL,
|
||||
"measurement_start_equity": round(metric_base_equity, 2),
|
||||
"measurement_start_positions": measurement_start_position_count or 0,
|
||||
"cost_per_side_pct": round(cost_rate * 100.0, 3),
|
||||
"fill_mode": fill_mode,
|
||||
"final_equity": round(final_equity, 2),
|
||||
@@ -2691,23 +2953,161 @@ def _simulate_portfolio(
|
||||
"n_returns": diag["n_returns"],
|
||||
"return_skew": diag["return_skew"],
|
||||
"return_kurtosis": diag["return_kurtosis"],
|
||||
"trades": len(trades),
|
||||
"win_rate": round(wins / len(trades) * 100.0, 1) if trades else None,
|
||||
"trades": len(metric_trades),
|
||||
"win_rate": (
|
||||
round(wins / len(metric_trades) * 100.0, 1)
|
||||
if metric_trades
|
||||
else None
|
||||
),
|
||||
"avg_trade_pnl": round(sum(pnls) / len(pnls), 2) if pnls else None,
|
||||
"best_trade_r": round(max(t["r"] for t in trades), 2) if trades else None,
|
||||
"worst_trade_r": round(min(t["r"] for t in trades), 2) if trades else None,
|
||||
"best_trade_r": (
|
||||
round(max(t["r"] for t in metric_trades), 2)
|
||||
if metric_trades
|
||||
else None
|
||||
),
|
||||
"worst_trade_r": (
|
||||
round(min(t["r"] for t in metric_trades), 2)
|
||||
if metric_trades
|
||||
else None
|
||||
),
|
||||
"best_trade_pnl": round(max(pnls), 2) if pnls else None,
|
||||
"worst_trade_pnl": round(min(pnls), 2) if pnls else None,
|
||||
"avg_hold_days": (
|
||||
round(sum(t["hold"] for t in trades) / len(trades), 1) if trades else None
|
||||
round(
|
||||
sum(t["hold"] for t in metric_trades) / len(metric_trades),
|
||||
1,
|
||||
)
|
||||
if metric_trades
|
||||
else None
|
||||
),
|
||||
"exit_reasons": reason_counts,
|
||||
"skipped_book_full": skipped_full,
|
||||
"spy_return_pct": round(spy_pct, 1) if spy_pct is not None else None,
|
||||
"yearly_returns": yearly,
|
||||
"start_date": date.fromordinal(calendar[0]).isoformat(),
|
||||
"start_date": date.fromordinal(metric_start_ord).isoformat(),
|
||||
"end_date": date.fromordinal(calendar[-1]).isoformat(),
|
||||
}
|
||||
if measurement_start_date is not None:
|
||||
result["simulation_start_date"] = date.fromordinal(calendar[0]).isoformat()
|
||||
if hard_end_date is not None:
|
||||
result["hard_end_date_exclusive"] = hard_end_date.isoformat()
|
||||
if measurement_start_date is not None:
|
||||
result["measurement_skipped_book_full"] = measurement_skipped_full
|
||||
result["measurement_opened_positions"] = measurement_opened_positions
|
||||
if min_initial_risk_fraction is not None:
|
||||
result["min_initial_risk_fraction"] = float(min_initial_risk_fraction)
|
||||
result["skipped_min_initial_risk"] = skipped_min_initial_risk
|
||||
result["measurement_skipped_min_initial_risk"] = (
|
||||
measurement_skipped_min_initial_risk
|
||||
)
|
||||
if include_capacity_diagnostics:
|
||||
measured_opened = (
|
||||
measurement_opened_positions
|
||||
if measurement_start_date is not None
|
||||
else opened_positions
|
||||
)
|
||||
measured_full = (
|
||||
measurement_skipped_full
|
||||
if measurement_start_date is not None
|
||||
else skipped_full
|
||||
)
|
||||
capacity_opportunities = measured_opened + measured_full
|
||||
result["opened_positions"] = measured_opened
|
||||
result["capacity_opportunities"] = capacity_opportunities
|
||||
result["blocked_fraction"] = (
|
||||
round(measured_full / capacity_opportunities, 6)
|
||||
if capacity_opportunities
|
||||
else 0.0
|
||||
)
|
||||
result["avg_positions"] = (
|
||||
round(
|
||||
sum(float(sample["positions"]) for sample in capacity_samples)
|
||||
/ len(capacity_samples),
|
||||
4,
|
||||
)
|
||||
if capacity_samples
|
||||
else 0.0
|
||||
)
|
||||
result["peak_positions"] = (
|
||||
max(int(sample["positions"]) for sample in capacity_samples)
|
||||
if capacity_samples
|
||||
else 0
|
||||
)
|
||||
result["sessions_at_capacity"] = sum(
|
||||
int(sample["at_capacity"]) for sample in capacity_samples
|
||||
)
|
||||
result["sessions_measured"] = len(capacity_samples)
|
||||
result["avg_cash_pct"] = (
|
||||
round(
|
||||
sum(float(sample["cash_pct"]) for sample in capacity_samples)
|
||||
/ len(capacity_samples),
|
||||
4,
|
||||
)
|
||||
if capacity_samples
|
||||
else None
|
||||
)
|
||||
result["avg_gross_exposure_pct"] = (
|
||||
round(
|
||||
sum(
|
||||
float(sample["gross_exposure_pct"])
|
||||
for sample in capacity_samples
|
||||
)
|
||||
/ len(capacity_samples),
|
||||
4,
|
||||
)
|
||||
if capacity_samples
|
||||
else None
|
||||
)
|
||||
if weekly_top_n_rebalance:
|
||||
measured_events = [
|
||||
event for event in weekly_rebalance_events if event["measurement"]
|
||||
]
|
||||
measured_reentries = [
|
||||
event for event in rebalance_reentry_events if event["measurement"]
|
||||
]
|
||||
result["weekly_rank_rejected_entries"] = (
|
||||
measurement_weekly_rank_rejected_entries
|
||||
if measurement_start_date is not None
|
||||
else weekly_rank_rejected_entries
|
||||
)
|
||||
result["weekly_rebalance_events"] = [
|
||||
{
|
||||
**{
|
||||
key: value
|
||||
for key, value in event.items()
|
||||
if key not in {"ord", "measurement"}
|
||||
},
|
||||
"date": date.fromordinal(event["ord"]).isoformat(),
|
||||
}
|
||||
for event in measured_events
|
||||
]
|
||||
result["rebalance_reentry_events"] = [
|
||||
{
|
||||
**{
|
||||
key: value
|
||||
for key, value in event.items()
|
||||
if key
|
||||
not in {
|
||||
"exit_ord",
|
||||
"reentry_ord",
|
||||
"measurement",
|
||||
"exit_calendar_index",
|
||||
"reentry_calendar_index",
|
||||
}
|
||||
},
|
||||
"exit_date": date.fromordinal(event["exit_ord"]).isoformat(),
|
||||
"reentry_date": date.fromordinal(
|
||||
event["reentry_ord"]
|
||||
).isoformat(),
|
||||
}
|
||||
for event in measured_reentries
|
||||
]
|
||||
for session_limit in (5, 10, 20):
|
||||
result[f"rebalance_reentries_within_{session_limit}_sessions"] = sum(
|
||||
1
|
||||
for event in measured_reentries
|
||||
if int(event["wait_sessions"]) <= session_limit
|
||||
)
|
||||
if vol_target is not None:
|
||||
result["vol_target"] = vol_target
|
||||
result["vol_lookback"] = int(vol_lookback)
|
||||
@@ -2782,7 +3182,7 @@ def _simulate_portfolio(
|
||||
"entry_date": date.fromordinal(trade["entry_ord"]).isoformat(),
|
||||
"exit_date": date.fromordinal(trade["exit_ord"]).isoformat(),
|
||||
}
|
||||
for trade in trades
|
||||
for trade in metric_trades
|
||||
]
|
||||
return result
|
||||
|
||||
|
||||
@@ -72,15 +72,25 @@ def _breadth_from_closes(
|
||||
return _breadth_with_counts(closes_by_symbol, window, min_tickers)[0]
|
||||
|
||||
|
||||
# Breadth deterioration counts fully when price masks it (true divergence, the
|
||||
# dangerous pre-top case) and at CONFIRMED_FLOOR when price falls with it.
|
||||
# v2 used a hard ``price_ret >= 0`` cliff, which zeroed the sensor during every
|
||||
# decline -- so on 2026-07-24, with the basket shedding 10 percentage points
|
||||
# above their 200-DMA in 20 sessions, Warning read exactly 0. Breadth *level*
|
||||
# lives in State but breadth *velocity* appears nowhere else, so partial credit
|
||||
# here is not double counting.
|
||||
DIVERGENCE_CONFIRMED_FLOOR = 0.35
|
||||
DIVERGENCE_TAPER_PCT = 3.0
|
||||
|
||||
|
||||
def compute_divergence_series(
|
||||
breadth: dict[date, float], benchmark_closes: Series, lookback: int = 20
|
||||
) -> dict[date, float]:
|
||||
"""Early-warning score (0-100, high = fragile) per date.
|
||||
|
||||
This is deliberately a pure divergence: it is positive only when benchmark
|
||||
price holds/rises while breadth falls. Absolute low breadth belongs in the
|
||||
State score, so it is not counted again here. A 20 percentage-point breadth
|
||||
deterioration maps to 100.
|
||||
A 20 percentage-point breadth deterioration maps to 100 when the benchmark
|
||||
is flat or rising, tapering to ``DIVERGENCE_CONFIRMED_FLOOR`` of that once
|
||||
the benchmark is down ``DIVERGENCE_TAPER_PCT`` or more over the window.
|
||||
"""
|
||||
bench = {d: c for d, c in benchmark_closes}
|
||||
common = sorted(d for d in bench if d in breadth)
|
||||
@@ -93,8 +103,9 @@ def compute_divergence_series(
|
||||
price_ret = (bench[d] / price_past - 1.0) * 100.0 # %
|
||||
breadth_chg = breadth[d] - breadth[d0] # percentage points
|
||||
deterioration = max(0.0, -breadth_chg)
|
||||
score = deterioration * 5.0 if price_ret >= 0 else 0.0
|
||||
out[d] = max(0.0, min(100.0, round(score, 2)))
|
||||
taper = max(0.0, min(1.0, (price_ret + DIVERGENCE_TAPER_PCT) / DIVERGENCE_TAPER_PCT))
|
||||
gate = DIVERGENCE_CONFIRMED_FLOOR + (1.0 - DIVERGENCE_CONFIRMED_FLOOR) * taper
|
||||
out[d] = max(0.0, min(100.0, round(deterioration * 5.0 * gate, 2)))
|
||||
return out
|
||||
|
||||
|
||||
|
||||
@@ -30,10 +30,10 @@ import hashlib
|
||||
import json
|
||||
import logging
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import date, datetime, timezone
|
||||
from datetime import date, datetime, timedelta, timezone
|
||||
from typing import Any, Protocol, runtime_checkable
|
||||
|
||||
from sqlalchemy import select, text
|
||||
from sqlalchemy import exists, select, text
|
||||
from sqlalchemy.engine import Engine # noqa: F401 (typing only)
|
||||
from sqlalchemy.ext.asyncio import AsyncEngine, AsyncSession
|
||||
|
||||
@@ -48,6 +48,7 @@ STATUS_RUNNING = "running"
|
||||
STATUS_VALIDATED = "validated"
|
||||
STATUS_PROMOTED = "promoted"
|
||||
STATUS_NO_OP = "no_op"
|
||||
STATUS_DEFERRED = "deferred"
|
||||
STATUS_FAILED = "failed"
|
||||
|
||||
_MAX_ERROR_LEN = 4000
|
||||
@@ -68,6 +69,12 @@ class ValidationResult:
|
||||
summary: dict[str, Any] = field(default_factory=dict)
|
||||
source_max_date: date | None = None
|
||||
messages: list[str] = field(default_factory=list)
|
||||
# Expected source-side lag: retry without an immediate error alert. Sources
|
||||
# can bound the quiet period with deferred_alert_after_days. Only meaningful
|
||||
# when ok=False.
|
||||
retryable: bool = False
|
||||
deferred_alert_after_days: int | None = None
|
||||
deferred_alert_messages: list[str] = field(default_factory=list)
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
@@ -123,19 +130,46 @@ async def _last_promoted_revision(db: AsyncSession, source: str) -> str | None:
|
||||
return row.scalar_one_or_none()
|
||||
|
||||
|
||||
async def _promotion_state_since(
|
||||
db: AsyncSession, source: str, cutoff: datetime
|
||||
) -> str:
|
||||
promoted = (
|
||||
DataImportRun.source == source,
|
||||
DataImportRun.status == STATUS_PROMOTED,
|
||||
)
|
||||
ever, recent = (
|
||||
await db.execute(
|
||||
select(
|
||||
exists().where(*promoted),
|
||||
exists().where(*promoted, DataImportRun.started_at >= cutoff),
|
||||
)
|
||||
)
|
||||
).one()
|
||||
return "recent" if recent else "stale" if ever else "never"
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now(timezone.utc)
|
||||
|
||||
|
||||
async def _alert(db: AsyncSession, source: str, code: str, messages: list[str]) -> None:
|
||||
async def _alert(
|
||||
db: AsyncSession,
|
||||
source: str,
|
||||
code: str,
|
||||
messages: list[str],
|
||||
*,
|
||||
severity: str = "error",
|
||||
dedup_hours: int = 24,
|
||||
) -> None:
|
||||
try:
|
||||
await system_event_service.log_event(
|
||||
db,
|
||||
severity="error",
|
||||
severity=severity,
|
||||
source="data_import",
|
||||
code=f"{source}_{code}",
|
||||
message=(("; ".join(messages)) or code)[:_MAX_ERROR_LEN],
|
||||
dedup_key=f"data_import:{source}:{code}",
|
||||
dedup_hours=dedup_hours,
|
||||
)
|
||||
except Exception: # noqa: BLE001 — alerting must never mask the real outcome
|
||||
logger.exception("Failed to emit data_import alert %s/%s", source, code)
|
||||
@@ -145,11 +179,18 @@ async def run_import(
|
||||
importer: SourceImporter,
|
||||
*,
|
||||
engine: AsyncEngine | None = None,
|
||||
force: bool = False,
|
||||
) -> DataImportRun | None:
|
||||
"""Run one import for ``importer``.
|
||||
|
||||
Returns the recorded ``DataImportRun`` (promoted / no_op / failed), or None
|
||||
Returns the recorded ``DataImportRun`` (promoted / no_op / deferred /
|
||||
failed), or None
|
||||
when the per-source advisory lock is already held (another run is active).
|
||||
|
||||
``force`` runs even when the revision is unchanged. The revision tracks the
|
||||
*source*, so a re-import driven by a change on our side — a parser fix that
|
||||
makes stored rows stale — is a no_op under the normal gate. Manually invoked
|
||||
only; scheduled jobs must leave it False so an unchanged source stays a no_op.
|
||||
"""
|
||||
engine = engine or app_engine
|
||||
source = importer.source
|
||||
@@ -189,7 +230,7 @@ async def run_import(
|
||||
revision = await importer.detect_revision(session)
|
||||
run.revision = revision
|
||||
last_rev = await _last_promoted_revision(session, source)
|
||||
if revision is not None and revision == last_rev:
|
||||
if not force and revision is not None and revision == last_rev:
|
||||
run.status = STATUS_NO_OP
|
||||
run.completed_at = _now()
|
||||
await session.commit()
|
||||
@@ -202,9 +243,46 @@ async def run_import(
|
||||
run.validation_json = json.dumps(result.summary, default=str)
|
||||
|
||||
if not result.ok:
|
||||
run.status = STATUS_FAILED
|
||||
run.error_details = ("; ".join(result.messages))[:_MAX_ERROR_LEN]
|
||||
run.completed_at = _now()
|
||||
if result.retryable:
|
||||
run.status = STATUS_DEFERRED
|
||||
await session.commit()
|
||||
alert_days = result.deferred_alert_after_days
|
||||
if alert_days is not None:
|
||||
alert_days = max(1, alert_days)
|
||||
cutoff = run.started_at - timedelta(days=alert_days)
|
||||
promotion_state = await _promotion_state_since(
|
||||
session, source, cutoff
|
||||
)
|
||||
if promotion_state != "recent":
|
||||
history = (
|
||||
f"{source} import has never promoted successfully"
|
||||
if promotion_state == "never"
|
||||
else f"{source} import has not promoted successfully "
|
||||
f"within {alert_days} day(s)"
|
||||
)
|
||||
await _alert(
|
||||
session,
|
||||
source,
|
||||
"deferred_stale",
|
||||
[
|
||||
f"{history}; import remains deferred",
|
||||
*result.deferred_alert_messages,
|
||||
f"Current deferral: "
|
||||
f"{run.error_details or 'validation deferred'}",
|
||||
],
|
||||
severity="warning",
|
||||
dedup_hours=alert_days * 24,
|
||||
)
|
||||
logger.info(
|
||||
"data_import %s: deferred for retry: %s",
|
||||
source,
|
||||
result.messages,
|
||||
)
|
||||
return run
|
||||
|
||||
run.status = STATUS_FAILED
|
||||
await session.commit()
|
||||
await _alert(session, source, "validation_failed", result.messages)
|
||||
logger.warning(
|
||||
|
||||
@@ -28,6 +28,10 @@ DRAWDOWN_LOOKBACK = 252
|
||||
HORIZON_DAYS = 20
|
||||
WARN_PERCENTILE = 80.0
|
||||
TRAIN_FRACTION = 0.70
|
||||
# Below this many holdout corrections, recall is one event away from a very
|
||||
# different headline and should not be read as a property of the score.
|
||||
MIN_EVENTS_FOR_CONFIDENCE = 8
|
||||
SENSOR_MISMATCH_TOLERANCE = 0.10
|
||||
|
||||
|
||||
def _median(values: list[float]) -> float | None:
|
||||
@@ -149,30 +153,72 @@ def _warning_series(
|
||||
breadth_divergence: dict[date, float],
|
||||
dates: list[date],
|
||||
config: dict,
|
||||
) -> dict[date, float]:
|
||||
"""Technical Warning score used historically (fundamentals have no PIT history)."""
|
||||
oas_series: rms.Series | None = None,
|
||||
) -> tuple[dict[date, float], dict[date, int]]:
|
||||
"""Warning score per session plus how many sensors backed it.
|
||||
|
||||
v2 re-derived this by hand from ``WARNING_WEIGHTS`` and so would have kept
|
||||
measuring the old construct after a scoring change. Since v3 dropped
|
||||
fundamentals from the score, this is now exactly the live Warning score
|
||||
rather than a technical-only approximation of it.
|
||||
|
||||
The sensor count matters because the score renormalises over whatever is
|
||||
available: a session backed by two sensors is not drawn from the same
|
||||
distribution as one backed by three, and the frozen threshold assumes it is.
|
||||
"""
|
||||
tickers = config["tickers"]
|
||||
smh_full = prices.get(tickers["leaders"][0], [])
|
||||
spy_full = prices.get(tickers["market"], [])
|
||||
out: dict[date, float] = {}
|
||||
backing: dict[date, int] = {}
|
||||
for session in dates:
|
||||
divergence = breadth_divergence.get(session)
|
||||
relative = rms.p4_relative_strength(
|
||||
sensors = rms.warning_sensor_scores(
|
||||
breadth_divergence.get(session),
|
||||
rms._closes_asof(smh_full, session),
|
||||
rms._closes_asof(spy_full, session),
|
||||
rms._window_asof(oas_series, session, rms.HY_OAS_WINDOW_DAYS),
|
||||
)
|
||||
values: list[tuple[float, float]] = []
|
||||
if divergence is not None:
|
||||
values.append((divergence, rms.WARNING_WEIGHTS["breadth_divergence"]))
|
||||
if relative is not None:
|
||||
values.append((relative, rms.WARNING_WEIGHTS["relative_strength"]))
|
||||
if values:
|
||||
out[session] = round(
|
||||
sum(value * weight for value, weight in values)
|
||||
/ sum(weight for _, weight in values),
|
||||
2,
|
||||
)
|
||||
return out
|
||||
score = rms.score_warning_sensors(sensors)
|
||||
if score is not None:
|
||||
out[session] = round(score, 2)
|
||||
backing[session] = sum(1 for value in sensors.values() if value is not None)
|
||||
return out, backing
|
||||
|
||||
|
||||
def _reliability(
|
||||
dates: list[date],
|
||||
split: int,
|
||||
backing: dict[date, int],
|
||||
events_detected: int,
|
||||
events_in_holdout: int,
|
||||
) -> dict:
|
||||
"""How far the headline metrics can actually be trusted.
|
||||
|
||||
Two things repeatedly invite over-reading this report:
|
||||
|
||||
* The holdout carries only the corrections that fall in the last 30% of the
|
||||
sample. A "2/4" is one event away from "3/4", and in practice the events
|
||||
that flip are decided by where the frozen threshold happens to land rather
|
||||
than by whether the score saw anything.
|
||||
* The score renormalises over available sensors, so a training window that
|
||||
predates a sensor's history freezes a threshold on a different construct
|
||||
than the holdout is measured against.
|
||||
"""
|
||||
expected = len(rms.WARNING_WEIGHTS)
|
||||
train = [backing[d] for d in dates[:split] if d in backing]
|
||||
holdout = [backing[d] for d in dates[split:] if d in backing]
|
||||
train_full = sum(1 for n in train if n == expected) / len(train) if train else 0.0
|
||||
holdout_full = sum(1 for n in holdout if n == expected) / len(holdout) if holdout else 0.0
|
||||
return {
|
||||
"events_detected": events_detected,
|
||||
"events_in_holdout": events_in_holdout,
|
||||
"minimum_events": MIN_EVENTS_FOR_CONFIDENCE,
|
||||
"underpowered": events_in_holdout < MIN_EVENTS_FOR_CONFIDENCE,
|
||||
"sensors_expected": expected,
|
||||
"train_full_sensor_share": round(train_full * 100, 1),
|
||||
"holdout_full_sensor_share": round(holdout_full * 100, 1),
|
||||
"sensor_coverage_mismatch": abs(train_full - holdout_full) > SENSOR_MISMATCH_TOLERANCE,
|
||||
}
|
||||
|
||||
|
||||
async def run_event_study(
|
||||
@@ -195,7 +241,13 @@ async def run_event_study(
|
||||
db, config["breadth_basket"], window=200, min_tickers=20
|
||||
)
|
||||
divergence = breadth_service.compute_divergence_series(breadth, benchmark)
|
||||
warning = _warning_series(prices, divergence, dates, config)
|
||||
oas_series = await rms._fetch_fred_series("BAMLH0A0HYM2", start, end)
|
||||
warning, backing = _warning_series(prices, divergence, dates, config, oas_series)
|
||||
# The credit sensor cannot reach back as far as the price history does (the
|
||||
# upstream series is capped at ~3 years), so the earlier part of the sample
|
||||
# scores on W1+W2 alone via renormalisation. Report where W3 starts rather
|
||||
# than letting the threshold quietly straddle two sensor sets.
|
||||
credit_from = oas_series[0][0].isoformat() if oas_series else None
|
||||
|
||||
split = max(1, min(len(dates) - 1, int(len(dates) * TRAIN_FRACTION)))
|
||||
train_values = [warning[d] for d in dates[:split] if d in warning]
|
||||
@@ -212,6 +264,8 @@ async def run_event_study(
|
||||
metrics["false_alarms"] / (holdout_sessions / 252.0), 2
|
||||
)
|
||||
|
||||
reliability = _reliability(dates, split, backing, len(all_events), len(holdout_events))
|
||||
|
||||
basket_asof = date.fromisoformat(config["basket_asof"])
|
||||
retrospective = dates[split] < basket_asof
|
||||
evaluation = "exploratory" if retrospective else "holdout"
|
||||
@@ -224,7 +278,14 @@ async def run_event_study(
|
||||
f"{evaluation.capitalize()} chronological test: warning episodes preceded "
|
||||
f"{metrics['events_warned']}/{metrics['events']} 10% corrections; "
|
||||
f"{metrics['events_missed']} missed, {metrics['false_alarms_per_year']:.1f} "
|
||||
f"false alarms/year, {lead_text}."
|
||||
f"false alarms/year, {lead_text}. "
|
||||
f"{metrics['events']} of {reliability['events_detected']} detected corrections "
|
||||
f"fall in the test period"
|
||||
+ (
|
||||
"; too few to read recall as a property of the score."
|
||||
if reliability["underpowered"]
|
||||
else "."
|
||||
)
|
||||
)
|
||||
per_event = metrics.pop("per_event")
|
||||
|
||||
@@ -243,6 +304,7 @@ async def run_event_study(
|
||||
"train_fraction": TRAIN_FRACTION,
|
||||
"warn_percentile": WARN_PERCENTILE,
|
||||
"warn_threshold": round(warn_threshold, 1),
|
||||
"credit_sensor_from": credit_from,
|
||||
"basket_hash": rms._basket_hash(config["breadth_basket"]),
|
||||
"basket_asof": config["basket_asof"],
|
||||
},
|
||||
@@ -255,6 +317,7 @@ async def run_event_study(
|
||||
"holdout_sessions": holdout_sessions,
|
||||
},
|
||||
"metrics": metrics,
|
||||
"reliability": reliability,
|
||||
"events": per_event,
|
||||
"recent_breadth": [
|
||||
{"date": d.isoformat(), "breadth": breadth[d], "warning": warning.get(d)}
|
||||
@@ -266,8 +329,11 @@ async def run_event_study(
|
||||
"event": "regime_event_study_complete",
|
||||
"evaluation": evaluation,
|
||||
"events": metrics["events"],
|
||||
"events_detected": reliability["events_detected"],
|
||||
"warned": metrics["events_warned"],
|
||||
"false_alarms_per_year": metrics["false_alarms_per_year"],
|
||||
"underpowered": reliability["underpowered"],
|
||||
"sensor_coverage_mismatch": reliability["sensor_coverage_mismatch"],
|
||||
}))
|
||||
return report
|
||||
|
||||
|
||||
@@ -0,0 +1,179 @@
|
||||
"""A5 activation: refresh the legacy fundamentals cache from local bulk data."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import date, datetime, timezone
|
||||
from typing import Any
|
||||
|
||||
from sqlalchemy import select, update
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.database import insert_for_session
|
||||
from app.models.fundamental import FundamentalData
|
||||
from app.models.score import CompositeScore, DimensionScore
|
||||
from app.services import fundamentals_candidate_service, settings_store
|
||||
|
||||
|
||||
# Absence is deliberately false. Production activation therefore requires one
|
||||
# explicit, durable SystemSetting change after the A5 evidence is approved.
|
||||
ACTIVATION_KEY = "fundamental_data_sec_dolt_cutover_enabled"
|
||||
_SCORE_FIELDS = ("pe_ratio", "revenue_growth", "earnings_surprise")
|
||||
|
||||
|
||||
async def is_enabled(db: AsyncSession) -> bool:
|
||||
raw = await settings_store.get_value(db, ACTIVATION_KEY, "false")
|
||||
return str(raw).strip().lower() == "true"
|
||||
|
||||
|
||||
async def refresh_if_enabled(
|
||||
db: AsyncSession,
|
||||
*,
|
||||
now: datetime | None = None,
|
||||
today: date | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Refresh atomically when activated; otherwise perform no writes."""
|
||||
if not await is_enabled(db):
|
||||
return {
|
||||
"enabled": False,
|
||||
"refreshed": 0,
|
||||
"score_inputs_changed": 0,
|
||||
"dimension_scores_staled": 0,
|
||||
"composite_scores_staled": 0,
|
||||
}
|
||||
return await refresh(db, now=now, today=today)
|
||||
|
||||
|
||||
async def refresh(
|
||||
db: AsyncSession,
|
||||
*,
|
||||
now: datetime | None = None,
|
||||
today: date | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Replace every ticker's compat-cache row in one database transaction.
|
||||
|
||||
Candidate values are assembled before the first write and use only local
|
||||
PostgreSQL tables. A failure rolls the whole refresh back. Only changes to
|
||||
the three scoring inputs invalidate cached scores; market cap and the next
|
||||
earnings date are display-only.
|
||||
"""
|
||||
refreshed_at = now or datetime.now(timezone.utc)
|
||||
candidates = await fundamentals_candidate_service.build_candidates(
|
||||
db, today=today
|
||||
)
|
||||
ticker_ids = [candidate.ticker_id for candidate in candidates]
|
||||
existing = await _existing_by_ticker(db, ticker_ids)
|
||||
changed_ids = {
|
||||
candidate.ticker_id
|
||||
for candidate in candidates
|
||||
if _score_inputs_changed(existing.get(candidate.ticker_id), candidate)
|
||||
}
|
||||
|
||||
for candidate in candidates:
|
||||
unavailable_json = json.dumps(
|
||||
candidate.unavailable_fields, sort_keys=True
|
||||
)
|
||||
stmt = insert_for_session(db, FundamentalData).values(
|
||||
ticker_id=candidate.ticker_id,
|
||||
pe_ratio=candidate.pe_ratio,
|
||||
revenue_growth=candidate.revenue_growth,
|
||||
earnings_surprise=candidate.earnings_surprise,
|
||||
market_cap=candidate.market_cap,
|
||||
next_earnings_date=candidate.next_earnings_date,
|
||||
fetched_at=refreshed_at,
|
||||
unavailable_fields_json=unavailable_json,
|
||||
)
|
||||
await db.execute(
|
||||
stmt.on_conflict_do_update(
|
||||
index_elements=["ticker_id"],
|
||||
set_={
|
||||
"pe_ratio": stmt.excluded.pe_ratio,
|
||||
"revenue_growth": stmt.excluded.revenue_growth,
|
||||
"earnings_surprise": stmt.excluded.earnings_surprise,
|
||||
"market_cap": stmt.excluded.market_cap,
|
||||
"next_earnings_date": stmt.excluded.next_earnings_date,
|
||||
"fetched_at": stmt.excluded.fetched_at,
|
||||
"unavailable_fields_json": (
|
||||
stmt.excluded.unavailable_fields_json
|
||||
),
|
||||
},
|
||||
)
|
||||
)
|
||||
|
||||
dimension_ids = await _fundamental_dimension_ids(db, changed_ids)
|
||||
composite_ids = await _composite_ids(db, changed_ids)
|
||||
if dimension_ids:
|
||||
await db.execute(
|
||||
update(DimensionScore)
|
||||
.where(DimensionScore.ticker_id.in_(dimension_ids))
|
||||
.values(is_stale=True)
|
||||
)
|
||||
if composite_ids:
|
||||
await db.execute(
|
||||
update(CompositeScore)
|
||||
.where(CompositeScore.ticker_id.in_(composite_ids))
|
||||
.values(is_stale=True)
|
||||
)
|
||||
|
||||
await db.commit()
|
||||
return {
|
||||
"enabled": True,
|
||||
"refreshed": len(candidates),
|
||||
"score_inputs_changed": len(changed_ids),
|
||||
"dimension_scores_staled": len(dimension_ids),
|
||||
"composite_scores_staled": len(composite_ids),
|
||||
}
|
||||
|
||||
|
||||
async def _existing_by_ticker(
|
||||
db: AsyncSession, ticker_ids: list[int]
|
||||
) -> dict[int, FundamentalData]:
|
||||
if not ticker_ids:
|
||||
return {}
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(FundamentalData).where(
|
||||
FundamentalData.ticker_id.in_(ticker_ids)
|
||||
)
|
||||
)
|
||||
).scalars()
|
||||
return {row.ticker_id: row for row in rows}
|
||||
|
||||
|
||||
async def _fundamental_dimension_ids(
|
||||
db: AsyncSession, ticker_ids: set[int]
|
||||
) -> set[int]:
|
||||
if not ticker_ids:
|
||||
return set()
|
||||
rows = await db.execute(
|
||||
select(DimensionScore.ticker_id).where(
|
||||
DimensionScore.ticker_id.in_(ticker_ids),
|
||||
DimensionScore.dimension == "fundamental",
|
||||
)
|
||||
)
|
||||
return set(rows.scalars())
|
||||
|
||||
|
||||
async def _composite_ids(
|
||||
db: AsyncSession, ticker_ids: set[int]
|
||||
) -> set[int]:
|
||||
if not ticker_ids:
|
||||
return set()
|
||||
rows = await db.execute(
|
||||
select(CompositeScore.ticker_id).where(
|
||||
CompositeScore.ticker_id.in_(ticker_ids)
|
||||
)
|
||||
)
|
||||
return set(rows.scalars())
|
||||
|
||||
|
||||
def _score_inputs_changed(
|
||||
existing: FundamentalData | None,
|
||||
candidate: fundamentals_candidate_service.CandidateFundamentals,
|
||||
) -> bool:
|
||||
if existing is None:
|
||||
return True
|
||||
return any(
|
||||
getattr(existing, field) != getattr(candidate, field)
|
||||
for field in _SCORE_FIELDS
|
||||
)
|
||||
@@ -121,6 +121,7 @@ def _build_metrics(derived, peer_derived, two: str | None) -> list[dict[str, Any
|
||||
"industry": industry,
|
||||
"period_end": _iso(series.period_end) if series else None,
|
||||
"filed_date": _iso(series.filed_date) if series else None,
|
||||
"caveat": series.caveat if series else None,
|
||||
"source": "sec",
|
||||
})
|
||||
return out
|
||||
@@ -154,6 +155,17 @@ def _build_valuation(derived, subject_price, peer_derived, peer_price_by_cik, tw
|
||||
"pe": _round(pe, 2),
|
||||
"fcf_yield": _round(fcf_yield, 2),
|
||||
"market_cap_est": _round(market_cap, 0),
|
||||
# market_cap_est and fcf_yield both rest on the share count. When it came
|
||||
# from the weighted-average diluted fallback (multi-class issuers, whose
|
||||
# per-class cover-page count is absent from companyfacts), say so rather
|
||||
# than presenting a period average as a point-in-time count.
|
||||
"shares_estimated": bool(
|
||||
market_cap is not None and derived.shares_outstanding_estimated
|
||||
),
|
||||
# A null P/E is ambiguous: no earnings data, or earnings we deliberately
|
||||
# suppressed. Only the latter carries a caveat, so a split-contaminated
|
||||
# TTM says why instead of looking like missing data.
|
||||
"pe_caveat": derived.ttm_diluted_eps_caveat if pe is None else None,
|
||||
"pe_industry": pe_industry,
|
||||
"fcf_yield_industry": fcf_yield_industry,
|
||||
"price_date": _iso(price_date),
|
||||
@@ -303,7 +315,8 @@ async def _latest_close(db, ticker_id: int) -> tuple[float, date] | None:
|
||||
|
||||
def _empty_metrics() -> list[dict[str, Any]]:
|
||||
return [{"key": k, "value": None, "history": [], "industry": None,
|
||||
"period_end": None, "filed_date": None, "source": "sec"} for k in METRIC_KEYS]
|
||||
"period_end": None, "filed_date": None, "caveat": None,
|
||||
"source": "sec"} for k in METRIC_KEYS]
|
||||
|
||||
|
||||
def _empty_earnings() -> dict[str, Any]:
|
||||
|
||||
@@ -0,0 +1,288 @@
|
||||
"""Local SEC/Dolt candidate values for the legacy fundamentals cache.
|
||||
|
||||
This is the single read path shared by the A5 parity report and the activated
|
||||
``fundamental_data`` refresh. It never contacts SEC or Dolt: every input comes
|
||||
from PostgreSQL, so price- and earnings-driven values can still refresh when an
|
||||
upstream import is unchanged or unavailable.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
from collections import defaultdict
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import date, datetime
|
||||
from typing import Any
|
||||
from zoneinfo import ZoneInfo
|
||||
|
||||
from sqlalchemy import func, select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.models.earnings_event import EarningsEvent
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.ticker import Ticker
|
||||
from app.services import fundamentals_derivation as deriv
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class CandidateFundamentals:
|
||||
ticker_id: int
|
||||
symbol: str
|
||||
cik: str | None
|
||||
pe_ratio: float | None
|
||||
revenue_growth: float | None
|
||||
earnings_surprise: float | None
|
||||
market_cap: float | None
|
||||
next_earnings_date: date | None
|
||||
price_date: date | None
|
||||
unavailable_fields: dict[str, str] = field(default_factory=dict)
|
||||
|
||||
|
||||
async def build_candidates(
|
||||
db: AsyncSession,
|
||||
*,
|
||||
today: date | None = None,
|
||||
) -> list[CandidateFundamentals]:
|
||||
"""Derive current cache candidates using only already-stored data."""
|
||||
today = today or datetime.now(ZoneInfo("America/New_York")).date()
|
||||
tickers = list(
|
||||
(await db.execute(select(Ticker).order_by(Ticker.symbol))).scalars()
|
||||
)
|
||||
if not tickers:
|
||||
return []
|
||||
|
||||
ticker_ids = [ticker.id for ticker in tickers]
|
||||
ciks = sorted({ticker.cik for ticker in tickers if ticker.cik})
|
||||
derived_by_cik = await _derived_by_cik(db, ciks)
|
||||
closes_by_ticker = await _latest_closes(db, ticker_ids)
|
||||
surprise_by_ticker, next_by_ticker = await _earnings_values(
|
||||
db, ticker_ids, today
|
||||
)
|
||||
|
||||
out: list[CandidateFundamentals] = []
|
||||
for ticker in tickers:
|
||||
derived = derived_by_cik.get(ticker.cik) if ticker.cik else None
|
||||
close = closes_by_ticker.get(ticker.id)
|
||||
price = close[0] if close is not None else None
|
||||
price_date = close[1] if close is not None else None
|
||||
growth_series = (
|
||||
derived.metrics.get("revenue_growth_yoy")
|
||||
if derived is not None
|
||||
else None
|
||||
)
|
||||
|
||||
pe_ratio = (
|
||||
_pe(price, derived.ttm_diluted_eps)
|
||||
if derived is not None
|
||||
else None
|
||||
)
|
||||
revenue_growth = (
|
||||
float(growth_series.value)
|
||||
if growth_series is not None and _finite(growth_series.value)
|
||||
else None
|
||||
)
|
||||
earnings_surprise = surprise_by_ticker.get(ticker.id)
|
||||
market_cap = (
|
||||
_market_cap(price, derived.shares_outstanding)
|
||||
if derived is not None
|
||||
else None
|
||||
)
|
||||
next_earnings_date = next_by_ticker.get(ticker.id)
|
||||
|
||||
out.append(
|
||||
CandidateFundamentals(
|
||||
ticker_id=ticker.id,
|
||||
symbol=ticker.symbol,
|
||||
cik=ticker.cik,
|
||||
pe_ratio=pe_ratio,
|
||||
revenue_growth=revenue_growth,
|
||||
earnings_surprise=earnings_surprise,
|
||||
market_cap=market_cap,
|
||||
next_earnings_date=next_earnings_date,
|
||||
price_date=price_date,
|
||||
unavailable_fields=_availability_metadata(
|
||||
derived=derived,
|
||||
price=price,
|
||||
pe_ratio=pe_ratio,
|
||||
revenue_growth=revenue_growth,
|
||||
earnings_surprise=earnings_surprise,
|
||||
market_cap=market_cap,
|
||||
next_earnings_date=next_earnings_date,
|
||||
),
|
||||
)
|
||||
)
|
||||
return out
|
||||
|
||||
|
||||
async def _derived_by_cik(
|
||||
db: AsyncSession, ciks: list[str]
|
||||
) -> dict[str, deriv.DerivedFundamentals]:
|
||||
if not ciks:
|
||||
return {}
|
||||
grouped: dict[str, list[FundamentalSnapshot]] = defaultdict(list)
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(FundamentalSnapshot).where(FundamentalSnapshot.cik.in_(ciks))
|
||||
)
|
||||
).scalars()
|
||||
for row in rows:
|
||||
grouped[row.cik].append(row)
|
||||
return {cik: deriv.derive(grouped.get(cik, [])) for cik in ciks}
|
||||
|
||||
|
||||
async def _latest_closes(
|
||||
db: AsyncSession, ticker_ids: list[int]
|
||||
) -> dict[int, tuple[float, date]]:
|
||||
latest = (
|
||||
select(
|
||||
OHLCVRecord.ticker_id,
|
||||
func.max(OHLCVRecord.date).label("max_date"),
|
||||
)
|
||||
.where(OHLCVRecord.ticker_id.in_(ticker_ids))
|
||||
.group_by(OHLCVRecord.ticker_id)
|
||||
.subquery()
|
||||
)
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(
|
||||
OHLCVRecord.ticker_id,
|
||||
OHLCVRecord.close,
|
||||
OHLCVRecord.date,
|
||||
).join(
|
||||
latest,
|
||||
(OHLCVRecord.ticker_id == latest.c.ticker_id)
|
||||
& (OHLCVRecord.date == latest.c.max_date),
|
||||
)
|
||||
)
|
||||
).all()
|
||||
return {
|
||||
ticker_id: (float(close), close_date)
|
||||
for ticker_id, close, close_date in rows
|
||||
if _finite(close)
|
||||
}
|
||||
|
||||
|
||||
async def _earnings_values(
|
||||
db: AsyncSession,
|
||||
ticker_ids: list[int],
|
||||
today: date,
|
||||
) -> tuple[dict[int, float], dict[int, date]]:
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(EarningsEvent)
|
||||
.where(EarningsEvent.ticker_id.in_(ticker_ids))
|
||||
.order_by(EarningsEvent.ticker_id, EarningsEvent.announce_date.desc())
|
||||
)
|
||||
).scalars()
|
||||
surprises: dict[int, float] = {}
|
||||
upcoming: dict[int, date] = {}
|
||||
for row in rows:
|
||||
if row.announce_date >= today:
|
||||
current = upcoming.get(row.ticker_id)
|
||||
if current is None or row.announce_date < current:
|
||||
upcoming[row.ticker_id] = row.announce_date
|
||||
continue
|
||||
if row.ticker_id in surprises:
|
||||
continue
|
||||
surprise = _surprise(row.eps_estimate, row.eps_actual)
|
||||
if surprise is not None:
|
||||
surprises[row.ticker_id] = surprise
|
||||
return surprises, upcoming
|
||||
|
||||
|
||||
def _availability_metadata(
|
||||
*,
|
||||
derived: deriv.DerivedFundamentals | None,
|
||||
price: float | None,
|
||||
pe_ratio: float | None,
|
||||
revenue_growth: float | None,
|
||||
earnings_surprise: float | None,
|
||||
market_cap: float | None,
|
||||
next_earnings_date: date | None,
|
||||
) -> dict[str, str]:
|
||||
metadata: dict[str, str] = {}
|
||||
|
||||
if pe_ratio is not None:
|
||||
metadata["source_pe_ratio"] = "sec_facts+ohlcv_records"
|
||||
elif derived is None or derived.latest_period_end is None:
|
||||
metadata["pe_ratio"] = "no SEC fundamental snapshots"
|
||||
elif not _finite(price) or price <= 0:
|
||||
metadata["pe_ratio"] = "no usable PostgreSQL close"
|
||||
elif derived.ttm_diluted_eps_caveat:
|
||||
metadata["pe_ratio"] = derived.ttm_diluted_eps_caveat
|
||||
else:
|
||||
metadata["pe_ratio"] = "no positive SEC-derived TTM diluted EPS"
|
||||
|
||||
if revenue_growth is not None:
|
||||
metadata["source_revenue_growth"] = "sec_facts"
|
||||
else:
|
||||
metadata["revenue_growth"] = "SEC-derived TTM revenue growth unavailable"
|
||||
|
||||
if earnings_surprise is not None:
|
||||
metadata["source_earnings_surprise"] = "dolt_earnings"
|
||||
else:
|
||||
metadata["earnings_surprise"] = (
|
||||
"no completed earnings event with actual and nonzero estimate"
|
||||
)
|
||||
|
||||
if market_cap is not None:
|
||||
metadata["source_market_cap"] = "sec_facts+ohlcv_records"
|
||||
if derived is not None and derived.shares_outstanding_estimated:
|
||||
metadata["market_cap_estimated"] = (
|
||||
"shares use the SEC weighted-average diluted fallback"
|
||||
)
|
||||
elif derived is None or derived.latest_period_end is None:
|
||||
metadata["market_cap"] = "no SEC fundamental snapshots"
|
||||
elif not _finite(price) or price <= 0:
|
||||
metadata["market_cap"] = "no usable PostgreSQL close"
|
||||
else:
|
||||
metadata["market_cap"] = "SEC-derived shares outstanding unavailable"
|
||||
|
||||
if next_earnings_date is not None:
|
||||
metadata["source_next_earnings_date"] = "dolt_earnings"
|
||||
else:
|
||||
metadata["next_earnings_date"] = "no upcoming earnings event"
|
||||
return metadata
|
||||
|
||||
|
||||
def _surprise(
|
||||
estimate: float | None,
|
||||
actual: float | None,
|
||||
) -> float | None:
|
||||
if not _finite(estimate) or not _finite(actual) or estimate == 0:
|
||||
return None
|
||||
return (float(actual) - float(estimate)) / abs(float(estimate)) * 100.0
|
||||
|
||||
|
||||
def _pe(price: float | None, ttm_eps: float | None) -> float | None:
|
||||
if (
|
||||
not _finite(price)
|
||||
or price <= 0
|
||||
or not _finite(ttm_eps)
|
||||
or ttm_eps <= 0
|
||||
):
|
||||
return None
|
||||
return float(price) / float(ttm_eps)
|
||||
|
||||
|
||||
def _market_cap(
|
||||
price: float | None,
|
||||
shares_outstanding: float | None,
|
||||
) -> float | None:
|
||||
if (
|
||||
not _finite(price)
|
||||
or price <= 0
|
||||
or not _finite(shares_outstanding)
|
||||
or shares_outstanding <= 0
|
||||
):
|
||||
return None
|
||||
return float(price) * float(shares_outstanding)
|
||||
|
||||
|
||||
def _finite(value: Any) -> bool:
|
||||
return (
|
||||
isinstance(value, (int, float))
|
||||
and not isinstance(value, bool)
|
||||
and math.isfinite(value)
|
||||
)
|
||||
@@ -8,8 +8,10 @@ schema decision. No I/O, no DB: it takes an issuer's snapshot rows (ORM rows or
|
||||
any objects with the same attributes) and returns structured metrics.
|
||||
|
||||
Rules:
|
||||
- **Amendment selection:** for each (fiscal_year, fiscal_period), the row with
|
||||
the newest `accepted_at` wins.
|
||||
- **Amendment selection:** for each (fiscal_year, fiscal_period), the newest
|
||||
`accepted_at` wins **per field**, falling back to the newest row that actually
|
||||
reports one. A partial amendment (a 10-K/A adding Part III carries no financial
|
||||
facts) must not blank the period.
|
||||
- **Discrete quarter** = YTD(Qn) − YTD(Qn−1); Q1 = YTD(Q1); **Q4 = YTD(FY) −
|
||||
YTD(Q3)**. Any missing period → the derived value is null, never partial.
|
||||
- **TTM** = sum of the trailing four discrete quarters ending at a period.
|
||||
@@ -21,18 +23,38 @@ from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import date
|
||||
from types import SimpleNamespace
|
||||
from typing import Any, Iterable
|
||||
|
||||
_FP_TO_Q = {"Q1": 1, "Q2": 2, "Q3": 3, "FY": 4}
|
||||
_Q_TO_FP = {1: "Q1", 2: "Q2", 3: "Q3", 4: "FY"}
|
||||
_PREV_FP = {"Q2": "Q1", "Q3": "Q2", "FY": "Q3"}
|
||||
TAPE_LEN = 4 # quarter-tape length
|
||||
SPLIT_SUSPECT_SHARE_CHANGE_PCT = 25.0
|
||||
SPLIT_SENSITIVE_CAVEAT = (
|
||||
"Not comparable: share count changed at least 25%; possible split or "
|
||||
"corporate action."
|
||||
)
|
||||
|
||||
# Duration (flow) fields differenced from YTD into discrete quarters + summed to TTM.
|
||||
_FLOW_FIELDS = (
|
||||
"revenue", "net_income", "operating_income", "diluted_eps", "cfo", "capex",
|
||||
"depreciation_amortization",
|
||||
)
|
||||
# Reported facts resolved independently across a period's accessions (see
|
||||
# _merge_amendments); period identity/provenance is taken from the newest one.
|
||||
_MERGED_FIELDS = (
|
||||
*_FLOW_FIELDS,
|
||||
"cash_and_st_investments", "total_debt", "shares_outstanding",
|
||||
"shares_outstanding_date", "weighted_avg_diluted_shares",
|
||||
# period_start is set alongside revenue by the parser, so it follows the same
|
||||
# fallback: a bare amendment reports neither and must not blank it.
|
||||
"period_start",
|
||||
)
|
||||
_CARRIED_FIELDS = (
|
||||
"fiscal_year", "fiscal_period", "period_end", "filed_date",
|
||||
"accepted_at", "form", "accession", "cik",
|
||||
)
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -47,6 +69,7 @@ class MetricSeries:
|
||||
history: list[MetricPoint] = field(default_factory=list) # oldest -> newest, <= TAPE_LEN
|
||||
period_end: date | None = None
|
||||
filed_date: date | None = None
|
||||
caveat: str | None = None
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -54,8 +77,15 @@ class DerivedFundamentals:
|
||||
metrics: dict[str, MetricSeries] = field(default_factory=dict)
|
||||
# request-time valuation inputs (ratios are computed in the API with price)
|
||||
ttm_diluted_eps: float | None = None
|
||||
# Set when ttm_diluted_eps was suppressed rather than simply unavailable.
|
||||
ttm_diluted_eps_caveat: str | None = None
|
||||
ttm_fcf: float | None = None
|
||||
shares_outstanding: float | None = None
|
||||
# True when shares_outstanding came from the weighted-average diluted count
|
||||
# because the point-in-time cover-page count was absent (always so for
|
||||
# multi-class issuers). Consumers must label anything derived from it as
|
||||
# estimated — it is a period average, not a point-in-time count.
|
||||
shares_outstanding_estimated: bool = False
|
||||
latest_period_end: date | None = None
|
||||
latest_filed_date: date | None = None
|
||||
|
||||
@@ -79,6 +109,15 @@ def derive(snapshots: Iterable[Any]) -> DerivedFundamentals:
|
||||
result.latest_period_end = latest_row.period_end
|
||||
result.latest_filed_date = latest_row.filed_date
|
||||
result.shares_outstanding = getattr(latest_row, "shares_outstanding", None)
|
||||
if result.shares_outstanding is None:
|
||||
# Multi-class issuers (META, CMCSA, BRK-B, CHTR, FOXA, NWSA, LEN) report
|
||||
# the cover-page count per class, which is dimensional and so absent from
|
||||
# companyfacts — leaving market cap and FCF yield silently unavailable for
|
||||
# some of the largest names. The weighted-average diluted count is always
|
||||
# present and within ~0.6% of the true count where both exist, so fall
|
||||
# back to it and mark the result estimated rather than show nothing.
|
||||
result.shares_outstanding = getattr(latest_row, "weighted_avg_diluted_shares", None)
|
||||
result.shares_outstanding_estimated = result.shares_outstanding is not None
|
||||
result.ttm_diluted_eps = _ttm(discrete["diluted_eps"], *latest)
|
||||
ttm_cfo = _ttm(discrete["cfo"], *latest)
|
||||
ttm_capex = _ttm(discrete["capex"], *latest)
|
||||
@@ -96,6 +135,14 @@ def derive(snapshots: Iterable[Any]) -> DerivedFundamentals:
|
||||
"net_debt_to_ebitda": _leverage_series(selected, discrete, tape),
|
||||
"share_count_change_yoy": _share_change_series(selected, tape),
|
||||
}
|
||||
# TTM EPS sums four quarters of *per-share* values, so a split inside that
|
||||
# window mixes pre- and post-split units — the same distortion the guard
|
||||
# already catches for the series, and the one that produced BKNG's P/E of
|
||||
# 1.10. Left unguarded it does not merely mislead: a nonsense-low P/E clamps
|
||||
# to a perfect 100 fundamental sub-score, so it must null out like the rest.
|
||||
if _guard_split_sensitive_metrics(result.metrics):
|
||||
result.ttm_diluted_eps = None
|
||||
result.ttm_diluted_eps_caveat = SPLIT_SENSITIVE_CAVEAT
|
||||
for series in result.metrics.values():
|
||||
series.period_end = latest_row.period_end
|
||||
series.filed_date = latest_row.filed_date
|
||||
@@ -105,17 +152,57 @@ def derive(snapshots: Iterable[Any]) -> DerivedFundamentals:
|
||||
# -- period selection --------------------------------------------------------
|
||||
|
||||
def _select_latest_per_period(snapshots: Iterable[Any]) -> dict[tuple[int, str], Any]:
|
||||
best: dict[tuple[int, str], Any] = {}
|
||||
grouped: dict[tuple[int, str], list[Any]] = {}
|
||||
for row in snapshots:
|
||||
fp = getattr(row, "fiscal_period", None)
|
||||
fy = getattr(row, "fiscal_year", None)
|
||||
if fp not in _FP_TO_Q or fy is None:
|
||||
continue
|
||||
key = (fy, fp)
|
||||
cur = best.get(key)
|
||||
if cur is None or _accepted(row) > _accepted(cur):
|
||||
best[key] = row
|
||||
return best
|
||||
grouped.setdefault((fy, fp), []).append(row)
|
||||
return {key: _merge_amendments(rows) for key, rows in grouped.items()}
|
||||
|
||||
|
||||
def _merge_amendments(rows: list[Any]) -> Any:
|
||||
"""Resolve one period from its accessions: newest wins, per field.
|
||||
|
||||
Amendments are frequently partial — a 10-K/A filed only to add Part III
|
||||
reports no financial facts at all. Taking the newest accession wholesale
|
||||
would blank every field it omits and null the period downstream (and with
|
||||
it TTM and YoY, which need an unbroken quarter chain), so each field falls
|
||||
back to the newest accession that actually reports it.
|
||||
|
||||
Only rows sharing the newest row's ``period_end`` are merged. A same-key row
|
||||
covering a *different* period is a mislabelled filing, not an amendment, and
|
||||
blending the two would silently mix fiscal years.
|
||||
"""
|
||||
if len(rows) == 1:
|
||||
return rows[0]
|
||||
ordered = sorted(rows, key=_amendment_order, reverse=True) # newest first
|
||||
newest = ordered[0]
|
||||
same_period = [
|
||||
row
|
||||
for row in ordered
|
||||
if getattr(row, "period_end", None) == getattr(newest, "period_end", None)
|
||||
]
|
||||
if len(same_period) == 1:
|
||||
return newest
|
||||
merged = SimpleNamespace(**{name: getattr(newest, name, None) for name in _CARRIED_FIELDS})
|
||||
for name in _MERGED_FIELDS:
|
||||
merged_value = None
|
||||
for row in same_period: # newest first
|
||||
value = getattr(row, name, None)
|
||||
if value is not None:
|
||||
merged_value = value
|
||||
break
|
||||
setattr(merged, name, merged_value)
|
||||
return merged
|
||||
|
||||
|
||||
def _amendment_order(row: Any) -> tuple[bool, Any]:
|
||||
# (has-timestamp, timestamp) so a row without one sorts oldest instead of
|
||||
# raising when compared against a row that has one.
|
||||
accepted = _accepted(row)
|
||||
return (accepted is not None, accepted)
|
||||
|
||||
|
||||
def _accepted(row: Any):
|
||||
@@ -248,6 +335,46 @@ def _share_change_series(selected, tape) -> MetricSeries:
|
||||
return _series(pts)
|
||||
|
||||
|
||||
def _guard_split_sensitive_metrics(metrics: dict[str, MetricSeries]) -> bool:
|
||||
"""Suppress historical comparisons likely distorted by a corporate action.
|
||||
|
||||
Company Facts has no point-in-time split factors. A large YoY share-count
|
||||
move can therefore make both the point-in-time share comparison and
|
||||
per-share EPS growth non-comparable. Keep the raw facts in snapshots, but
|
||||
expose nulls plus an explicit caveat in the user-facing derived series.
|
||||
|
||||
Returns True when the *latest* period is suspect, so callers can apply the
|
||||
same suppression to per-share scalars derived from that window.
|
||||
"""
|
||||
shares = metrics.get("share_count_change_yoy")
|
||||
eps = metrics.get("eps_growth_yoy")
|
||||
if shares is None or eps is None:
|
||||
return False
|
||||
|
||||
suspect_periods = {
|
||||
point.period_end
|
||||
for point in shares.history
|
||||
if point.value is not None
|
||||
and abs(point.value) >= SPLIT_SUSPECT_SHARE_CHANGE_PCT
|
||||
}
|
||||
if not suspect_periods:
|
||||
return False
|
||||
|
||||
latest_suspect = False
|
||||
for series in (shares, eps):
|
||||
latest_guarded = bool(
|
||||
series.history and series.history[-1].period_end in suspect_periods
|
||||
)
|
||||
latest_suspect = latest_suspect or latest_guarded
|
||||
for point in series.history:
|
||||
if point.period_end in suspect_periods:
|
||||
point.value = None
|
||||
series.value = series.history[-1].value if series.history else None
|
||||
if latest_guarded:
|
||||
series.caveat = SPLIT_SENSITIVE_CAVEAT
|
||||
return latest_suspect
|
||||
|
||||
|
||||
def _net_debt(row: Any) -> float | None:
|
||||
if row is None:
|
||||
return None
|
||||
|
||||
@@ -0,0 +1,498 @@
|
||||
"""Read-only A5 comparison of legacy and SEC/Dolt fundamental inputs.
|
||||
|
||||
The report deliberately does not write ``fundamental_data`` or score tables.
|
||||
It reconstructs the current legacy and candidate fundamental scores, projects
|
||||
their composite-score/rank effect with the active weights, and archives a
|
||||
timestamped JSON + CSV bundle for explicit human approval.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import csv
|
||||
import io
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import statistics
|
||||
from datetime import date, datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any, Iterable
|
||||
from zoneinfo import ZoneInfo
|
||||
|
||||
from sqlalchemy import select, text
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.models.data_import_run import DataImportRun
|
||||
from app.models.fundamental import FundamentalData
|
||||
from app.services import fundamentals_candidate_service as candidate_service
|
||||
|
||||
REPORT_VERSION = 1
|
||||
APPROVAL_STATUS = "pending_explicit_approval"
|
||||
FIELD_KEYS = ("pe_ratio", "revenue_growth", "earnings_surprise")
|
||||
MIN_SCORE_METRICS = 2
|
||||
|
||||
# Materiality is a review aid, never an automatic cutover verdict. Definition
|
||||
# changes remain visible even when a delta falls inside these bands.
|
||||
FIELD_TOLERANCES = {
|
||||
"pe_ratio": {"absolute": 1.0, "relative_pct": 10.0},
|
||||
"revenue_growth": {"absolute": 2.0, "relative_pct": None},
|
||||
"earnings_surprise": {"absolute": 2.0, "relative_pct": None},
|
||||
}
|
||||
DEFINITION_NOTES = {
|
||||
"pe_ratio": (
|
||||
"Legacy provider P/E convention versus latest close divided by "
|
||||
"SEC-derived TTM diluted EPS."
|
||||
),
|
||||
"revenue_growth": (
|
||||
"Legacy provider growth convention versus SEC-derived TTM revenue YoY."
|
||||
),
|
||||
"earnings_surprise": (
|
||||
"Legacy provider latest surprise versus latest completed Dolt earnings "
|
||||
"event with actual and estimate."
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def fundamental_score(
|
||||
pe_ratio: float | None,
|
||||
revenue_growth: float | None,
|
||||
earnings_surprise: float | None,
|
||||
) -> float | None:
|
||||
"""Match the production fundamental-dimension formula without persistence."""
|
||||
scores: list[float] = []
|
||||
if _finite(pe_ratio) and pe_ratio > 0:
|
||||
scores.append(max(0.0, min(100.0, 100.0 - (pe_ratio - 15.0) * (100.0 / 30.0))))
|
||||
if _finite(revenue_growth):
|
||||
scores.append(max(0.0, min(100.0, 50.0 + revenue_growth * 2.5)))
|
||||
if _finite(earnings_surprise):
|
||||
scores.append(max(0.0, min(100.0, 50.0 + earnings_surprise * 5.0)))
|
||||
return sum(scores) / len(scores) if len(scores) >= MIN_SCORE_METRICS else None
|
||||
|
||||
|
||||
async def build_report(
|
||||
db: AsyncSession,
|
||||
*,
|
||||
generated_at: datetime | None = None,
|
||||
today: date | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Build a point-in-time parity report from one database session."""
|
||||
generated_at = generated_at or datetime.now(timezone.utc)
|
||||
today = today or datetime.now(ZoneInfo("America/New_York")).date()
|
||||
|
||||
# A report must not mix rows from before and after a concurrent import
|
||||
# promotion. The scheduled job provides a fresh session, so establish the
|
||||
# production snapshot before its first query and have Postgres enforce the
|
||||
# no-write contract as well. SQLite tests retain their normal transaction.
|
||||
if db.get_bind().dialect.name == "postgresql":
|
||||
connection = await db.connection(
|
||||
execution_options={"isolation_level": "REPEATABLE READ"}
|
||||
)
|
||||
await connection.execute(text("SET TRANSACTION READ ONLY"))
|
||||
|
||||
candidates = await candidate_service.build_candidates(db, today=today)
|
||||
ticker_ids = [candidate.ticker_id for candidate in candidates]
|
||||
legacy_by_ticker = await _legacy_values(db, ticker_ids)
|
||||
source_runs = await _source_runs(db)
|
||||
|
||||
rows: list[dict[str, Any]] = []
|
||||
for candidate in candidates:
|
||||
legacy = legacy_by_ticker.get(candidate.ticker_id)
|
||||
candidate_values = {
|
||||
"pe_ratio": candidate.pe_ratio,
|
||||
"revenue_growth": candidate.revenue_growth,
|
||||
"earnings_surprise": candidate.earnings_surprise,
|
||||
}
|
||||
legacy_values = {
|
||||
"pe_ratio": legacy.pe_ratio if legacy else None,
|
||||
"revenue_growth": legacy.revenue_growth if legacy else None,
|
||||
"earnings_surprise": legacy.earnings_surprise if legacy else None,
|
||||
}
|
||||
fields = {
|
||||
key: _field_comparison(key, legacy_values[key], candidate_values[key])
|
||||
for key in FIELD_KEYS
|
||||
}
|
||||
legacy_score = fundamental_score(**legacy_values)
|
||||
candidate_score = fundamental_score(**candidate_values)
|
||||
rows.append(
|
||||
{
|
||||
"symbol": candidate.symbol,
|
||||
"cik": candidate.cik,
|
||||
"legacy_fetched_at": _iso(legacy.fetched_at) if legacy else None,
|
||||
"price_date": _iso(candidate.price_date),
|
||||
"fields": fields,
|
||||
"scores": {
|
||||
"legacy_fundamental": _round(legacy_score),
|
||||
"candidate_fundamental": _round(candidate_score),
|
||||
"fundamental_delta": _delta(legacy_score, candidate_score),
|
||||
"legacy_fundamental_rank": None,
|
||||
"candidate_fundamental_rank": None,
|
||||
"fundamental_rank_change": None,
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
_attach_ranks(rows, "legacy_fundamental", "legacy_fundamental_rank")
|
||||
_attach_ranks(rows, "candidate_fundamental", "candidate_fundamental_rank")
|
||||
for row in rows:
|
||||
scores = row["scores"]
|
||||
scores["fundamental_rank_change"] = _rank_change(
|
||||
scores["legacy_fundamental_rank"], scores["candidate_fundamental_rank"]
|
||||
)
|
||||
|
||||
return {
|
||||
"report_version": REPORT_VERSION,
|
||||
"generated_at": generated_at.isoformat(),
|
||||
"as_of_date": today.isoformat(),
|
||||
"approval_status": APPROVAL_STATUS,
|
||||
"read_only": True,
|
||||
"fundamental_score_formula": (
|
||||
"Equal-weighted mean of 2+ available sub-scores: P/E = "
|
||||
"clamp(100-(pe-15)*(100/30)); revenue growth = "
|
||||
"clamp(50+growth*2.5); earnings surprise = "
|
||||
"clamp(50+surprise*5)."
|
||||
),
|
||||
"source_runs": source_runs,
|
||||
"definition_notes": DEFINITION_NOTES,
|
||||
"materiality_notes": {
|
||||
"fields": FIELD_TOLERANCES,
|
||||
"fundamental_score_absolute": 5.0,
|
||||
"automatic_cutover": False,
|
||||
},
|
||||
"summary": _summary(rows),
|
||||
"rows": rows,
|
||||
}
|
||||
|
||||
|
||||
def store_report(report: dict[str, Any], report_dir: str | Path) -> dict[str, str]:
|
||||
"""Atomically archive JSON/CSV artifacts and update the latest manifest."""
|
||||
directory = Path(report_dir).expanduser().resolve()
|
||||
directory.mkdir(parents=True, exist_ok=True)
|
||||
stamp = _artifact_stamp(report["generated_at"])
|
||||
json_name = f"fundamentals-parity-{stamp}.json"
|
||||
csv_name = f"fundamentals-parity-{stamp}.csv"
|
||||
json_path = directory / json_name
|
||||
csv_path = directory / csv_name
|
||||
|
||||
_atomic_write(json_path, json.dumps(report, indent=2, sort_keys=True) + "\n")
|
||||
_atomic_write(csv_path, report_csv(report))
|
||||
manifest = {
|
||||
"generated_at": report["generated_at"],
|
||||
"json_file": json_name,
|
||||
"csv_file": csv_name,
|
||||
}
|
||||
_atomic_write(
|
||||
directory / "latest.json",
|
||||
json.dumps(manifest, indent=2, sort_keys=True) + "\n",
|
||||
)
|
||||
return {
|
||||
"json": str(json_path),
|
||||
"csv": str(csv_path),
|
||||
"manifest": str(directory / "latest.json"),
|
||||
}
|
||||
|
||||
|
||||
async def generate_and_store(
|
||||
db: AsyncSession,
|
||||
report_dir: str | Path,
|
||||
*,
|
||||
generated_at: datetime | None = None,
|
||||
today: date | None = None,
|
||||
) -> tuple[dict[str, Any], dict[str, str]]:
|
||||
report = await build_report(db, generated_at=generated_at, today=today)
|
||||
return report, store_report(report, report_dir)
|
||||
|
||||
|
||||
def load_latest(report_dir: str | Path) -> dict[str, Any] | None:
|
||||
manifest = _load_manifest(report_dir)
|
||||
if manifest is None:
|
||||
return None
|
||||
try:
|
||||
path = _manifest_artifact(report_dir, manifest, "json_file")
|
||||
loaded = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError, TypeError, ValueError):
|
||||
return None
|
||||
return loaded if isinstance(loaded, dict) else None
|
||||
|
||||
|
||||
def load_latest_csv(report_dir: str | Path) -> tuple[str, str] | None:
|
||||
return _load_latest_text_artifact(report_dir, "csv_file")
|
||||
|
||||
|
||||
def load_latest_json(report_dir: str | Path) -> tuple[str, str] | None:
|
||||
return _load_latest_text_artifact(report_dir, "json_file")
|
||||
|
||||
|
||||
def _load_latest_text_artifact(
|
||||
report_dir: str | Path, manifest_key: str
|
||||
) -> tuple[str, str] | None:
|
||||
manifest = _load_manifest(report_dir)
|
||||
if manifest is None:
|
||||
return None
|
||||
try:
|
||||
path = _manifest_artifact(report_dir, manifest, manifest_key)
|
||||
return path.name, path.read_text(encoding="utf-8")
|
||||
except (OSError, TypeError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
def report_csv(report: dict[str, Any]) -> str:
|
||||
output = io.StringIO(newline="")
|
||||
columns = [
|
||||
"symbol",
|
||||
"cik",
|
||||
"legacy_fetched_at",
|
||||
"price_date",
|
||||
*(
|
||||
f"{field}_{suffix}"
|
||||
for field in FIELD_KEYS
|
||||
for suffix in ("legacy", "candidate", "absolute_delta", "relative_delta_pct", "material")
|
||||
),
|
||||
"legacy_fundamental",
|
||||
"candidate_fundamental",
|
||||
"fundamental_delta",
|
||||
"legacy_fundamental_rank",
|
||||
"candidate_fundamental_rank",
|
||||
"fundamental_rank_change",
|
||||
]
|
||||
writer = csv.DictWriter(output, fieldnames=columns)
|
||||
writer.writeheader()
|
||||
for row in report.get("rows", []):
|
||||
flat = {
|
||||
"symbol": row["symbol"],
|
||||
"cik": row.get("cik"),
|
||||
"legacy_fetched_at": row.get("legacy_fetched_at"),
|
||||
"price_date": row.get("price_date"),
|
||||
**row["scores"],
|
||||
}
|
||||
for field in FIELD_KEYS:
|
||||
comparison = row["fields"][field]
|
||||
for suffix in (
|
||||
"legacy",
|
||||
"candidate",
|
||||
"absolute_delta",
|
||||
"relative_delta_pct",
|
||||
"material",
|
||||
):
|
||||
flat[f"{field}_{suffix}"] = comparison.get(suffix)
|
||||
writer.writerow(flat)
|
||||
return output.getvalue()
|
||||
|
||||
|
||||
async def _legacy_values(
|
||||
db: AsyncSession, ticker_ids: list[int]
|
||||
) -> dict[int, FundamentalData]:
|
||||
if not ticker_ids:
|
||||
return {}
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(FundamentalData).where(FundamentalData.ticker_id.in_(ticker_ids))
|
||||
)
|
||||
).scalars()
|
||||
return {row.ticker_id: row for row in rows}
|
||||
|
||||
|
||||
async def _source_runs(db: AsyncSession) -> dict[str, dict[str, Any] | None]:
|
||||
sources = ("sec_facts", "dolt_earnings")
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(DataImportRun)
|
||||
.where(
|
||||
DataImportRun.source.in_(sources),
|
||||
DataImportRun.status.in_(("promoted", "no_op")),
|
||||
)
|
||||
.order_by(DataImportRun.id.desc())
|
||||
)
|
||||
).scalars()
|
||||
latest: dict[str, dict[str, Any] | None] = {source: None for source in sources}
|
||||
for row in rows:
|
||||
if latest[row.source] is None:
|
||||
latest[row.source] = {
|
||||
"run_id": row.id,
|
||||
"status": row.status,
|
||||
"revision": row.revision,
|
||||
"source_max_date": _iso(row.source_max_date),
|
||||
"completed_at": _iso(row.completed_at),
|
||||
}
|
||||
return latest
|
||||
|
||||
|
||||
def _field_comparison(
|
||||
key: str, legacy: float | None, candidate: float | None
|
||||
) -> dict[str, Any]:
|
||||
legacy = float(legacy) if _finite(legacy) else None
|
||||
candidate = float(candidate) if _finite(candidate) else None
|
||||
absolute = _delta(legacy, candidate)
|
||||
relative = (
|
||||
None
|
||||
if absolute is None or legacy in (None, 0)
|
||||
else round(absolute / abs(legacy) * 100.0, 4)
|
||||
)
|
||||
tolerance = FIELD_TOLERANCES[key]
|
||||
material = False
|
||||
if absolute is not None:
|
||||
material = abs(absolute) > tolerance["absolute"]
|
||||
relative_limit = tolerance["relative_pct"]
|
||||
if relative_limit is not None:
|
||||
material = material and relative is not None and abs(relative) > relative_limit
|
||||
return {
|
||||
"legacy": _round(legacy),
|
||||
"candidate": _round(candidate),
|
||||
"absolute_delta": absolute,
|
||||
"relative_delta_pct": relative,
|
||||
"material": material,
|
||||
"definition_changed": True,
|
||||
}
|
||||
|
||||
|
||||
def _attach_ranks(rows: list[dict[str, Any]], value_key: str, rank_key: str) -> None:
|
||||
values = [
|
||||
row["scores"][value_key]
|
||||
for row in rows
|
||||
if _finite(row["scores"][value_key])
|
||||
]
|
||||
for row in rows:
|
||||
value = row["scores"][value_key]
|
||||
row["scores"][rank_key] = (
|
||||
1 + sum(other > value for other in values) if _finite(value) else None
|
||||
)
|
||||
|
||||
|
||||
def _summary(rows: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
field_stats = {}
|
||||
for key in FIELD_KEYS:
|
||||
comparisons = [row["fields"][key] for row in rows]
|
||||
deltas = [
|
||||
abs(item["absolute_delta"])
|
||||
for item in comparisons
|
||||
if item["absolute_delta"] is not None
|
||||
]
|
||||
field_stats[key] = {
|
||||
"legacy_available": sum(item["legacy"] is not None for item in comparisons),
|
||||
"candidate_available": sum(
|
||||
item["candidate"] is not None for item in comparisons
|
||||
),
|
||||
"both_available": len(deltas),
|
||||
"material_differences": sum(item["material"] for item in comparisons),
|
||||
"median_absolute_delta": _round(statistics.median(deltas) if deltas else None),
|
||||
"p95_absolute_delta": _round(_percentile(deltas, 0.95)),
|
||||
"max_absolute_delta": _round(max(deltas) if deltas else None),
|
||||
}
|
||||
|
||||
fundamental_deltas = _score_deltas(rows, "fundamental_delta")
|
||||
changed_rows = sorted(
|
||||
(
|
||||
{
|
||||
"symbol": row["symbol"],
|
||||
"fundamental_delta": row["scores"]["fundamental_delta"],
|
||||
"fundamental_rank_change": row["scores"]["fundamental_rank_change"],
|
||||
}
|
||||
for row in rows
|
||||
if row["scores"]["fundamental_delta"] is not None
|
||||
),
|
||||
key=lambda item: (
|
||||
abs(item["fundamental_delta"] or 0),
|
||||
),
|
||||
reverse=True,
|
||||
)[:20]
|
||||
return {
|
||||
"universe_count": len(rows),
|
||||
"legacy_fundamental_score_available": _count_score(
|
||||
rows, "legacy_fundamental"
|
||||
),
|
||||
"candidate_fundamental_score_available": _count_score(
|
||||
rows, "candidate_fundamental"
|
||||
),
|
||||
"fundamental_scores_compared": len(fundamental_deltas),
|
||||
"fundamental_score_material_changes": sum(
|
||||
abs(delta) > 5.0 for delta in fundamental_deltas
|
||||
),
|
||||
"fundamental_rank_changes": _rank_change_count(
|
||||
rows, "fundamental_rank_change"
|
||||
),
|
||||
"field_stats": field_stats,
|
||||
"largest_changes": changed_rows,
|
||||
}
|
||||
|
||||
|
||||
def _score_deltas(rows: Iterable[dict[str, Any]], key: str) -> list[float]:
|
||||
return [
|
||||
row["scores"][key]
|
||||
for row in rows
|
||||
if row["scores"][key] is not None
|
||||
]
|
||||
|
||||
|
||||
def _count_score(rows: Iterable[dict[str, Any]], key: str) -> int:
|
||||
return sum(row["scores"][key] is not None for row in rows)
|
||||
|
||||
|
||||
def _rank_change_count(rows: Iterable[dict[str, Any]], key: str) -> int:
|
||||
return sum(
|
||||
row["scores"][key] not in (None, 0)
|
||||
for row in rows
|
||||
)
|
||||
|
||||
|
||||
def _rank_change(legacy: int | None, candidate: int | None) -> int | None:
|
||||
# Positive means the candidate improved its rank.
|
||||
return legacy - candidate if legacy is not None and candidate is not None else None
|
||||
|
||||
|
||||
def _delta(legacy: float | None, candidate: float | None) -> float | None:
|
||||
if not _finite(legacy) or not _finite(candidate):
|
||||
return None
|
||||
return round(candidate - legacy, 4)
|
||||
|
||||
|
||||
def _round(value: float | None, digits: int = 4) -> float | None:
|
||||
return round(float(value), digits) if _finite(value) else None
|
||||
|
||||
|
||||
def _percentile(values: list[float], quantile: float) -> float | None:
|
||||
if not values:
|
||||
return None
|
||||
ordered = sorted(values)
|
||||
index = max(0, math.ceil(quantile * len(ordered)) - 1)
|
||||
return ordered[index]
|
||||
|
||||
|
||||
def _finite(value: Any) -> bool:
|
||||
return (
|
||||
isinstance(value, (int, float))
|
||||
and not isinstance(value, bool)
|
||||
and math.isfinite(value)
|
||||
)
|
||||
|
||||
|
||||
def _iso(value: Any) -> str | None:
|
||||
return value.isoformat() if value is not None else None
|
||||
|
||||
|
||||
def _artifact_stamp(raw: str) -> str:
|
||||
parsed = datetime.fromisoformat(raw.replace("Z", "+00:00"))
|
||||
return parsed.astimezone(timezone.utc).strftime("%Y%m%dT%H%M%S%fZ")
|
||||
|
||||
|
||||
def _atomic_write(path: Path, content: str) -> None:
|
||||
temp = path.with_name(f".{path.name}.{os.getpid()}.tmp")
|
||||
temp.write_text(content, encoding="utf-8", newline="")
|
||||
os.replace(temp, path)
|
||||
|
||||
|
||||
def _load_manifest(report_dir: str | Path) -> dict[str, Any] | None:
|
||||
path = Path(report_dir).expanduser().resolve() / "latest.json"
|
||||
try:
|
||||
loaded = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError, TypeError, ValueError):
|
||||
return None
|
||||
return loaded if isinstance(loaded, dict) else None
|
||||
|
||||
|
||||
def _manifest_artifact(
|
||||
report_dir: str | Path, manifest: dict[str, Any], key: str
|
||||
) -> Path:
|
||||
directory = Path(report_dir).expanduser().resolve()
|
||||
name = Path(str(manifest.get(key, ""))).name
|
||||
if not name:
|
||||
raise ValueError(f"Latest parity manifest has no {key}")
|
||||
return directory / name
|
||||
@@ -0,0 +1,164 @@
|
||||
"""Actionability gate for incomplete SEC fundamentals."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass
|
||||
|
||||
from sqlalchemy import exists, func, select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.models.data_import_run import DataImportRun
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.sec_filing_gap import SecFilingGap
|
||||
from app.models.ticker import Ticker
|
||||
from app.services import fundamental_data_refresh_service
|
||||
|
||||
_SEC_FORMS = ("10-K", "10-Q", "10-K/A", "10-Q/A")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SetupQuality:
|
||||
eligible: bool
|
||||
code: str | None = None
|
||||
message: str | None = None
|
||||
|
||||
|
||||
async def active_gaps(
|
||||
db: AsyncSession,
|
||||
ciks: set[str] | None = None,
|
||||
) -> list[SecFilingGap]:
|
||||
"""Unresolved gaps that have not been superseded by a later filing."""
|
||||
matching_snapshot = exists().where(
|
||||
FundamentalSnapshot.accession == SecFilingGap.accession
|
||||
)
|
||||
gap_date = func.coalesce(
|
||||
SecFilingGap.index_date,
|
||||
func.date(SecFilingGap.first_seen_at),
|
||||
)
|
||||
later_snapshot = exists().where(
|
||||
FundamentalSnapshot.cik == SecFilingGap.cik,
|
||||
FundamentalSnapshot.form.in_(_SEC_FORMS),
|
||||
FundamentalSnapshot.filed_date > gap_date,
|
||||
)
|
||||
stmt = select(SecFilingGap).where(
|
||||
~matching_snapshot,
|
||||
~later_snapshot,
|
||||
)
|
||||
if ciks is not None:
|
||||
if not ciks:
|
||||
return []
|
||||
stmt = stmt.where(SecFilingGap.cik.in_(ciks))
|
||||
return list((await db.execute(stmt)).scalars().all())
|
||||
|
||||
|
||||
async def _latest_validation(db: AsyncSession) -> dict:
|
||||
payload = (
|
||||
await db.execute(
|
||||
select(DataImportRun.validation_json)
|
||||
.where(
|
||||
DataImportRun.source == "sec_facts",
|
||||
DataImportRun.validation_json.is_not(None),
|
||||
)
|
||||
.order_by(DataImportRun.id.desc())
|
||||
.limit(1)
|
||||
)
|
||||
).scalar_one_or_none()
|
||||
if not payload:
|
||||
return {}
|
||||
try:
|
||||
summary = json.loads(payload)
|
||||
except (TypeError, ValueError):
|
||||
return {}
|
||||
return summary if isinstance(summary, dict) else {}
|
||||
|
||||
|
||||
async def blocked_reasons_by_cik(
|
||||
db: AsyncSession,
|
||||
ciks: set[str] | None = None,
|
||||
) -> dict[str, str]:
|
||||
"""Current SEC blocker code by CIK; no historical audit scan."""
|
||||
if not await fundamental_data_refresh_service.is_enabled(db):
|
||||
return {}
|
||||
if ciks is not None and not ciks:
|
||||
return {}
|
||||
|
||||
reasons = {
|
||||
gap.cik: "sec_filing_gap" for gap in await active_gaps(db, ciks)
|
||||
}
|
||||
summary = await _latest_validation(db)
|
||||
|
||||
def wanted(cik: str) -> bool:
|
||||
return ciks is None or cik in ciks
|
||||
|
||||
# New summaries carry the complete compact CIK set while the detailed lists
|
||||
# stay capped for audit readability. Detailed entries supply the reason.
|
||||
for cik in summary.get("setup_blocked_ciks") or []:
|
||||
normalized = str(cik) if cik else ""
|
||||
if normalized and wanted(normalized):
|
||||
reasons.setdefault(normalized, "sec_filing_gap")
|
||||
for item in summary.get("missing_xbrl") or []:
|
||||
normalized = str(item.get("cik") or "")
|
||||
if normalized and wanted(normalized):
|
||||
reasons.setdefault(normalized, "sec_filing_gap")
|
||||
for cik in summary.get("no_xbrl_ciks") or []:
|
||||
normalized = str(cik) if cik else ""
|
||||
if normalized and wanted(normalized):
|
||||
reasons[normalized] = "no_xbrl_filings"
|
||||
for item in summary.get("no_xbrl_filings") or []:
|
||||
normalized = str(item.get("cik") or "")
|
||||
if normalized and wanted(normalized):
|
||||
reasons[normalized] = "no_xbrl_filings"
|
||||
return reasons
|
||||
|
||||
|
||||
async def blocked_ciks(db: AsyncSession) -> set[str]:
|
||||
return set(await blocked_reasons_by_cik(db))
|
||||
|
||||
|
||||
async def blocked_ticker_ids(db: AsyncSession) -> set[int]:
|
||||
ciks = await blocked_ciks(db)
|
||||
if not ciks:
|
||||
return set()
|
||||
rows = await db.execute(select(Ticker.id).where(Ticker.cik.in_(ciks)))
|
||||
return {int(ticker_id) for ticker_id in rows.scalars()}
|
||||
|
||||
|
||||
async def ticker_quality(db: AsyncSession, symbol: str) -> SetupQuality:
|
||||
ticker = (
|
||||
await db.execute(
|
||||
select(Ticker).where(Ticker.symbol == symbol.strip().upper())
|
||||
)
|
||||
).scalar_one_or_none()
|
||||
if ticker is None or not ticker.cik:
|
||||
return SetupQuality(eligible=True)
|
||||
reason = (await blocked_reasons_by_cik(db, {ticker.cik})).get(ticker.cik)
|
||||
if reason == "no_xbrl_filings":
|
||||
return SetupQuality(
|
||||
eligible=False,
|
||||
code=reason,
|
||||
message=(
|
||||
"No SEC 10-K/10-Q is available for this registrant, so new setups "
|
||||
"are paused. New registrants clear automatically after their first "
|
||||
"filing; a successor shell needs an SEC CIK override."
|
||||
),
|
||||
)
|
||||
if reason:
|
||||
return SetupQuality(
|
||||
eligible=False,
|
||||
code=reason,
|
||||
message=(
|
||||
"A recent SEC filing is still being reconciled, so new setups are "
|
||||
"paused. The scheduled fundamentals import retries it automatically."
|
||||
),
|
||||
)
|
||||
return SetupQuality(eligible=True)
|
||||
|
||||
|
||||
async def ticker_is_eligible(db: AsyncSession, ticker_id: int) -> bool:
|
||||
cik = (
|
||||
await db.execute(select(Ticker.cik).where(Ticker.id == ticker_id))
|
||||
).scalar_one_or_none()
|
||||
if not cik:
|
||||
return True
|
||||
return cik not in await blocked_reasons_by_cik(db, {cik})
|
||||
@@ -100,6 +100,8 @@ async def fetch_and_ingest(
|
||||
symbol: str,
|
||||
start_date: date | None = None,
|
||||
end_date: date | None = None,
|
||||
*,
|
||||
refresh_sr: bool = True,
|
||||
) -> IngestionResult:
|
||||
"""Fetch OHLCV data from provider and upsert into Price Store.
|
||||
|
||||
@@ -129,7 +131,12 @@ async def fetch_and_ingest(
|
||||
if bar_count < minimum_backfill_bars:
|
||||
start_date = backfill_start
|
||||
elif progress is not None:
|
||||
start_date = progress.last_ingested_date + timedelta(days=1)
|
||||
# Re-fetch the latest stored session so an in-progress daily bar can
|
||||
# be overwritten as the market moves. Starting one day later makes
|
||||
# every subsequent intraday, near-close, and manual refresh skip
|
||||
# today's bar once the first partial snapshot has been stored.
|
||||
# The price-store upsert keeps this one-session overlap idempotent.
|
||||
start_date = progress.last_ingested_date
|
||||
else:
|
||||
start_date = backfill_start
|
||||
|
||||
@@ -239,7 +246,7 @@ async def fetch_and_ingest(
|
||||
ticker.symbol,
|
||||
ingested_count,
|
||||
)
|
||||
if ingested_count > 0:
|
||||
if ingested_count > 0 and refresh_sr:
|
||||
await _refresh_structural_sr(db, ticker.symbol)
|
||||
return IngestionResult(
|
||||
symbol=ticker.symbol,
|
||||
@@ -249,9 +256,28 @@ async def fetch_and_ingest(
|
||||
message=f"Rate limited. Ingested {ingested_count} records. Resume available.",
|
||||
)
|
||||
|
||||
if ingested_count > 0:
|
||||
if ingested_count > 0 and refresh_sr:
|
||||
await _refresh_structural_sr(db, ticker.symbol)
|
||||
|
||||
# Incremental fetches deliberately overlap the latest stored session so an
|
||||
# in-progress bar can be updated. A halted/delisted symbol can therefore
|
||||
# return one old bar forever; non-empty no longer means fresh. Judge stale
|
||||
# state from the newest stored session after the upserts instead.
|
||||
latest = await _get_latest_ohlcv_date(db, ticker.id)
|
||||
gap_days = (end_date - latest).days if latest is not None else None
|
||||
if gap_days is not None and gap_days > _STALE_OHLCV_GAP_DAYS:
|
||||
return IngestionResult(
|
||||
symbol=ticker.symbol,
|
||||
records_ingested=ingested_count,
|
||||
last_date=latest,
|
||||
status="stale",
|
||||
message=(
|
||||
f"No new bars since {latest.isoformat()} ({gap_days}d gap). "
|
||||
"The symbol may be halted, delisted, or renamed under a new ticker — "
|
||||
"check the listing and add/fetch the current symbol if it changed."
|
||||
),
|
||||
)
|
||||
|
||||
return IngestionResult(
|
||||
symbol=ticker.symbol,
|
||||
records_ingested=ingested_count,
|
||||
|
||||
@@ -352,6 +352,7 @@ def _to_dict(
|
||||
current_price: float | None,
|
||||
benchmark_closes: dict[date, float] | None = None,
|
||||
trailing: tuple[float, float | None] | None = None,
|
||||
holding_sessions: tuple[int, int] | None = None,
|
||||
) -> dict:
|
||||
# For open trades, mark to market; for closed, the realized exit price.
|
||||
ref = current_price if trade.status == "open" else trade.close_price
|
||||
@@ -395,6 +396,8 @@ def _to_dict(
|
||||
"fill_mode": trade.fill_mode,
|
||||
"trailing_stop": trailing[0] if trailing else None,
|
||||
"trailing_distance_pct": trailing[1] if trailing else None,
|
||||
"sessions_held": holding_sessions[0] if holding_sessions else None,
|
||||
"sessions_remaining": holding_sessions[1] if holding_sessions else None,
|
||||
}
|
||||
|
||||
|
||||
@@ -435,6 +438,35 @@ async def list_trades(
|
||||
# Current trailing-stop level + distance for open trades (when a trailing
|
||||
# policy is active).
|
||||
policy = await get_exit_policy(db)
|
||||
holding_sessions: dict[int, tuple[int, int]] = {}
|
||||
if policy["mode"] in ("time", "atr_trailing"):
|
||||
hold_days = int(policy["hold_days"])
|
||||
open_trades = [trade for trade, _ in rows if trade.status == "open"]
|
||||
if open_trades:
|
||||
ticker_ids = {trade.ticker_id for trade in open_trades}
|
||||
earliest_opened = min(trade.opened_at.date() for trade in open_trades)
|
||||
session_rows = (
|
||||
await db.execute(
|
||||
select(OHLCVRecord.ticker_id, OHLCVRecord.date)
|
||||
.where(
|
||||
OHLCVRecord.ticker_id.in_(ticker_ids),
|
||||
OHLCVRecord.date > earliest_opened,
|
||||
)
|
||||
.order_by(OHLCVRecord.ticker_id, OHLCVRecord.date)
|
||||
)
|
||||
).all()
|
||||
dates_by_ticker: dict[int, list[date]] = {}
|
||||
for ticker_id, session_date in session_rows:
|
||||
dates_by_ticker.setdefault(int(ticker_id), []).append(session_date)
|
||||
for trade in open_trades:
|
||||
dates = dates_by_ticker.get(trade.ticker_id, [])
|
||||
held = len(dates) - bisect.bisect_right(
|
||||
dates, trade.opened_at.date()
|
||||
)
|
||||
# Do not clamp: a policy shortened below the current holding
|
||||
# period must remain visible as overdue until the exit pass runs.
|
||||
holding_sessions[trade.id] = (held, hold_days - held)
|
||||
|
||||
trailing_info: dict[int, tuple[float, float | None]] = {}
|
||||
if policy["mode"] == "trailing":
|
||||
trail_frac = policy["trailing_pct"] / 100.0
|
||||
@@ -483,7 +515,14 @@ async def list_trades(
|
||||
trailing_info[t.id] = (level, dist)
|
||||
|
||||
return [
|
||||
_to_dict(t, sym, prices.get(t.ticker_id), benchmark_closes, trailing_info.get(t.id))
|
||||
_to_dict(
|
||||
t,
|
||||
sym,
|
||||
prices.get(t.ticker_id),
|
||||
benchmark_closes,
|
||||
trailing_info.get(t.id),
|
||||
holding_sessions.get(t.id),
|
||||
)
|
||||
for t, sym in rows
|
||||
]
|
||||
|
||||
|
||||
@@ -1,16 +1,23 @@
|
||||
"""AI/Tech Regime Monitor v2.
|
||||
"""AI/Tech Regime Monitor v3.
|
||||
|
||||
The monitor is a risk thermometer, not a probability or trading rule. It keeps
|
||||
two deliberately separate outputs:
|
||||
|
||||
* State: current structural stress (price, breadth, credit, volatility).
|
||||
* Warning: deterioration/divergence that may precede State (breadth, relative
|
||||
strength, and sourced fundamental observations).
|
||||
* Warning: deterioration/divergence that may precede State (breadth divergence,
|
||||
relative strength, credit impulse).
|
||||
|
||||
Daily snapshots are the point-in-time record. The first v2 run rewrites the
|
||||
latest ``REBUILD_SESSIONS`` trading sessions once; ordinary runs thereafter only
|
||||
upsert the latest trading date. Fundamental observations are never replayed
|
||||
before their effective date.
|
||||
Both scores are quantitative and daily. The sourced hyperscaler capex and
|
||||
earnings-reaction observations are a qualitative *overlay* in v3 rather than
|
||||
weighted sensors: at a combined 20 points they could not reach the event
|
||||
study's alarm threshold even when both pegged, so refreshing them appeared to
|
||||
do nothing. They are reported next to the scores instead of inside them.
|
||||
|
||||
Daily snapshots are the point-in-time record. The first run under a new
|
||||
``METHODOLOGY`` rewrites the latest ``REBUILD_SESSIONS`` trading sessions once;
|
||||
ordinary runs thereafter only upsert the latest trading date. The overlay is
|
||||
still gated by its effective date so a rebuild cannot stamp today's observation
|
||||
onto historical snapshots.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -41,19 +48,51 @@ _CA_BUNDLE = os.environ.get("SSL_CERT_FILE", "")
|
||||
KEY_CONFIG = "regime_monitor_config"
|
||||
KEY_FUNDAMENTALS = "regime_fundamental_overrides"
|
||||
|
||||
METHODOLOGY = "v2"
|
||||
METHODOLOGY = "v3"
|
||||
# Snapshots are reseeded on a methodology bump, but fundamental observations are
|
||||
# collected by hand/LLM and carried across it when the format is compatible.
|
||||
CATEGORICAL_FUNDAMENTAL_METHODOLOGIES = frozenset({"v2", "v3"})
|
||||
REBUILD_SESSIONS = 400
|
||||
MIN_COVERAGE = 75.0
|
||||
SOURCE_MAX_LAG_DAYS = 7
|
||||
|
||||
QUADRANT_STATE_DIVIDER = 60.0
|
||||
QUADRANT_WARNING_DIVIDER = 60.0
|
||||
# Bands are per axis: the two scores have genuinely different realized ranges,
|
||||
# so one shared set made Warning's top bands unreachable (v2 Warning never
|
||||
# exceeded 64.9 in 408 sessions while State reached 91.2). Thresholds are round
|
||||
# numbers chosen so each band covers a sane share of history, not percentile
|
||||
# fits -- percentile-derived bands would drift on every rebuild and silently
|
||||
# rewrite what past snapshots meant. Realized shares over the 408 sessions to
|
||||
# 2026-07-24: State 73/15/8/3%, Warning 69/20/8/3%.
|
||||
STATE_BANDS = (20.0, 50.0, 80.0)
|
||||
WARNING_BANDS = (20.0, 40.0, 60.0)
|
||||
|
||||
QUADRANT_STATE_DIVIDER = 50.0
|
||||
QUADRANT_WARNING_DIVIDER = 40.0
|
||||
QUADRANT_MARGIN = 5.0
|
||||
|
||||
HY_OAS_MILD = 3.5
|
||||
HY_OAS_ELEVATED = 5.0
|
||||
HY_OAS_STRESSED = 7.0
|
||||
HY_OAS_REFERENCE_YEARS = 10.0
|
||||
# ICE restricted FRED to a rolling 3-year window for BAMLH0A0HYM2 in April 2026
|
||||
# ("Starting in April 2026, this series will only include 3 years of
|
||||
# observations"), so v2's 10-year reference window silently became 3. Against 3
|
||||
# years of uniformly tight spreads (2.59-4.61 over the calibration window) the
|
||||
# blended upper-tail percentile saturated at an OAS of ~4.5 and scored 20 points
|
||||
# of stress at 3.5 -- the level these anchors call "mild". The anchors already
|
||||
# encode the long-run distribution, so the credit *level* is now purely anchored
|
||||
# and credit *dynamics* live in W3 on the Warning axis where they belong.
|
||||
HY_OAS_WINDOW_DAYS = 400 # only W3's lookback plus slack is needed now
|
||||
W3_OAS_LOOKBACK = 20
|
||||
W3_OAS_FULL_SCALE_PCT = 35.0
|
||||
|
||||
# Drawdown anchors (drawdown %, stress score). v2 used a bare ``dd_pct * 5``,
|
||||
# which pegged at a 20% drawdown -- the 90th percentile of the observed
|
||||
# distribution -- so 39 of 408 sessions sat at exactly 100 with no resolution
|
||||
# left during the part of a selloff that matters most. These anchors keep
|
||||
# headroom past the observed 36% maximum.
|
||||
P3_DRAWDOWN_ANCHORS = (
|
||||
(0.0, 0.0), (4.0, 10.0), (8.0, 25.0), (16.0, 50.0), (28.0, 78.0), (40.0, 100.0),
|
||||
)
|
||||
|
||||
STATE_WEIGHTS = {
|
||||
"price": 40.0,
|
||||
@@ -61,11 +100,15 @@ STATE_WEIGHTS = {
|
||||
"credit": 20.0,
|
||||
"volatility": 15.0,
|
||||
}
|
||||
# Fundamentals left the score in v3. At 12 + 8 points they could not reach the
|
||||
# event study's alarm threshold even when both pegged at 100, so the LLM read was
|
||||
# decorative; it is now a separate qualitative overlay. Credit *impulse* takes
|
||||
# their place because the OAS level is pinned at zero below the 3.5 anchor while
|
||||
# its rate of change is not.
|
||||
WARNING_WEIGHTS = {
|
||||
"breadth_divergence": 50.0,
|
||||
"breadth_divergence": 45.0,
|
||||
"relative_strength": 30.0,
|
||||
"capex": 12.0,
|
||||
"earnings_reaction": 8.0,
|
||||
"credit_impulse": 25.0,
|
||||
}
|
||||
|
||||
# Fixed at the v2 launch. These are liquid S&P 500/Nasdaq AI, semiconductor,
|
||||
@@ -92,7 +135,10 @@ DEFAULT_CONFIG: dict = {
|
||||
|
||||
CAPEX_STATES = ("raising", "holding", "cutting", "unknown")
|
||||
GNSD_STATES = ("yes", "no", "mixed")
|
||||
_CAPEX_STATE_SCORES = {"raising": 0.0, "holding": 0.0, "cutting": 100.0}
|
||||
# v2 scored raising and holding identically at 0, so in a capex boom the reading
|
||||
# was pinned at 0 and could not express the raising -> holding deceleration that
|
||||
# is the actual early warning. Display-only in v3, but it should still describe.
|
||||
_CAPEX_STATE_SCORES = {"raising": 0.0, "holding": 50.0, "cutting": 100.0}
|
||||
_GNSD_SCORES = {"yes": 100.0, "no": 0.0}
|
||||
|
||||
Series = list[tuple[date, float]]
|
||||
@@ -127,12 +173,23 @@ def _blend(leader: float | None, confirm: float | None, leader_weight: float = 2
|
||||
return sum(v * w for v, w in parts) / sum(w for _, w in parts)
|
||||
|
||||
|
||||
def band_for(score: float) -> str:
|
||||
if score < 30:
|
||||
def _interpolate(x: float, anchors: tuple[tuple[float, float], ...]) -> float:
|
||||
"""Piecewise-linear lookup, flat outside the first and last anchor."""
|
||||
if x <= anchors[0][0]:
|
||||
return anchors[0][1]
|
||||
for (x0, y0), (x1, y1) in zip(anchors, anchors[1:]):
|
||||
if x <= x1:
|
||||
return y0 + (y1 - y0) * (x - x0) / (x1 - x0)
|
||||
return anchors[-1][1]
|
||||
|
||||
|
||||
def band_for(score: float, bands: tuple[float, float, float] = STATE_BANDS) -> str:
|
||||
watch, elevated, breaking = bands
|
||||
if score < watch:
|
||||
return "stable"
|
||||
if score < 60:
|
||||
if score < elevated:
|
||||
return "watch"
|
||||
if score < 80:
|
||||
if score < breaking:
|
||||
return "elevated"
|
||||
return "breaking"
|
||||
|
||||
@@ -167,19 +224,29 @@ def p2_death_cross(smh: list[float], qqq: list[float], leader_weight: float = 2.
|
||||
return _blend(_death_cross(smh), _death_cross(qqq), leader_weight)
|
||||
|
||||
|
||||
def _drawdown(closes: list[float]) -> float | None:
|
||||
def drawdown_pct(closes: list[float]) -> float | None:
|
||||
"""Percentage below the trailing 52-week closing high."""
|
||||
if len(closes) < 30:
|
||||
return None
|
||||
peak = max(closes[-252:])
|
||||
if peak <= 0:
|
||||
return None
|
||||
dd_pct = (peak - closes[-1]) / peak * 100.0
|
||||
return _clamp(dd_pct * 5.0)
|
||||
return (peak - closes[-1]) / peak * 100.0
|
||||
|
||||
|
||||
def p3_drawdown(smh: list[float], qqq: list[float]) -> float | None:
|
||||
vals = [v for v in (_drawdown(smh), _drawdown(qqq)) if v is not None]
|
||||
return max(vals) if vals else None
|
||||
def _drawdown(closes: list[float]) -> float | None:
|
||||
dd_pct = drawdown_pct(closes)
|
||||
return None if dd_pct is None else _clamp(_interpolate(dd_pct, P3_DRAWDOWN_ANCHORS))
|
||||
|
||||
|
||||
def p3_drawdown(smh: list[float], qqq: list[float], leader_weight: float = 2.0) -> float | None:
|
||||
"""Anchored drawdown stress on the same 2:1 leader/confirm blend P1 and P2 use.
|
||||
|
||||
v2 took ``max()`` here, which meant the more volatile leader always won and
|
||||
the price pillar reduced to this one sensor: its realized share of State was
|
||||
65% against a nominal 40% weight. Blending brings that back to 40%.
|
||||
"""
|
||||
return _blend(_drawdown(smh), _drawdown(qqq), leader_weight)
|
||||
|
||||
|
||||
def p4_relative_strength(smh: list[float], spy: list[float], lookback: int = 60) -> float | None:
|
||||
@@ -220,18 +287,68 @@ def _oas_absolute_score(value: float) -> float:
|
||||
|
||||
|
||||
def f2_credit_spreads(oas_values: list[float]) -> float | None:
|
||||
"""HY OAS stress: 70% named absolute anchors + 30% upper-tail percentile."""
|
||||
"""HY OAS level against named absolute anchors (3.5 mild / 5.0 / 7.0).
|
||||
|
||||
v2 blended 70% of this with a 30% upper-tail percentile over the available
|
||||
history. That leg was always a second, noisier estimate of what the anchors
|
||||
already encode -- and once the usable window shrank to 3 uniformly tight
|
||||
years it saturated far below any real stress level. Removed rather than
|
||||
repaired: see ``HY_OAS_WINDOW_DAYS``.
|
||||
"""
|
||||
if not oas_values:
|
||||
return None
|
||||
latest = oas_values[-1]
|
||||
absolute = _oas_absolute_score(latest)
|
||||
if len(oas_values) < 30:
|
||||
return round(absolute, 2)
|
||||
less = sum(1 for v in oas_values if v < latest)
|
||||
equal = sum(1 for v in oas_values if v == latest)
|
||||
percentile = (less + 0.5 * equal) / len(oas_values) * 100.0
|
||||
relative = _clamp((percentile - 50.0) / 45.0 * 100.0)
|
||||
return round(absolute * 0.7 + relative * 0.3, 2)
|
||||
return round(_oas_absolute_score(oas_values[-1]), 2)
|
||||
|
||||
|
||||
def w3_credit_impulse(
|
||||
oas_values: list[float], lookback: int = W3_OAS_LOOKBACK
|
||||
) -> float | None:
|
||||
"""HY OAS rate of change: widening only, relative so it works at any level.
|
||||
|
||||
The credit *level* (C1) sits at zero for as long as spreads stay under the
|
||||
3.5 mild anchor -- 2.77 as of the v3 cutover -- so it contributes nothing to
|
||||
State in a calm tape. The rate of change still does, and spread widening is
|
||||
a classic lead, which is what Warning is for. Relative rather than absolute
|
||||
because +0.5pp means something very different at 2.7 than at 8.0.
|
||||
"""
|
||||
if len(oas_values) < lookback + 1:
|
||||
return None
|
||||
past = oas_values[-lookback - 1]
|
||||
if past <= 0:
|
||||
return None
|
||||
change_pct = (oas_values[-1] / past - 1.0) * 100.0
|
||||
return _clamp(change_pct / W3_OAS_FULL_SCALE_PCT * 100.0)
|
||||
|
||||
|
||||
def warning_sensor_scores(
|
||||
divergence: float | None,
|
||||
leader_closes: list[float],
|
||||
market_closes: list[float],
|
||||
oas_window: list[float],
|
||||
) -> dict[str, float | None]:
|
||||
"""The three Warning sensors, by pillar id.
|
||||
|
||||
Single definition so the live monitor and the event study cannot drift apart
|
||||
-- in v2 the study re-derived the score from ``WARNING_WEIGHTS`` by hand and
|
||||
would have silently kept measuring the old construct through this change.
|
||||
"""
|
||||
return {
|
||||
"breadth_divergence": divergence,
|
||||
"relative_strength": p4_relative_strength(leader_closes, market_closes),
|
||||
"credit_impulse": w3_credit_impulse(oas_window),
|
||||
}
|
||||
|
||||
|
||||
def score_warning_sensors(sensors: dict[str, float | None]) -> float | None:
|
||||
"""Weighted Warning score, renormalised over the sensors that are available."""
|
||||
live = [
|
||||
(float(score), float(WARNING_WEIGHTS[key]))
|
||||
for key, score in sensors.items()
|
||||
if score is not None and key in WARNING_WEIGHTS
|
||||
]
|
||||
if not live:
|
||||
return None
|
||||
return sum(s * w for s, w in live) / sum(w for _, w in live)
|
||||
|
||||
|
||||
def _sensor(sensor_id: str, label: str, score: float | None, **details: object) -> dict:
|
||||
@@ -244,7 +361,11 @@ def _sensor(sensor_id: str, label: str, score: float | None, **details: object)
|
||||
}
|
||||
|
||||
|
||||
def _score_pillars(pillars: list[dict], weights: dict[str, float]) -> dict:
|
||||
def _score_pillars(
|
||||
pillars: list[dict],
|
||||
weights: dict[str, float],
|
||||
bands: tuple[float, float, float] = STATE_BANDS,
|
||||
) -> dict:
|
||||
expected = sum(max(0.0, float(w)) for w in weights.values())
|
||||
available_weight = sum(
|
||||
max(0.0, float(weights.get(p["id"], 0.0)))
|
||||
@@ -276,7 +397,12 @@ def _score_pillars(pillars: list[dict], weights: dict[str, float]) -> dict:
|
||||
rounded = round(score, 1) if score is not None else None
|
||||
return {
|
||||
"score": rounded,
|
||||
"band": band_for(rounded) if rounded is not None and coverage >= MIN_COVERAGE else None,
|
||||
"band": (
|
||||
band_for(rounded, bands)
|
||||
if rounded is not None and coverage >= MIN_COVERAGE
|
||||
else None
|
||||
),
|
||||
"bands": {"watch": bands[0], "elevated": bands[1], "breaking": bands[2]},
|
||||
"coverage": round(coverage, 1),
|
||||
"minimum_coverage": MIN_COVERAGE,
|
||||
"available_pillars": [p["id"] for p in rows if p["available"]],
|
||||
@@ -309,13 +435,24 @@ def _value_asof(series: Series | None, as_of: date) -> float | None:
|
||||
return item[1] if item else None
|
||||
|
||||
|
||||
def _window_asof(series: Series | None, as_of: date, years: float) -> list[float]:
|
||||
def _window_asof(series: Series | None, as_of: date, days: int) -> list[float]:
|
||||
if not series:
|
||||
return []
|
||||
start = as_of - timedelta(days=int(365.25 * years))
|
||||
start = as_of - timedelta(days=days)
|
||||
return [v for d, v in series if start <= d <= as_of]
|
||||
|
||||
|
||||
def _coverage_days(series: Series | None, as_of: date) -> int | None:
|
||||
"""Span of history actually available at ``as_of``.
|
||||
|
||||
Recorded in every snapshot because the v2 credit percentile degraded from a
|
||||
10-year to a 3-year reference silently when the upstream licence changed --
|
||||
nothing asserted the window it claimed, so nothing noticed for months.
|
||||
"""
|
||||
dates = [d for d, _ in series or [] if d <= as_of]
|
||||
return (as_of - dates[0]).days if dates else None
|
||||
|
||||
|
||||
def _next_weekday(d: date) -> date:
|
||||
candidate = d + timedelta(days=1)
|
||||
while candidate.weekday() >= 5:
|
||||
@@ -340,19 +477,31 @@ def _fundamental_effective_date(overrides: dict) -> date | None:
|
||||
return _next_weekday(fetched) if fetched else None
|
||||
|
||||
|
||||
def _fundamental_scores_asof(overrides: dict, config: dict, as_of: date) -> tuple[float | None, float | None, dict]:
|
||||
def fundamental_overlay(overrides: dict, config: dict, as_of: date) -> dict:
|
||||
"""Point-in-time qualitative overlay. Never feeds State or Warning in v3.
|
||||
|
||||
The effective-date gate stays even though nothing is scored from this: the
|
||||
400-session rebuild replays historical dates, and stamping today's LLM read
|
||||
onto 2024 snapshots would be plain lookahead in the stored record.
|
||||
"""
|
||||
effective = _fundamental_effective_date(overrides)
|
||||
if effective is None or as_of < effective:
|
||||
return None, None, {"effective_date": effective.isoformat() if effective else None, "age_days": None}
|
||||
age = (as_of - effective).days
|
||||
stale = age > int(config.get("fundamental_staleness_days", 80))
|
||||
f1 = overrides.get("f1_score")
|
||||
f3 = overrides.get("f3_score")
|
||||
return (
|
||||
None if stale or f1 is None else _clamp(float(f1)),
|
||||
None if stale or f3 is None else _clamp(float(f3)),
|
||||
{"effective_date": effective.isoformat(), "age_days": age, "stale": stale},
|
||||
)
|
||||
pending = effective is None or as_of < effective
|
||||
age = None if pending else (as_of - effective).days
|
||||
stale = bool(age is not None and age > int(config.get("fundamental_staleness_days", 80)))
|
||||
return {
|
||||
"available": not pending and not stale,
|
||||
"pending": pending,
|
||||
"stale": stale,
|
||||
"effective_date": effective.isoformat() if effective else None,
|
||||
"age_days": age,
|
||||
"capex": None if pending else overrides.get("capex"),
|
||||
"good_news_stock_down": None if pending else overrides.get("good_news_stock_down"),
|
||||
"capex_stress": None if pending else overrides.get("f1_score"),
|
||||
"earnings_stress": None if pending else overrides.get("f3_score"),
|
||||
"reasoning": None if pending else overrides.get("reasoning"),
|
||||
"source": overrides.get("source"),
|
||||
"fetched_at": overrides.get("fetched_at"),
|
||||
}
|
||||
|
||||
|
||||
def _basket_hash(symbols: list[str]) -> str:
|
||||
@@ -393,12 +542,14 @@ def _compute_index(
|
||||
vix_item = _item_asof(vix_series, as_of)
|
||||
vix_score = p5_volatility(vix_item[1] if vix_item else None)
|
||||
oas_item = _item_asof(oas_series, as_of)
|
||||
oas_window = _window_asof(oas_series, as_of, HY_OAS_REFERENCE_YEARS)
|
||||
oas_window = _window_asof(oas_series, as_of, HY_OAS_WINDOW_DAYS)
|
||||
credit_score = f2_credit_spreads(oas_window)
|
||||
|
||||
divergence = _value_asof(divergence_series, as_of)
|
||||
relative_strength = p4_relative_strength(smh, spy)
|
||||
f1, f3, fundamental_meta = _fundamental_scores_asof(overrides, config, as_of)
|
||||
sensors = warning_sensor_scores(divergence, smh, spy, oas_window)
|
||||
relative_strength = sensors["relative_strength"]
|
||||
credit_impulse = sensors["credit_impulse"]
|
||||
overlay = fundamental_overlay(overrides, config, as_of)
|
||||
|
||||
state_pillars = [
|
||||
{
|
||||
@@ -444,21 +595,22 @@ def _compute_index(
|
||||
"sensors": [_sensor("W2", "60-session relative-strength deterioration", relative_strength)],
|
||||
},
|
||||
{
|
||||
"id": "capex",
|
||||
"label": "Hyperscaler capex revisions",
|
||||
"score": round(f1, 1) if f1 is not None else None,
|
||||
"sensors": [_sensor("F1", "Capex guidance cuts", f1)],
|
||||
},
|
||||
{
|
||||
"id": "earnings_reaction",
|
||||
"label": "Good news, stock down",
|
||||
"score": round(f3, 1) if f3 is not None else None,
|
||||
"sensors": [_sensor("F3", "Abnormal earnings reaction", f3)],
|
||||
"id": "credit_impulse",
|
||||
"label": "Credit impulse",
|
||||
"score": round(credit_impulse, 1) if credit_impulse is not None else None,
|
||||
"sensors": [
|
||||
_sensor(
|
||||
"W3",
|
||||
f"HY OAS {W3_OAS_LOOKBACK}-session widening",
|
||||
credit_impulse,
|
||||
oas=oas_item[1] if oas_item else None,
|
||||
)
|
||||
],
|
||||
},
|
||||
]
|
||||
|
||||
state = _score_pillars(state_pillars, STATE_WEIGHTS)
|
||||
warning = _score_pillars(warning_pillars, WARNING_WEIGHTS)
|
||||
state = _score_pillars(state_pillars, STATE_WEIGHTS, STATE_BANDS)
|
||||
warning = _score_pillars(warning_pillars, WARNING_WEIGHTS, WARNING_BANDS)
|
||||
|
||||
price_item = _item_asof(prices.get(tickers["leaders"][0]), as_of)
|
||||
dated_sources = {
|
||||
@@ -481,6 +633,7 @@ def _compute_index(
|
||||
"date": as_of.isoformat(),
|
||||
"state": state,
|
||||
"warning": warning,
|
||||
"fundamental_overlay": overlay,
|
||||
"quadrant_config": {
|
||||
"state_divider": QUADRANT_STATE_DIVIDER,
|
||||
"warning_divider": QUADRANT_WARNING_DIVIDER,
|
||||
@@ -502,14 +655,18 @@ def _compute_index(
|
||||
"breadth_pct_above_200": round(breadth_pct, 1) if breadth_pct is not None else None,
|
||||
"breadth_date": breadth_item[0].isoformat() if breadth_item else None,
|
||||
"fundamentals_fetched_at": overrides.get("fetched_at"),
|
||||
"fundamentals_effective_date": fundamental_meta.get("effective_date"),
|
||||
"fundamentals_age_days": fundamental_meta.get("age_days"),
|
||||
"fundamentals_effective_date": overlay.get("effective_date"),
|
||||
"fundamentals_age_days": overlay.get("age_days"),
|
||||
},
|
||||
"data_quality": {
|
||||
"minimum_coverage": MIN_COVERAGE,
|
||||
"oldest_market_input_age_days": max(source_ages.values()) if source_ages else None,
|
||||
"stale_inputs": stale_inputs,
|
||||
"inputs_fresh": not stale_inputs,
|
||||
# Upstream history spans, so a provider silently truncating a series
|
||||
# shows up in the record instead of quietly reshaping a sensor.
|
||||
"credit_history_days": _coverage_days(oas_series, as_of),
|
||||
"vix_history_days": _coverage_days(vix_series, as_of),
|
||||
},
|
||||
}
|
||||
|
||||
@@ -581,7 +738,11 @@ async def get_fundamental_overrides(db: AsyncSession) -> dict:
|
||||
stored = json.loads(raw)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
if stored.get("methodology") != METHODOLOGY:
|
||||
# The guard rejects pre-v2 blobs, where f1/f3 were arbitrary numbers with no
|
||||
# categorical source. v2 and v3 share the categorical format and both derive
|
||||
# f1/f3 from it below, so a methodology bump must not discard a live
|
||||
# observation -- only the capex *scale* changed, and that is recomputed.
|
||||
if stored.get("methodology") not in CATEGORICAL_FUNDAMENTAL_METHODOLOGIES:
|
||||
return default
|
||||
capex = _normalise_capex_states(stored.get("capex"), names)
|
||||
reaction = str(stored.get("good_news_stock_down", "mixed")).strip().lower()
|
||||
@@ -745,7 +906,7 @@ async def _upsert_snapshot(
|
||||
created_at=datetime.now(timezone.utc),
|
||||
))
|
||||
else:
|
||||
existing_v2 = _parse_v2(row.breakdown_json)
|
||||
existing_v2 = _parse_snapshot(row.breakdown_json)
|
||||
if existing_v2 is not None and not rewrite_existing_v2:
|
||||
return False, existing_v2
|
||||
row.total_score = float(state_score or 0.0)
|
||||
@@ -754,7 +915,7 @@ async def _upsert_snapshot(
|
||||
return True, result
|
||||
|
||||
|
||||
def _parse_v2(raw: str) -> dict | None:
|
||||
def _parse_snapshot(raw: str) -> dict | None:
|
||||
try:
|
||||
parsed = json.loads(raw)
|
||||
except (TypeError, ValueError):
|
||||
@@ -762,12 +923,12 @@ def _parse_v2(raw: str) -> dict | None:
|
||||
return parsed if parsed.get("methodology") == METHODOLOGY else None
|
||||
|
||||
|
||||
async def _latest_v2_row(db: AsyncSession) -> tuple[RegimeSnapshot, dict] | None:
|
||||
async def _latest_snapshot_row(db: AsyncSession) -> tuple[RegimeSnapshot, dict] | None:
|
||||
result = await db.execute(
|
||||
select(RegimeSnapshot).order_by(RegimeSnapshot.date.desc()).limit(1000)
|
||||
)
|
||||
for row in result.scalars().all():
|
||||
parsed = _parse_v2(row.breakdown_json)
|
||||
parsed = _parse_snapshot(row.breakdown_json)
|
||||
if parsed is not None:
|
||||
return row, parsed
|
||||
return None
|
||||
@@ -791,8 +952,10 @@ async def update_regime_monitor(db: AsyncSession, rebuild_sessions: int = REBUIL
|
||||
latest_date = leader_series[-1][0]
|
||||
|
||||
vix_series = await _fetch_fred_series("VIXCLS", end - timedelta(days=1200), end)
|
||||
# Asking for 13 years was misleading once the licence capped the series at 3;
|
||||
# the level needs the latest point and W3 needs its lookback, nothing more.
|
||||
oas_series = await _fetch_fred_series(
|
||||
"BAMLH0A0HYM2", end - timedelta(days=int(365.25 * 13)), end
|
||||
"BAMLH0A0HYM2", end - timedelta(days=HY_OAS_WINDOW_DAYS), end
|
||||
)
|
||||
|
||||
basket = config["breadth_basket"]
|
||||
@@ -805,7 +968,7 @@ async def update_regime_monitor(db: AsyncSession, rebuild_sessions: int = REBUIL
|
||||
logger.warning("Regime monitor: fixed-basket breadth skipped: %s", exc)
|
||||
breadth, breadth_counts, divergence = {}, {}, {}
|
||||
|
||||
latest_v2 = await _latest_v2_row(db)
|
||||
latest_v2 = await _latest_snapshot_row(db)
|
||||
rebuilding = latest_v2 is None and bool(leader_series)
|
||||
if rebuilding:
|
||||
dates = [d for d, _ in leader_series[-max(1, rebuild_sessions):]]
|
||||
@@ -860,7 +1023,7 @@ async def _result_at_or_before(
|
||||
.limit(1000)
|
||||
)
|
||||
for raw in result.scalars().all():
|
||||
parsed = _parse_v2(raw)
|
||||
parsed = _parse_snapshot(raw)
|
||||
parsed_hash = ((parsed or {}).get("basket") or {}).get("hash")
|
||||
if parsed is not None and (basket_hash is None or parsed_hash == basket_hash):
|
||||
return parsed
|
||||
@@ -877,7 +1040,7 @@ def _delta(current: dict, previous: dict | None) -> float | None:
|
||||
|
||||
|
||||
async def get_regime_monitor(db: AsyncSession) -> dict:
|
||||
latest = await _latest_v2_row(db)
|
||||
latest = await _latest_snapshot_row(db)
|
||||
if latest is None:
|
||||
return {"available": False, "reason": "v2 not computed yet"}
|
||||
row, result = latest
|
||||
@@ -902,6 +1065,15 @@ async def get_regime_monitor(db: AsyncSession) -> dict:
|
||||
quality["snapshot_age_days"] = snapshot_age
|
||||
quality["is_fresh"] = bool(quality.get("inputs_fresh")) and snapshot_age <= 4
|
||||
result["data_quality"] = quality
|
||||
|
||||
# The snapshot's overlay is the point-in-time record; the reader also wants
|
||||
# the current observation even when it is not effective until the next
|
||||
# session, because otherwise refreshing it looks like it did nothing.
|
||||
config = await get_regime_config(db)
|
||||
overrides = await get_fundamental_overrides(db)
|
||||
live = fundamental_overlay(overrides, config, date.today())
|
||||
live["observed_in_snapshot"] = bool((result.get("fundamental_overlay") or {}).get("available"))
|
||||
result["fundamental_context"] = live
|
||||
result["available"] = True
|
||||
return result
|
||||
|
||||
@@ -915,7 +1087,7 @@ async def get_regime_history(db: AsyncSession, days: int = 800) -> list[dict]:
|
||||
)
|
||||
out: list[dict] = []
|
||||
for row in result.scalars().all():
|
||||
data = _parse_v2(row.breakdown_json)
|
||||
data = _parse_snapshot(row.breakdown_json)
|
||||
if data is None:
|
||||
continue
|
||||
state, warning = data.get("state") or {}, data.get("warning") or {}
|
||||
|
||||
@@ -27,6 +27,7 @@ from app.models.signal_context_snapshot import SignalContextSnapshot
|
||||
from app.models.ticker import Ticker
|
||||
from app.models.trade_setup import TradeSetup
|
||||
from app.services.indicator_service import _extract_ohlcv, compute_atr
|
||||
from app.services import fundamentals_quality_service, system_event_service
|
||||
from app.services.price_service import query_ohlcv
|
||||
from app.services.qualification import setup_qualifies
|
||||
from app.services.sr_service import detect_gate_target_ladder
|
||||
@@ -526,6 +527,7 @@ async def scan_ticker(
|
||||
primary_min_rr: float | None = None,
|
||||
gate_levels_override: list[Any] | None = None,
|
||||
scan_run_id: str | None = None,
|
||||
fundamentals_eligible: bool | None = None,
|
||||
) -> list[TradeSetup]:
|
||||
"""Scan a single ticker for trade setups meeting the R:R threshold.
|
||||
|
||||
@@ -542,6 +544,17 @@ async def scan_ticker(
|
||||
"""
|
||||
ticker = await _get_ticker(db, symbol)
|
||||
|
||||
if fundamentals_eligible is None:
|
||||
fundamentals_eligible = await fundamentals_quality_service.ticker_is_eligible(
|
||||
db, ticker.id
|
||||
)
|
||||
if not fundamentals_eligible:
|
||||
logger.info(
|
||||
"Skipping %s: unresolved or unavailable SEC fundamentals",
|
||||
ticker.symbol,
|
||||
)
|
||||
return []
|
||||
|
||||
if primary_min_rr is None:
|
||||
primary_min_rr = PRIMARY_TARGET_MIN_RR
|
||||
|
||||
@@ -726,6 +739,29 @@ async def scan_all_tickers(
|
||||
ticker_rows = [(int(ticker_id), symbol) for ticker_id, symbol in result.all()]
|
||||
total = len(ticker_rows)
|
||||
|
||||
# Data-quality failures are not weak signals: they make a ticker ineligible.
|
||||
# Resolve once for the universe scan and pass the decision into scan_ticker.
|
||||
try:
|
||||
fundamentals_blocked_ids = (
|
||||
await fundamentals_quality_service.blocked_ticker_ids(db)
|
||||
)
|
||||
except Exception:
|
||||
await db.rollback()
|
||||
logger.exception(
|
||||
"Could not resolve fundamentals quality; blocking this scan closed"
|
||||
)
|
||||
await system_event_service.log_event_standalone(
|
||||
severity="error",
|
||||
source="rr_scanner",
|
||||
code="fundamentals_quality_unavailable",
|
||||
message=(
|
||||
"The fundamentals quality gate could not be evaluated; the "
|
||||
"universe scan was blocked to avoid issuing unchecked setups."
|
||||
),
|
||||
dedup_key="rr_scanner:fundamentals_quality_unavailable",
|
||||
)
|
||||
fundamentals_blocked_ids = {ticker_id for ticker_id, _ in ticker_rows}
|
||||
|
||||
# Gate-reset observations must use the same runtime activation settings as
|
||||
# the live setup list. If the config cannot be loaded, scan normally but do
|
||||
# not mutate reset state from an evaluation whose rules are unknown.
|
||||
@@ -765,6 +801,12 @@ async def scan_all_tickers(
|
||||
for index, (ticker_id, symbol) in enumerate(ticker_rows):
|
||||
if progress_callback is not None:
|
||||
progress_callback(index, total, symbol)
|
||||
if ticker_id in fundamentals_blocked_ids:
|
||||
logger.info(
|
||||
"Skipping %s: unresolved or unavailable SEC fundamentals",
|
||||
symbol,
|
||||
)
|
||||
continue
|
||||
# Refresh Structural S/R once, then scores. get_sr_levels is read-only;
|
||||
# without this recalculate the score path would see yesterday's zones.
|
||||
# A refresh failure still scans the ticker: qualification re-gates on
|
||||
@@ -795,6 +837,7 @@ async def scan_all_tickers(
|
||||
volatility_percentile=(ranks.get(symbol) or {}).get("volatility_percentile"),
|
||||
primary_min_rr=PRIMARY_TARGET_MIN_RR,
|
||||
scan_run_id=scan_run_id,
|
||||
fundamentals_eligible=True,
|
||||
)
|
||||
all_setups.extend(setups)
|
||||
if activation is not None:
|
||||
@@ -882,6 +925,26 @@ async def get_trade_setups(
|
||||
stmt = stmt.where(TradeSetup.recommended_action == recommended_action)
|
||||
excluded_ticker_ids: set[int] = set()
|
||||
reentry_gate_locks: dict[int, datetime] = {}
|
||||
try:
|
||||
excluded_ticker_ids.update(
|
||||
await fundamentals_quality_service.blocked_ticker_ids(db)
|
||||
)
|
||||
except Exception:
|
||||
await db.rollback()
|
||||
logger.exception(
|
||||
"Could not resolve fundamentals quality; hiding actionable setups"
|
||||
)
|
||||
await system_event_service.log_event_standalone(
|
||||
severity="error",
|
||||
source="rr_scanner",
|
||||
code="fundamentals_quality_unavailable",
|
||||
message=(
|
||||
"The fundamentals quality gate could not be evaluated; actionable "
|
||||
"setups were hidden until the metadata check recovers."
|
||||
),
|
||||
dedup_key="rr_scanner:fundamentals_quality_unavailable",
|
||||
)
|
||||
return []
|
||||
if exclude_open_trade_tickers:
|
||||
# Manual book only. The shadow book holds the *top-ranked* names by
|
||||
# construction, so letting its positions hide setups would leave the
|
||||
|
||||
@@ -8,7 +8,9 @@ the daily filing index — behind one client that honors SEC's fair-access polic
|
||||
- request spacing well under the 10 req/s limit;
|
||||
- exponential backoff + retry on 429;
|
||||
- **403 → alert and stop** (raise ``SecForbiddenError``), never a retry-loop — a
|
||||
403 means the UA or request pattern is wrong and retrying won't fix it.
|
||||
403 means the UA or request pattern is wrong and retrying won't fix it. The one
|
||||
exception is S3's ``AccessDenied`` on an ``/Archives/`` path, which is how the
|
||||
bucket reports an absent file (``_is_absent_archive_key``).
|
||||
|
||||
Parsing lives here (index fixed-width, submissions pagination); DB writes and the
|
||||
snapshot mapping live in the importer. No conditional GETs — the companyfacts
|
||||
@@ -53,11 +55,44 @@ class SecForbiddenError(SecError):
|
||||
|
||||
|
||||
class SecNotFoundError(SecError):
|
||||
"""SEC returned 404 — the resource does not exist (e.g. no index for a day).
|
||||
"""The resource does not exist (e.g. no daily index published for a day).
|
||||
|
||||
The *only* error a caller may treat as 'missing' — every other SecError
|
||||
(403, exhausted retries, 5xx, timeout) must propagate so a fetch failure is
|
||||
never mistaken for an empty result."""
|
||||
(fair-access rejection, exhausted retries, 5xx, timeout) must propagate so a
|
||||
fetch failure is never mistaken for an empty result.
|
||||
|
||||
Raised for a 404, and for the one 403 that also means "absent": see
|
||||
``_is_absent_archive_key``."""
|
||||
|
||||
|
||||
def _is_absent_archive_key(url: str, resp: httpx.Response) -> bool:
|
||||
"""True when a 403 means "this file does not exist", not "you are blocked".
|
||||
|
||||
``www.sec.gov/Archives`` is served straight out of an S3 bucket that grants
|
||||
no ``s3:ListBucket``, so a missing key cannot be answered with 404 — S3
|
||||
returns **403 with its ``AccessDenied`` XML** instead. SEC publishes a daily
|
||||
index only for business days, so every weekend and market holiday inside an
|
||||
incremental walk lands on exactly this response (verified 2026-07-30:
|
||||
``form.20260725.idx``, a Saturday, 403s while the Friday and Monday files
|
||||
return 200 on the same User-Agent).
|
||||
|
||||
A genuine fair-access rejection is distinguishable and must stay fatal: it is
|
||||
SEC's WAF interstitial — ``text/html``, "Your Request Originates from an
|
||||
Undeclared Automated Tool" — and it is returned for files that *do* exist,
|
||||
on any path. Hence the narrow gate: the Archives prefix plus S3's own error
|
||||
document. Nothing else may be downgraded to "missing"."""
|
||||
try:
|
||||
parsed = httpx.URL(url)
|
||||
except (TypeError, ValueError): # pragma: no cover — url comes from us
|
||||
return False
|
||||
if parsed.host != "www.sec.gov" or not parsed.path.startswith("/Archives/"):
|
||||
return False
|
||||
if "xml" not in resp.headers.get("Content-Type", "").lower():
|
||||
return False
|
||||
try:
|
||||
return "<Code>AccessDenied</Code>" in resp.text
|
||||
except (UnicodeDecodeError, httpx.HTTPError): # pragma: no cover
|
||||
return False
|
||||
|
||||
|
||||
def _looks_like_contact_email(ua: str) -> bool:
|
||||
@@ -155,6 +190,8 @@ class SecClient:
|
||||
|
||||
code = resp.status_code
|
||||
if code == 403:
|
||||
if _is_absent_archive_key(url, resp):
|
||||
raise SecNotFoundError(f"SEC 403/AccessDenied (absent) for {url}")
|
||||
raise SecForbiddenError(
|
||||
f"SEC 403 for {url} — User-Agent/pattern rejected; set a real "
|
||||
"sec_user_agent contact email"
|
||||
@@ -252,7 +289,17 @@ class SecClient:
|
||||
try:
|
||||
text = await self.get_text(url)
|
||||
except SecNotFoundError:
|
||||
logger.info("no daily index for %s (404)", day)
|
||||
# Absent on a weekend is routine (SEC publishes business days only); on a
|
||||
# weekday it is either a market holiday or something worth a look — a SEC
|
||||
# hiccup, or a rejection page misread as absent, would otherwise let the
|
||||
# importer advance past real filings silently. Log-level only, no alert:
|
||||
# cheaper than carrying a holiday calendar just to stay quiet ~10 days/yr.
|
||||
logger.log(
|
||||
logging.INFO if day.weekday() >= 5 else logging.WARNING,
|
||||
"no daily index published for %s (%s)",
|
||||
day,
|
||||
f"{day:%a}",
|
||||
)
|
||||
return [] # weekend/holiday/not-yet-published; other errors propagate
|
||||
return _parse_form_index(text)
|
||||
|
||||
|
||||
@@ -8,6 +8,9 @@ fixture and verifiable against a real companyfacts pull.
|
||||
The load-bearing rules (design Decision 2 + review):
|
||||
- Period identity comes from `end == submissions.reportDate`, never `fy/fp`
|
||||
(fy/fp is the *filing's* context; comparatives inside a filing repeat it).
|
||||
This applies to the stored `fiscal_year`/`fiscal_period` too: they are derived
|
||||
from `reportDate` against the issuer's `fiscalYearEnd` (see `_period_identity`),
|
||||
because SEC's fy/fp collide and invert often enough to break the quarter chain.
|
||||
- Duration facts are stored as **cumulative YTD**: pick the fact whose span
|
||||
matches the fiscal-period-to-date length (Q1≈3mo … FY≈12mo) within tolerance.
|
||||
If no YTD-length fact exists, store null — never a discrete masquerading as YTD.
|
||||
@@ -15,7 +18,10 @@ The load-bearing rules (design Decision 2 + review):
|
||||
is a single consolidated value: the cover-page `dei` fact (its own cover-date
|
||||
`end` stored separately) if present, else `us-gaap:CommonStockSharesOutstanding`
|
||||
at period end (e.g. Alphabet has no `dei` fact) — never a class sum or the
|
||||
weighted-average/diluted count.
|
||||
weighted-average/diluted count. Multi-class issuers report it per class, which
|
||||
is dimensional and therefore absent from companyfacts entirely, so
|
||||
`weighted_avg_diluted_shares` is stored alongside as an explicit fallback for
|
||||
market cap — a separate column, never backfilled into `shares_outstanding`.
|
||||
- Cash and debt composites are aggregate-first and mutually exclusive (each
|
||||
source tag counted at most once).
|
||||
|
||||
@@ -29,7 +35,7 @@ from __future__ import annotations
|
||||
import logging
|
||||
import math
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import date, datetime
|
||||
from datetime import date, datetime, timedelta
|
||||
from typing import Any, NamedTuple
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
@@ -37,14 +43,35 @@ logger = logging.getLogger(__name__)
|
||||
# Expected YTD span (days) per fiscal period; a duration fact must land within
|
||||
# tolerance of this to count as the period's cumulative value.
|
||||
_EXPECTED_YTD_DAYS = {"Q1": 91, "Q2": 182, "Q3": 273, "FY": 365}
|
||||
_YTD_TOLERANCE_DAYS = 20 # covers 52/53-week fiscal calendars
|
||||
# Period identity (see _period_identity): how far a quarter end sits before its
|
||||
# fiscal-year end, and how far a fiscal-year end may drift from the nominal MMDD.
|
||||
# The quarter bands are 91 days apart, so ±35 stays unambiguous even for a 4-4-5
|
||||
# filer whose 16-week Q4 puts Q3 112 days out.
|
||||
_QUARTER_DAYS_TO_FY_END = {"Q1": 273, "Q2": 182, "Q3": 91}
|
||||
_QUARTER_TOLERANCE_DAYS = 35
|
||||
_FYE_DRIFT_TOLERANCE_DAYS = 21
|
||||
# Covers 52/53-week calendars *and* 4-4-5 retail ones (12/12/12/16 weeks), whose
|
||||
# YTD-Q3 is 36 weeks = 251-252 days and missed a 20-day tolerance by ~2 -- so
|
||||
# COST/PEP lost Q3 every year, breaking the quarter chain and nulling TTM + YoY.
|
||||
# Q1 84d, Q2 168d and FY 364d were always inside. Adjacent periods stay
|
||||
# unambiguous at 25 (66-116, 157-207, 248-298, 340-390).
|
||||
_YTD_TOLERANCE_DAYS = 25
|
||||
|
||||
# us-gaap duration concepts (money), priority order; first present wins.
|
||||
_DURATION_USD = {
|
||||
# Order is load-bearing (first present wins) and the tail entries are
|
||||
# deliberately *appended*: every issuer that already resolved keeps the same
|
||||
# concept, and only issuers that resolved to nothing gain a value.
|
||||
# - IncludingAssessedTax: REITs/consumer filers that tag only this variant
|
||||
# (e.g. ARE, KHC) reported no revenue at all.
|
||||
# - RevenuesNetOfInterestExpense: the banks' total-revenue tag. JPM/GS/WFC
|
||||
# tag it in every 10-Q and `Revenues` only (if at all) in the 10-K.
|
||||
"revenue": [
|
||||
"RevenueFromContractWithCustomerExcludingAssessedTax",
|
||||
"Revenues",
|
||||
"SalesRevenueNet",
|
||||
"RevenueFromContractWithCustomerIncludingAssessedTax",
|
||||
"RevenuesNetOfInterestExpense",
|
||||
],
|
||||
"net_income": ["NetIncomeLoss"],
|
||||
"operating_income": ["OperatingIncomeLoss"],
|
||||
@@ -62,7 +89,27 @@ _DURATION_USD = {
|
||||
"DepreciationAndAmortization",
|
||||
],
|
||||
}
|
||||
_EPS_CONCEPTS = ["EarningsPerShareDiluted"] # unit USD/shares
|
||||
# unit USD/shares. Appended (not reordered) so any issuer that already resolved
|
||||
# keeps the same concept. REG tags only the continuing-operations variant on every
|
||||
# filing; FCX switches by form type -- EarningsPerShareDiluted in its 10-Qs, the
|
||||
# continuing-ops tag in its 10-K -- which nulled the FY row and killed Q4 + TTM.
|
||||
# The basic variants are a last resort for a period that tags no diluted EPS at
|
||||
# all (PPL's 2026 Q1). Basic ignores option/convert dilution so it slightly
|
||||
# overstates EPS (~1.2% for PPL), but only fires when diluted is entirely absent,
|
||||
# and high-dilution names always tag diluted -- so it never displaces a real one.
|
||||
_EPS_CONCEPTS = [
|
||||
"EarningsPerShareDiluted",
|
||||
"IncomeLossFromContinuingOperationsPerDilutedShare",
|
||||
"EarningsPerShareBasic",
|
||||
"IncomeLossFromContinuingOperationsPerBasicShare",
|
||||
]
|
||||
# Weighted-average diluted share count (unit "shares"), the market-cap fallback
|
||||
# for multi-class issuers whose cover-page count is dimensional and therefore
|
||||
# absent from companyfacts. Always present, since EPS is computed from it.
|
||||
_WEIGHTED_AVG_SHARE_CONCEPTS = [
|
||||
"WeightedAverageNumberOfDilutedSharesOutstanding",
|
||||
"WeightedAverageNumberOfSharesOutstandingBasicAndDiluted",
|
||||
]
|
||||
# us-gaap instant (balance-sheet) concepts, at end == reportDate.
|
||||
_CASH = ["CashAndCashEquivalentsAtCarryingValue"]
|
||||
_ST_INVESTMENTS = ["ShortTermInvestments", "MarketableSecuritiesCurrent"] # pick one
|
||||
@@ -104,6 +151,7 @@ class SnapshotRow:
|
||||
total_debt: float | None = None
|
||||
shares_outstanding: float | None = None
|
||||
shares_outstanding_date: date | None = None
|
||||
weighted_avg_diluted_shares: float | None = None
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -127,14 +175,23 @@ def parse_snapshots(
|
||||
companyfacts: dict[str, Any],
|
||||
filings: dict[str, FilingMeta],
|
||||
accessions: set[str],
|
||||
fiscal_year_end: str | None = None,
|
||||
) -> ParseResult:
|
||||
"""Build snapshot rows for ``accessions`` (those with facts + filing meta).
|
||||
|
||||
``fiscal_year_end`` is the issuer's declared ``submissions.fiscalYearEnd``
|
||||
(MMDD) and seeds period identity (see ``_period_identity``), making it
|
||||
independent of SEC's unreliable fy/fp fields. It is only a hint: the issuer's
|
||||
own 10-K period ends override it (see ``resolve_fiscal_year_end``). With
|
||||
neither available, the old fy/fp behaviour is used.
|
||||
|
||||
``skipped_filings`` = no row produced (missing facts/meta or no usable period
|
||||
identity); ``field_issues`` = a row was produced but a field is null/ambiguous.
|
||||
Callers must not use field issues as failed-row coverage.
|
||||
"""
|
||||
cik = f"{int(companyfacts['cik']):010d}"
|
||||
# The declared value is only a hint; the issuer's own 10-Ks are authoritative.
|
||||
fiscal_year_end = resolve_fiscal_year_end(filings, fiscal_year_end)
|
||||
by_accn = _index_by_accession(companyfacts)
|
||||
result = ParseResult()
|
||||
for accn in accessions:
|
||||
@@ -143,7 +200,7 @@ def parse_snapshots(
|
||||
if meta is None or not facts:
|
||||
result.skipped_filings.append({"accession": accn, "reason": "no facts or filing metadata"})
|
||||
continue
|
||||
row, note = _parse_one(cik, accn, facts, meta)
|
||||
row, note = _parse_one(cik, accn, facts, meta, fiscal_year_end)
|
||||
if row is None:
|
||||
result.skipped_filings.append({"accession": accn, "reason": note or "unparseable"})
|
||||
continue
|
||||
@@ -190,12 +247,18 @@ def _index_by_accession(companyfacts: dict[str, Any]) -> dict[str, list[Fact]]:
|
||||
|
||||
|
||||
def _parse_one(
|
||||
cik: str, accn: str, facts: list[Fact], meta: FilingMeta
|
||||
cik: str, accn: str, facts: list[Fact], meta: FilingMeta,
|
||||
fiscal_year_end: str | None = None,
|
||||
) -> tuple[SnapshotRow | None, str | None]:
|
||||
"""Returns (row, note). row is None when there's no usable period identity;
|
||||
note is a validation reason (row-skip reason when row is None, else a
|
||||
field-level issue such as ambiguous shares)."""
|
||||
fy, fp = _fiscal_context(facts, meta.report_date)
|
||||
fy, fp = _period_identity(meta, fiscal_year_end)
|
||||
if fy is None or fp is None:
|
||||
# No fiscal calendar, or a period the calendar cannot place (a transition
|
||||
# period). Fall back to the filing's own context: an imperfect label still
|
||||
# beats dropping the filing entirely.
|
||||
fy, fp = _fiscal_context(facts, meta.report_date)
|
||||
if fy is None or fp not in _EXPECTED_YTD_DAYS:
|
||||
return None, "no usable period identity"
|
||||
|
||||
@@ -227,9 +290,109 @@ def _parse_one(
|
||||
shares, shares_date, ambiguous = _select_shares(facts, meta.report_date)
|
||||
row.shares_outstanding = shares
|
||||
row.shares_outstanding_date = shares_date
|
||||
row.weighted_avg_diluted_shares = _select_weighted_avg_shares(facts, meta.report_date)
|
||||
|
||||
return row, ("ambiguous shares outstanding" if ambiguous else None)
|
||||
|
||||
|
||||
def resolve_fiscal_year_end(
|
||||
filings: dict[str, FilingMeta], declared: str | None
|
||||
) -> str | None:
|
||||
"""The issuer's fiscal-year-end MMDD, preferring its own 10-K period ends.
|
||||
|
||||
``submissions.fiscalYearEnd`` is *not* reliable: Franklin Resources (BEN)
|
||||
declares 1231 while every one of its 10-Ks ends 09-30. Trusting it put BEN's
|
||||
fiscal Q1 (Dec) 0 days from the claimed year end — matching no quarter band —
|
||||
and labelled its fiscal Q2 (Mar) as Q1, colliding two periods on one key and
|
||||
destroying the quarter chain.
|
||||
|
||||
A 10-K's reportDate **is** the fiscal year end by definition, so it wins
|
||||
whenever one is available; the declared value is only a fallback for an issuer
|
||||
with no annual filing in the set. The most recent 10-K is used, so an issuer
|
||||
that changed its year end is measured against its current calendar.
|
||||
"""
|
||||
annual = [m.report_date for m in filings.values() if m.form.startswith("10-K")]
|
||||
if annual:
|
||||
latest = max(annual)
|
||||
return f"{latest.month:02d}{latest.day:02d}"
|
||||
return declared
|
||||
|
||||
|
||||
def _period_identity(
|
||||
meta: FilingMeta, fiscal_year_end: str | None
|
||||
) -> tuple[int | None, str | None]:
|
||||
"""(fiscal_year, fiscal_period) from the period end and the issuer's fiscal
|
||||
calendar — never from the fy/fp fields.
|
||||
|
||||
SEC's fy/fp describe the *filing*, and they are unreliable as period identity:
|
||||
observed in production, a 10-Q labelled ``FY`` (BXP), a year ending 2025-12-31
|
||||
labelled 2024 (FRT, a December filer), a year ending 2025-06-27 labelled 2027
|
||||
(STX), and four different period ends all labelled 2022 Q3 (PPL). Because
|
||||
readers key on (fiscal_year, fiscal_period), colliding labels silently discard
|
||||
a period and inverted ones scramble the quarter chain — nulling TTM and YoY.
|
||||
|
||||
``period_end`` is authoritative, so identity is derived from it: the form
|
||||
decides FY vs quarter, and distance to the fiscal-year end decides which
|
||||
quarter. Labels need not match the issuer's own naming — a filer whose year
|
||||
ends in early January (DPZ) shifts by one — they need to be unique, monotonic
|
||||
and YoY-aligned, which is all the derivation asks of them. Nothing outside the
|
||||
derivation reads these columns.
|
||||
|
||||
Known limitation: ``fiscalYearEnd`` is the issuer's *current* calendar, so a
|
||||
company that has changed its fiscal year end gets its historical periods
|
||||
measured against the new one. The quarter tolerance shunts most of those to
|
||||
the fy/fp fallback, and a same-key collision resolves newest-wins, so the
|
||||
failure mode is a degraded old year rather than a scrambled current one.
|
||||
"""
|
||||
fy = _fiscal_year_of(meta.report_date, fiscal_year_end)
|
||||
if fy is None:
|
||||
return None, None
|
||||
if meta.form.startswith("10-K"):
|
||||
return fy, "FY"
|
||||
nominal_end = _nominal_fy_end(fy, fiscal_year_end)
|
||||
if nominal_end is None:
|
||||
return None, None
|
||||
remaining = (nominal_end - meta.report_date).days
|
||||
best = min(
|
||||
_QUARTER_DAYS_TO_FY_END,
|
||||
key=lambda k: abs(_QUARTER_DAYS_TO_FY_END[k] - remaining),
|
||||
)
|
||||
if abs(_QUARTER_DAYS_TO_FY_END[best] - remaining) > _QUARTER_TOLERANCE_DAYS:
|
||||
return None, None # transition period or odd filing — let the caller fall back
|
||||
return fy, best
|
||||
|
||||
|
||||
def _nominal_fy_end(year: int, fiscal_year_end: str | None) -> date | None:
|
||||
"""The issuer's nominal fiscal-year end in ``year`` from a MMDD string."""
|
||||
if not fiscal_year_end or len(fiscal_year_end) != 4 or not fiscal_year_end.isdigit():
|
||||
return None
|
||||
month, day = int(fiscal_year_end[:2]), int(fiscal_year_end[2:])
|
||||
if not 1 <= month <= 12 or not 1 <= day <= 31:
|
||||
return None
|
||||
while day > 28: # 52/53-week ends land on 0229/0230/0231 in some filings
|
||||
try:
|
||||
return date(year, month, day)
|
||||
except ValueError:
|
||||
day -= 1
|
||||
return date(year, month, day)
|
||||
|
||||
|
||||
def _fiscal_year_of(period_end: date, fiscal_year_end: str | None) -> int | None:
|
||||
"""Which fiscal year ``period_end`` belongs to.
|
||||
|
||||
A 52/53-week calendar's real year end drifts around the nominal MMDD (and can
|
||||
cross the calendar year), so allow drift before rolling into the next year.
|
||||
"""
|
||||
nominal = _nominal_fy_end(period_end.year, fiscal_year_end)
|
||||
if nominal is None:
|
||||
return None
|
||||
return (
|
||||
period_end.year
|
||||
if period_end <= nominal + timedelta(days=_FYE_DRIFT_TOLERANCE_DAYS)
|
||||
else period_end.year + 1
|
||||
)
|
||||
|
||||
|
||||
def _fiscal_context(facts: list[Fact], report_date: date) -> tuple[int | None, str | None]:
|
||||
"""The filing's (fy, fp) taken as the majority context among the facts that
|
||||
end at reportDate (the current-period facts, which share the filing's
|
||||
@@ -352,6 +515,34 @@ def _select_shares(
|
||||
return None, None, False # simply absent — not a conflict
|
||||
|
||||
|
||||
def _select_weighted_avg_shares(facts: list[Fact], report_date: date) -> float | None:
|
||||
"""The most recent quarter's weighted-average diluted share count.
|
||||
|
||||
Deliberately the **shortest** duration ending at reportDate, not the YTD one:
|
||||
the shorter the window the closer the average sits to the current count, which
|
||||
is what a market cap wants. Measured against issuers where the true
|
||||
point-in-time count is available, the quarter average is within ~0.6%.
|
||||
"""
|
||||
best: tuple[int, float] | None = None
|
||||
for concept in _WEIGHTED_AVG_SHARE_CONCEPTS:
|
||||
for f in facts:
|
||||
if (
|
||||
f.taxonomy != "us-gaap"
|
||||
or f.concept != concept
|
||||
or f.unit != "shares"
|
||||
or f.start is None
|
||||
or f.end != report_date
|
||||
or f.val <= 0
|
||||
):
|
||||
continue
|
||||
span = (f.end - f.start).days
|
||||
if best is None or span < best[0]:
|
||||
best = (span, float(f.val))
|
||||
if best is not None:
|
||||
return best[1] # first present concept wins, as elsewhere
|
||||
return None
|
||||
|
||||
|
||||
def _d(value: Any) -> date | None:
|
||||
if not value:
|
||||
return None
|
||||
|
||||
@@ -18,25 +18,50 @@ Guardrails (design + reviews):
|
||||
a tracked XBRL index accession missing from Company Facts fails the run (the two
|
||||
are separate SEC products that can lag) so we retry rather than record a
|
||||
null/partial snapshot. Non-XBRL amendments are skipped with a recorded reason.
|
||||
A failure here blocks every later run (``source_max_date`` only advances on a
|
||||
promoted run), so it names the offending filings in the alert and separates the
|
||||
causes — ``not_in_companyfacts`` (facts lag) vs ``not_in_submissions`` (the
|
||||
index row is absent from the issuer's own filing list, which no retry fixes).
|
||||
- **Co-registrant recovery**, because "missing from Company Facts" is often not
|
||||
missing at all: SEC files some combined parent/subsidiary filings' XBRL under
|
||||
the co-registrant's CIK, so the ticker-carrying parent's own file never gets
|
||||
that accession. The daily index lists every co-registrant of an accession, so
|
||||
the facts are found there and re-stamped to the real filer — guarded by a
|
||||
share-count continuity check so a subsidiary's standalone numbers can never be
|
||||
stored as the parent's. Confirmed 2026-07-27 (NEE via FPL, DOW via Dow Chemical)
|
||||
and it is not transient: an NEE filing misattributed in 2014 is still misfiled.
|
||||
- **Bounded blocking.** Anything still unresolvable after ``MISSING_XBRL_RETRY_DAYS``
|
||||
stops failing the whole import and enters a durable retry queue. The scheduled
|
||||
importer retries queued accessions automatically, while the affected issuer is
|
||||
excluded from actionable setups until its filing is recovered.
|
||||
- ``promote`` inserts snapshots ``ON CONFLICT (accession) DO NOTHING`` (immutable),
|
||||
reports differing existing accessions, and applies ticker updates in the same
|
||||
transaction.
|
||||
- ``reparse=True`` is the one exception to immutability, and it is deliberate:
|
||||
it restages every accession with the current parser and **rewrites** the rows
|
||||
that now reconstruct differently. Immutability protects SEC's record (one row
|
||||
per accession, amendments retained) — but the stored row is *our* reconstruction,
|
||||
so after a parser fix, keeping it is preserving a stale cache, not history.
|
||||
Manually invoked through ``scripts/reparse_fundamentals.py``; never scheduled.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
from collections import defaultdict
|
||||
from dataclasses import dataclass, field
|
||||
from collections import Counter, defaultdict
|
||||
from dataclasses import dataclass, field, replace
|
||||
from datetime import date, datetime, timedelta, timezone
|
||||
from typing import Any, Callable
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy import delete, select, update
|
||||
|
||||
from app.database import insert_for_session
|
||||
from app.models.data_import_run import DataImportRun
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.sec_filing_gap import SecFilingGap
|
||||
from app.models.system_event import SystemEvent
|
||||
from app.services import fundamentals_quality_service
|
||||
from app.services import sec_facts_parser as parser
|
||||
from app.services import sec_universe
|
||||
from app.services.data_import import STATUS_PROMOTED, ValidationResult
|
||||
@@ -51,13 +76,24 @@ _XBRL_FORMS = {"10-K", "10-Q", "10-K/A", "10-Q/A"}
|
||||
# On the one-time backfill, require this fraction of tracked issuers to yield at
|
||||
# least one snapshot (guards a broken fetch/parse from promoting a hollow table).
|
||||
MIN_BACKFILL_COVERAGE = 0.5
|
||||
# How long an index accession may stay unresolvable before the run stops failing
|
||||
# on it. Genuine index↔facts lag clears within a day (a weekend stretches it to
|
||||
# three); past that it is misfiled, not late, and blocking forever costs more
|
||||
# than the missing filing does — see the unresolved-filing guardrail below.
|
||||
MISSING_XBRL_RETRY_DAYS = 3
|
||||
FILING_GAP_ESCALATE_DAYS = 14
|
||||
# Share-count band a co-registrant-recovered row must land in, relative to the
|
||||
# issuer's own last snapshot. Wide enough for buybacks/issuance, nowhere near
|
||||
# wide enough to let a subsidiary shell's token float through (see _shares_continuous).
|
||||
RECOVERY_SHARES_MIN = 0.5
|
||||
RECOVERY_SHARES_MAX = 2.0
|
||||
|
||||
_SNAPSHOT_COLS = (
|
||||
"cik", "accession", "form", "filed_date", "accepted_at", "period_start",
|
||||
"period_end", "fiscal_year", "fiscal_period", "revenue", "net_income",
|
||||
"operating_income", "diluted_eps", "cfo", "capex", "depreciation_amortization",
|
||||
"cash_and_st_investments", "total_debt", "shares_outstanding",
|
||||
"shares_outstanding_date",
|
||||
"shares_outstanding_date", "weighted_avg_diluted_shares",
|
||||
)
|
||||
# Compare ALL source fields (every column except the accession key) to flag a
|
||||
# differing existing accession — immutable, so we report, never mutate.
|
||||
@@ -72,9 +108,17 @@ class StagedFundamentals:
|
||||
skipped_filings: list[dict[str, str]] = field(default_factory=list)
|
||||
field_issues: list[dict[str, str]] = field(default_factory=list)
|
||||
skipped_non_xbrl: list[dict[str, str]] = field(default_factory=list)
|
||||
missing_xbrl: list[dict[str, str]] = field(default_factory=list)
|
||||
# Index rows we could not resolve to Company Facts, with a per-row reason
|
||||
# (not_in_companyfacts | not_in_submissions | ...) — see _missing().
|
||||
missing_xbrl: list[dict[str, Any]] = field(default_factory=list)
|
||||
# Accessions parsed out of a co-registrant's Company Facts file.
|
||||
recovered: list[dict[str, Any]] = field(default_factory=list)
|
||||
invalid_payloads: list[dict[str, str]] = field(default_factory=list)
|
||||
existing_accessions: set[str] = field(default_factory=set)
|
||||
# Tracked issuers whose registrant has NO XBRL 10-K/10-Q at all: they can
|
||||
# never yield a snapshot, so this is a resolution problem (a ticker pointed
|
||||
# at a successor shell), not missing data. See sec_universe.CIK_OVERRIDES_KEY.
|
||||
no_xbrl_filings: list[dict[str, Any]] = field(default_factory=list)
|
||||
discrepancies: list[dict[str, Any]] = field(default_factory=list)
|
||||
backfill: bool = False
|
||||
issuers_fetched: int = 0
|
||||
@@ -93,12 +137,24 @@ class SecFundamentalsImporter:
|
||||
*,
|
||||
client_factory: Callable[[], SecClient] | None = None,
|
||||
today: date | None = None,
|
||||
reparse: bool = False,
|
||||
) -> None:
|
||||
self._client_factory = client_factory or (lambda: SecClient())
|
||||
self.today = today or _now().date()
|
||||
# Reparse: re-derive every stored accession with the CURRENT parser and
|
||||
# rewrite the ones that now reconstruct differently. Snapshots are
|
||||
# immutable with respect to SEC (one row per accession, amendments kept),
|
||||
# but the stored row is *our reconstruction* — when a parser bug is fixed,
|
||||
# leaving it stale is not immutability, it is a stale cache. Manually
|
||||
# invoked via scripts/reparse_fundamentals.py; never scheduled.
|
||||
self.reparse = reparse
|
||||
# cached by detect_revision, consumed by stage:
|
||||
self._resolved: ResolvedUniverse | None = None
|
||||
self._index_rows: list[dict[str, Any]] = []
|
||||
# accession -> the OTHER CIKs the daily index lists it under (co-registrants
|
||||
# of a combined filing). Only populated for accessions a tracked issuer filed.
|
||||
self._coregistrants: dict[str, list[int]] = {}
|
||||
self._retry_rows: list[dict[str, Any]] = []
|
||||
self._latest_index_date: date | None = None
|
||||
self._backfill = False
|
||||
|
||||
@@ -111,7 +167,11 @@ class SecFundamentalsImporter:
|
||||
self._latest_index_date = await client.latest_index_date(self.today)
|
||||
if self._latest_index_date is None:
|
||||
raise SecError("no EDGAR daily index available")
|
||||
if last_processed is None:
|
||||
# Reparse needs every accession restaged, not just those filed since
|
||||
# the last run — the facts a fixed parser now accepts were never
|
||||
# stored, so a reparse cannot be served from the database.
|
||||
self._coregistrants = {}
|
||||
if last_processed is None or self.reparse:
|
||||
self._backfill = True
|
||||
self._index_rows = []
|
||||
else:
|
||||
@@ -120,9 +180,19 @@ class SecFundamentalsImporter:
|
||||
client, last_processed, self._latest_index_date
|
||||
)
|
||||
content = sec_universe.index_content_hash(self._index_rows)
|
||||
return sec_universe.compose_revision(
|
||||
revision = sec_universe.compose_revision(
|
||||
self._latest_index_date, content, self._resolved.symbol_to_cik
|
||||
)
|
||||
self._retry_rows = []
|
||||
if not self._backfill:
|
||||
self._retry_rows = await self._retry_backlog(
|
||||
db,
|
||||
set(self._resolved.cik_to_ticker_ids),
|
||||
)
|
||||
# Company Facts can change while the daily index revision stays fixed.
|
||||
# Returning None deliberately bypasses the framework's no-op gate so a
|
||||
# scheduled run retries every active gap.
|
||||
return None if self._retry_rows else revision
|
||||
|
||||
async def stage(self, db) -> StagedFundamentals:
|
||||
assert self._resolved is not None, "detect_revision must run first"
|
||||
@@ -130,10 +200,32 @@ class SecFundamentalsImporter:
|
||||
staged = StagedFundamentals(resolved=resolved, backfill=self._backfill)
|
||||
|
||||
cik_to_tids = resolved.cik_to_ticker_ids
|
||||
filed_by_cik: dict[int, list[str]] = defaultdict(list)
|
||||
# Whole index rows (not bare accessions): form + index date are what make
|
||||
# an unresolvable filing diagnosable without re-walking the index by hand.
|
||||
filed_by_cik: dict[int, list[dict[str, Any]]] = defaultdict(list)
|
||||
for r in self._index_rows:
|
||||
if r["cik"] in cik_to_tids:
|
||||
filed_by_cik[r["cik"]].append(r["accession"])
|
||||
filed_by_cik[r["cik"]].append(r)
|
||||
|
||||
# Promoted-around filings live in a small durable retry queue, including
|
||||
# the one-time migration backfill. Merge them into normal incremental
|
||||
# work so the scheduled importer heals them without operator action.
|
||||
if not self._backfill:
|
||||
seen = {
|
||||
(int(cik), row["accession"])
|
||||
for cik, rows in filed_by_cik.items()
|
||||
for row in rows
|
||||
}
|
||||
for row in self._retry_rows:
|
||||
cik = int(row["cik"])
|
||||
key = (cik, row["accession"])
|
||||
if key in seen:
|
||||
continue
|
||||
filed_by_cik[cik].append(row)
|
||||
seen.add(key)
|
||||
coregistrants = [int(value) for value in row.get("coregistrants") or []]
|
||||
if coregistrants:
|
||||
self._coregistrants[row["accession"]] = coregistrants
|
||||
|
||||
existing = await self._ciks_with_snapshots(db, set(cik_to_tids))
|
||||
if self._backfill:
|
||||
@@ -144,10 +236,15 @@ class SecFundamentalsImporter:
|
||||
backfill_ciks = {c for c in cik_to_tids if c not in existing}
|
||||
incremental_ciks = set(filed_by_cik) - backfill_ciks
|
||||
|
||||
# Continuity reference for co-registrant recovery, read once up front.
|
||||
last_shares = await self._last_shares_outstanding(db, set(cik_to_tids))
|
||||
|
||||
async with self._client_factory() as client:
|
||||
for cik in sorted(backfill_ciks | incremental_ciks):
|
||||
is_backfill = cik in backfill_ciks
|
||||
await self._stage_issuer(client, cik, is_backfill, filed_by_cik, staged)
|
||||
await self._stage_issuer(
|
||||
client, cik, is_backfill, filed_by_cik, staged, last_shares
|
||||
)
|
||||
|
||||
# Read-only discrepancy detection: an accession we reconstructed that is
|
||||
# already stored, differing in ANY source field (immutable → report in
|
||||
@@ -164,7 +261,9 @@ class SecFundamentalsImporter:
|
||||
staged.discrepancies.append({"accession": row.accession, "fields": fields})
|
||||
return staged
|
||||
|
||||
async def _stage_issuer(self, client, cik, is_backfill, filed_by_cik, staged) -> None:
|
||||
async def _stage_issuer(
|
||||
self, client, cik, is_backfill, filed_by_cik, staged, last_shares
|
||||
) -> None:
|
||||
cf = await client.companyfacts(cik)
|
||||
bad = _companyfacts_structure_error(cf)
|
||||
if bad is not None:
|
||||
@@ -175,28 +274,90 @@ class SecFundamentalsImporter:
|
||||
return
|
||||
sub = await client.submissions(cik, include_history=is_backfill)
|
||||
xbrl_meta, nonxbrl = _filing_meta(sub)
|
||||
if not xbrl_meta:
|
||||
staged.no_xbrl_filings.append(
|
||||
{"cik": cik10(cik), "name": sub.get("name"), "tickers": sub.get("tickers")}
|
||||
)
|
||||
|
||||
fiscal_year_end = sub.get("fiscal_year_end")
|
||||
recovered_rows: list[SnapshotRow] = []
|
||||
index_rows = {
|
||||
row["accession"]: row for row in filed_by_cik.get(cik, [])
|
||||
}
|
||||
if is_backfill:
|
||||
accns = set(xbrl_meta)
|
||||
else:
|
||||
present = parser.companyfacts_accessions(cf)
|
||||
accns = set()
|
||||
for accn in filed_by_cik.get(cik, []):
|
||||
for index_row in filed_by_cik.get(cik, []):
|
||||
accn = index_row["accession"]
|
||||
if accn in nonxbrl:
|
||||
staged.skipped_non_xbrl.append({"cik": cik10(cik), "accession": accn})
|
||||
elif accn in xbrl_meta and accn in present:
|
||||
elif accn not in xbrl_meta:
|
||||
# The daily index lists it but the issuer's own filing list does
|
||||
# not (submissions lagging the index, no usable period metadata,
|
||||
# or a co-registrant filing). NOT a Company-Facts lag — separate
|
||||
# cause, separate fix, so it gets its own reason.
|
||||
staged.missing_xbrl.append(
|
||||
_missing(
|
||||
cik,
|
||||
index_row,
|
||||
"not_in_submissions",
|
||||
self.today,
|
||||
self._coregistrants.get(accn),
|
||||
)
|
||||
)
|
||||
elif accn in present:
|
||||
accns.add(accn)
|
||||
else:
|
||||
# XBRL (or unknown) filing not yet in Company Facts → the products
|
||||
# have lagged; fail+retry rather than record nothing for it.
|
||||
staged.missing_xbrl.append({"cik": cik10(cik), "accession": accn})
|
||||
# Filed, XBRL, but absent from this issuer's Company Facts. Try
|
||||
# the co-registrant file before treating it as missing data.
|
||||
row, source_cik = await self._recover_from_coregistrant(
|
||||
client, cik, accn, xbrl_meta, fiscal_year_end,
|
||||
last_shares.get(cik10(cik)), staged,
|
||||
)
|
||||
if row is not None:
|
||||
recovered_rows.append(row)
|
||||
staged.recovered.append({
|
||||
"cik": cik10(cik),
|
||||
"accession": accn,
|
||||
"source_cik": source_cik,
|
||||
"form": index_row.get("form"),
|
||||
})
|
||||
else:
|
||||
staged.missing_xbrl.append(_missing(
|
||||
cik, index_row,
|
||||
# Found, but it did not look like this issuer's own
|
||||
# numbers — say so; it is not the same as absent.
|
||||
"coregistrant_facts_rejected" if source_cik
|
||||
else "not_in_companyfacts",
|
||||
self.today,
|
||||
self._coregistrants.get(accn),
|
||||
))
|
||||
|
||||
result = parser.parse_snapshots(cf, xbrl_meta, accns)
|
||||
# fiscalYearEnd (MMDD) is what lets the parser derive period identity from
|
||||
# reportDate instead of SEC's unreliable fy/fp fields.
|
||||
result = parser.parse_snapshots(cf, xbrl_meta, accns, fiscal_year_end=fiscal_year_end)
|
||||
for skipped in result.skipped_filings:
|
||||
index_row = index_rows.get(skipped["accession"])
|
||||
if index_row is not None:
|
||||
# Facts are present but our parser cannot construct a snapshot.
|
||||
# A new index row keeps the normal grace period before promotion;
|
||||
# a row already read from the queue retains its _retry_queue marker
|
||||
# so later imports promote and retry without wedging the index.
|
||||
staged.missing_xbrl.append(_missing(
|
||||
cik,
|
||||
index_row,
|
||||
"parser_unusable",
|
||||
self.today,
|
||||
self._coregistrants.get(skipped["accession"]),
|
||||
))
|
||||
staged.rows.extend(result.rows)
|
||||
staged.rows.extend(recovered_rows)
|
||||
staged.skipped_filings.extend(result.skipped_filings)
|
||||
staged.field_issues.extend(result.field_issues)
|
||||
staged.issuers_fetched += 1
|
||||
if result.rows:
|
||||
if result.rows or recovered_rows:
|
||||
staged.issuers_with_rows += 1
|
||||
|
||||
# SIC proposal for this issuer's tickers (read-only; applied in promote).
|
||||
@@ -205,14 +366,66 @@ class SecFundamentalsImporter:
|
||||
for tid in staged.resolved.cik_to_ticker_ids.get(cik, []):
|
||||
staged.sic_updates.append((tid, sic, desc))
|
||||
|
||||
async def _recover_from_coregistrant(
|
||||
self, client, cik: int, accn: str, xbrl_meta, fiscal_year_end, reference, staged,
|
||||
) -> tuple[SnapshotRow | None, str | None]:
|
||||
"""Look for ``accn``'s facts in a co-registrant's Company Facts file.
|
||||
|
||||
SEC sometimes files a combined parent/subsidiary filing's XBRL under the
|
||||
co-registrant's CIK rather than the filer's — the ticker-carrying parent's
|
||||
own file simply never gets that accession. Verified 2026-07-27 for NEE
|
||||
(facts under Florida Power & Light) and DOW (under Dow Chemical); an NEE
|
||||
filing misattributed the same way in **2014** is still misattributed, so
|
||||
this does not self-correct and no amount of retrying recovers it.
|
||||
|
||||
Returns ``(row, source_cik)`` on success, ``(None, source_cik)`` when the
|
||||
facts were found but rejected by the continuity guard, ``(None, None)``
|
||||
when no co-registrant has them.
|
||||
|
||||
Incremental path only: the co-registrant map comes from the daily index,
|
||||
which a backfill/reparse does not walk. A reparse therefore recovers a
|
||||
filing only once SEC re-files it under the filer's own CIK.
|
||||
"""
|
||||
for co in self._coregistrants.get(accn, []):
|
||||
try:
|
||||
cf_co = await client.companyfacts(co)
|
||||
except SecError:
|
||||
continue # a co-registrant shell often has no facts file at all
|
||||
if _companyfacts_structure_error(cf_co) is not None:
|
||||
continue
|
||||
result = parser.parse_snapshots(
|
||||
cf_co, xbrl_meta, {accn}, fiscal_year_end=fiscal_year_end
|
||||
)
|
||||
if not result.rows:
|
||||
continue
|
||||
row = result.rows[0]
|
||||
if not _shares_continuous(row.shares_outstanding, reference):
|
||||
return None, cik10(co)
|
||||
# A recovered row is the one most worth flagging, so its parser caveats
|
||||
# travel with it rather than being dropped on the way out.
|
||||
staged.field_issues.extend(result.field_issues)
|
||||
# parse_snapshots stamps the CIK of the payload it read — re-stamp to
|
||||
# the issuer that actually filed, or the row lands under the shell.
|
||||
return replace(row, cik=cik10(cik)), cik10(co)
|
||||
return None, None
|
||||
|
||||
async def validate(self, db, staged: StagedFundamentals) -> ValidationResult:
|
||||
messages: list[str] = []
|
||||
|
||||
# Consistency gate — before any write.
|
||||
if staged.missing_xbrl:
|
||||
# Consistency gate — before any write. Only filings still inside the retry
|
||||
# window block: a failure here stops every later run too (source_max_date
|
||||
# advances on promotion alone), so blocking forever on a filing SEC has
|
||||
# misfiled would cost far more than the one filing it withholds. Older
|
||||
# ones are carried by promote() as a warning instead. The message names
|
||||
# the filings: "which ones" has to be in the alert itself, not merely
|
||||
# reconstructible by re-walking the index.
|
||||
blocking = _within_retry_window(staged.missing_xbrl)
|
||||
aged_out = _past_retry_window(staged.missing_xbrl)
|
||||
if blocking:
|
||||
messages.append(
|
||||
f"{len(staged.missing_xbrl)} tracked XBRL filing(s) not yet in "
|
||||
"Company Facts (index/facts lag) — retry"
|
||||
f"{len(blocking)} tracked XBRL filing(s) unresolved within the "
|
||||
f"{MISSING_XBRL_RETRY_DAYS}-day retry window "
|
||||
f"({_reason_counts(blocking)}) — retry: {_missing_detail(blocking)}"
|
||||
)
|
||||
# Malformed companyfacts payloads must fail, not degrade to skipped rows.
|
||||
if staged.invalid_payloads:
|
||||
@@ -241,7 +454,24 @@ class SecFundamentalsImporter:
|
||||
"skipped_filings": len(staged.skipped_filings),
|
||||
"field_issues": len(staged.field_issues),
|
||||
"skipped_non_xbrl": len(staged.skipped_non_xbrl),
|
||||
"missing_xbrl": len(staged.missing_xbrl),
|
||||
"no_xbrl_filings": staged.no_xbrl_filings[:50],
|
||||
"no_xbrl_filings_count": len(staged.no_xbrl_filings),
|
||||
"no_xbrl_ciks": sorted({
|
||||
str(item["cik"])
|
||||
for item in staged.no_xbrl_filings
|
||||
if item.get("cik")
|
||||
}),
|
||||
"missing_xbrl": staged.missing_xbrl[:50],
|
||||
"missing_xbrl_count": len(staged.missing_xbrl),
|
||||
"missing_xbrl_blocking": len(blocking),
|
||||
"recovered_from_coregistrant": staged.recovered[:50],
|
||||
"recovered_count": len(staged.recovered),
|
||||
# Complete compact gate input; detailed audit lists above stay capped.
|
||||
"setup_blocked_ciks": sorted({
|
||||
str(item["cik"])
|
||||
for item in [*staged.missing_xbrl, *staged.no_xbrl_filings]
|
||||
if item.get("cik")
|
||||
}),
|
||||
"invalid_payloads": staged.invalid_payloads,
|
||||
"cik_updates": len(staged.resolved.cik_updates),
|
||||
# differing existing accessions (immutable — kept, reported here)
|
||||
@@ -253,46 +483,280 @@ class SecFundamentalsImporter:
|
||||
summary=summary,
|
||||
source_max_date=self._latest_index_date,
|
||||
messages=messages,
|
||||
# Company-Facts absence is usually publication lag, but can also be a
|
||||
# permanent co-registrant misfile that the daily index did not expose.
|
||||
# Defer quietly at first; the framework warns if promotions stay stale.
|
||||
retryable=(
|
||||
len(messages) == 1
|
||||
and bool(blocking)
|
||||
and all(
|
||||
m.get("reason") in {"not_in_companyfacts", "parser_unusable"}
|
||||
for m in blocking
|
||||
)
|
||||
),
|
||||
deferred_alert_after_days=MISSING_XBRL_RETRY_DAYS,
|
||||
deferred_alert_messages=(
|
||||
[
|
||||
f"{len(aged_out)} tracked SEC filing(s) remain unresolved past "
|
||||
f"the {MISSING_XBRL_RETRY_DAYS}-day retry window. They will "
|
||||
f"enter automatic retry and block affected symbols from setups: "
|
||||
f"{_missing_detail(aged_out)}"
|
||||
]
|
||||
if aged_out
|
||||
else []
|
||||
),
|
||||
)
|
||||
|
||||
async def promote(self, db, staged: StagedFundamentals, run_id: int) -> dict[str, int]:
|
||||
inserted = 0
|
||||
updated = 0
|
||||
# Only accessions whose reconstruction actually changed are rewritten;
|
||||
# an unchanged stored row is left completely alone.
|
||||
changed = {d["accession"] for d in staged.discrepancies} if self.reparse else set()
|
||||
for row in staged.rows:
|
||||
if row.accession in staged.existing_accessions:
|
||||
continue # immutable — keep the original row
|
||||
if row.accession in changed:
|
||||
# Write the FULL column set (_row_values covers _SNAPSHOT_COLS)
|
||||
# so a rewritten row is never half old-parse, half new-parse.
|
||||
# created_at stays at the original insert; import_run_id
|
||||
# attributes the rewrite.
|
||||
values = _row_values(row, run_id)
|
||||
values.pop("created_at", None)
|
||||
await db.execute(
|
||||
update(FundamentalSnapshot)
|
||||
.where(FundamentalSnapshot.accession == row.accession)
|
||||
.values(**values)
|
||||
)
|
||||
updated += 1
|
||||
continue # otherwise immutable — keep the original row
|
||||
stmt = insert_for_session(db, FundamentalSnapshot).values(**_row_values(row, run_id))
|
||||
stmt = stmt.on_conflict_do_nothing(index_elements=["accession"]) # race belt-and-suspenders
|
||||
await db.execute(stmt)
|
||||
inserted += 1
|
||||
|
||||
# Synchronize the retry queue in the snapshot-promotion transaction.
|
||||
existing_gaps = (await db.execute(select(SecFilingGap))).scalars().all()
|
||||
existing_gap_accessions = {gap.accession for gap in existing_gaps}
|
||||
resolved_accessions = {row.accession for row in staged.rows}
|
||||
# A filing now classified non-XBRL can never yield a snapshot and is no
|
||||
# longer a fundamentals completeness gap.
|
||||
resolved_accessions.update(
|
||||
item["accession"] for item in staged.skipped_non_xbrl
|
||||
)
|
||||
queue_resolved = 0
|
||||
if resolved_accessions:
|
||||
result = await db.execute(
|
||||
delete(SecFilingGap).where(
|
||||
SecFilingGap.accession.in_(resolved_accessions)
|
||||
)
|
||||
)
|
||||
queue_resolved = int(result.rowcount or 0)
|
||||
|
||||
now = _now()
|
||||
tolerated = _past_retry_window(staged.missing_xbrl)
|
||||
for gap in tolerated:
|
||||
stmt = insert_for_session(db, SecFilingGap).values(
|
||||
cik=gap["cik"],
|
||||
accession=gap["accession"],
|
||||
form=gap.get("form"),
|
||||
index_date=gap.get("index_date"),
|
||||
reason=gap["reason"],
|
||||
coregistrant_ciks_json=json.dumps(gap.get("coregistrants") or []),
|
||||
first_seen_at=now,
|
||||
last_attempted_at=now,
|
||||
)
|
||||
await db.execute(
|
||||
stmt.on_conflict_do_update(
|
||||
index_elements=["accession"],
|
||||
set_={
|
||||
"cik": stmt.excluded.cik,
|
||||
"form": stmt.excluded.form,
|
||||
"index_date": stmt.excluded.index_date,
|
||||
"reason": stmt.excluded.reason,
|
||||
"coregistrant_ciks_json": stmt.excluded.coregistrant_ciks_json,
|
||||
"last_attempted_at": stmt.excluded.last_attempted_at,
|
||||
},
|
||||
)
|
||||
)
|
||||
|
||||
# Remove gaps made irrelevant by a later valid 10-K/10-Q. Quality reads
|
||||
# already ignore them; physical cleanup keeps the queue small.
|
||||
active_ids = {gap.id for gap in await fundamentals_quality_service.active_gaps(db)}
|
||||
obsolete_ids = {
|
||||
gap.id for gap in existing_gaps
|
||||
if gap.id not in active_ids and gap.accession not in resolved_accessions
|
||||
}
|
||||
if obsolete_ids:
|
||||
result = await db.execute(
|
||||
delete(SecFilingGap).where(SecFilingGap.id.in_(obsolete_ids))
|
||||
)
|
||||
queue_resolved += int(result.rowcount or 0)
|
||||
|
||||
newly_queued = [
|
||||
gap for gap in tolerated
|
||||
if gap["accession"] not in existing_gap_accessions
|
||||
]
|
||||
|
||||
# Warn (in-transaction, so it commits atomically with the promotion) when
|
||||
# any existing accession reconstructed differently — kept immutable.
|
||||
if staged.discrepancies:
|
||||
accns = ", ".join(d["accession"] for d in staged.discrepancies[:10])
|
||||
disposition = (
|
||||
f"REWRITTEN by reparse run {run_id}" if self.reparse else "kept immutable"
|
||||
)
|
||||
db.add(SystemEvent(
|
||||
severity="warning",
|
||||
source="sec_facts",
|
||||
code="snapshot_discrepancy",
|
||||
code="snapshot_reparse" if self.reparse else "snapshot_discrepancy",
|
||||
message=(
|
||||
f"{len(staged.discrepancies)} stored accession(s) reconstructed "
|
||||
f"differently; kept immutable: {accns}"
|
||||
f"differently; {disposition}: {accns}"
|
||||
)[:4000],
|
||||
dedup_key=f"sec_facts:discrepancy:{run_id}",
|
||||
created_at=_now(),
|
||||
))
|
||||
|
||||
# Persistent current gaps get one actionable escalation rather than a
|
||||
# daily warning. The nullable marker makes this durable and noise-free.
|
||||
escalation_cutoff = now - timedelta(days=FILING_GAP_ESCALATE_DAYS)
|
||||
aged_gaps = (
|
||||
await db.execute(
|
||||
select(SecFilingGap).where(
|
||||
SecFilingGap.first_seen_at <= escalation_cutoff,
|
||||
SecFilingGap.escalated_at.is_(None),
|
||||
)
|
||||
)
|
||||
).scalars().all()
|
||||
if aged_gaps:
|
||||
named = ", ".join(
|
||||
f"{gap.cik}/{gap.accession} ({gap.reason})"
|
||||
for gap in aged_gaps[:10]
|
||||
)
|
||||
db.add(SystemEvent(
|
||||
severity="warning",
|
||||
source="sec_facts",
|
||||
code="filing_gap_aged",
|
||||
message=(
|
||||
f"{len(aged_gaps)} SEC filing gap(s) remain unresolved after "
|
||||
f"{FILING_GAP_ESCALATE_DAYS} days; affected setups remain paused. "
|
||||
f"Review the filing/CIK mapping or parser: {named}"
|
||||
)[:4000],
|
||||
dedup_key=f"sec_facts:filing_gap_aged:{run_id}",
|
||||
created_at=now,
|
||||
))
|
||||
await db.execute(
|
||||
update(SecFilingGap)
|
||||
.where(SecFilingGap.id.in_([gap.id for gap in aged_gaps]))
|
||||
.values(escalated_at=now)
|
||||
)
|
||||
|
||||
# Recovered rows are real data from an unexpected place — record where they
|
||||
# came from, so a wrong recovery is auditable rather than invisible.
|
||||
if staged.recovered:
|
||||
named = ", ".join(
|
||||
f"{r['accession']} <- CIK {r['source_cik']}" for r in staged.recovered[:10]
|
||||
)
|
||||
logger.info(
|
||||
"sec_facts: recovered %d filing(s) from co-registrants: %s",
|
||||
len(staged.recovered),
|
||||
named,
|
||||
)
|
||||
|
||||
# One warning when a gap first enters automatic retry. Repeating it every
|
||||
# day adds noise; the queue remains the durable actionable state.
|
||||
if newly_queued:
|
||||
symbols_by_cik: dict[str, list[str]] = defaultdict(list)
|
||||
for symbol, cik in staged.resolved.symbol_to_cik.items():
|
||||
symbols_by_cik[cik10(cik)].append(symbol)
|
||||
named = ", ".join(
|
||||
f"{'/'.join(symbols_by_cik.get(gap['cik'], [])) or gap['cik']}"
|
||||
f"/{gap['accession']}"
|
||||
for gap in newly_queued[:10]
|
||||
)
|
||||
db.add(SystemEvent(
|
||||
severity="warning",
|
||||
source="sec_facts",
|
||||
code="unresolved_filing",
|
||||
message=(
|
||||
f"{len(newly_queued)} filing(s) entered automatic SEC retry. "
|
||||
f"Affected symbols are blocked from new actionable setups until "
|
||||
f"their filing is recovered: {named}"
|
||||
)[:4000],
|
||||
dedup_key=f"sec_facts:unresolved_filing:{run_id}",
|
||||
created_at=_now(),
|
||||
))
|
||||
|
||||
# A new registrant may have no XBRL filing yet. Keep it out of actionable
|
||||
# setups, but log it instead of raising a recurring operator warning.
|
||||
if staged.no_xbrl_filings:
|
||||
named = ", ".join(
|
||||
f"{e['cik']} ({e.get('name') or '?'})" for e in staged.no_xbrl_filings[:10]
|
||||
)
|
||||
logger.info(
|
||||
"sec_facts: %d registrant(s) have no XBRL history yet: %s",
|
||||
len(staged.no_xbrl_filings),
|
||||
named,
|
||||
)
|
||||
|
||||
ticker_counts = await sec_universe.apply_ticker_updates(
|
||||
db, staged.resolved, staged.sic_updates
|
||||
)
|
||||
return {
|
||||
"inserted": inserted,
|
||||
"existing_unchanged": len(staged.existing_accessions),
|
||||
"updated": updated,
|
||||
"existing_unchanged": len(staged.existing_accessions) - updated,
|
||||
"discrepancies": len(staged.discrepancies),
|
||||
"retry_queue_added": len(newly_queued),
|
||||
"retry_queue_resolved": queue_resolved,
|
||||
**ticker_counts,
|
||||
}
|
||||
|
||||
# -- helpers -----------------------------------------------------------
|
||||
|
||||
async def _retry_backlog(
|
||||
self,
|
||||
db,
|
||||
tracked_ciks: set[int],
|
||||
) -> list[dict[str, Any]]:
|
||||
"""Active typed gaps; migration 028 owns historical bootstrap."""
|
||||
if not tracked_ciks:
|
||||
return []
|
||||
tracked = {cik10(cik) for cik in tracked_ciks}
|
||||
candidates: dict[str, dict[str, Any]] = {}
|
||||
|
||||
queued = await fundamentals_quality_service.active_gaps(db, tracked)
|
||||
for gap in queued:
|
||||
try:
|
||||
coregistrants = json.loads(gap.coregistrant_ciks_json or "[]")
|
||||
except (TypeError, ValueError):
|
||||
coregistrants = []
|
||||
candidates[gap.accession] = {
|
||||
"cik": gap.cik,
|
||||
"accession": gap.accession,
|
||||
"form": gap.form,
|
||||
"index_date": gap.index_date,
|
||||
"reason": gap.reason,
|
||||
"coregistrants": coregistrants,
|
||||
"_retry_queue": True,
|
||||
}
|
||||
|
||||
if not candidates:
|
||||
return []
|
||||
resolved = set(
|
||||
(
|
||||
await db.execute(
|
||||
select(FundamentalSnapshot.accession).where(
|
||||
FundamentalSnapshot.accession.in_(list(candidates))
|
||||
)
|
||||
)
|
||||
).scalars().all()
|
||||
)
|
||||
return [
|
||||
item
|
||||
for accession, item in candidates.items()
|
||||
if accession not in resolved
|
||||
]
|
||||
|
||||
async def _last_processed_index_date(self, db) -> date | None:
|
||||
return (
|
||||
await db.execute(
|
||||
@@ -316,9 +780,25 @@ class SecFundamentalsImporter:
|
||||
rows: list[dict[str, Any]] = []
|
||||
day = last_processed + timedelta(days=1)
|
||||
while day <= latest:
|
||||
# Group the whole day first: a combined filing is listed once per
|
||||
# co-registrant CIK, and those sibling CIKs are the only pointer to
|
||||
# where SEC may have put the XBRL (see _recover_from_coregistrant).
|
||||
by_accession: dict[str, list[dict[str, Any]]] = defaultdict(list)
|
||||
for r in await client.daily_index(day):
|
||||
if r["form"] in _XBRL_FORMS and r["cik"] in tracked:
|
||||
rows.append(r)
|
||||
if r["form"] in _XBRL_FORMS:
|
||||
by_accession[r["accession"]].append(r)
|
||||
for accession, group in by_accession.items():
|
||||
filers = {r["cik"] for r in group}
|
||||
tracked_filers = filers & tracked
|
||||
if not tracked_filers:
|
||||
continue
|
||||
siblings = sorted(filers - tracked_filers)
|
||||
if siblings:
|
||||
self._coregistrants[accession] = siblings
|
||||
for r in group:
|
||||
if r["cik"] in tracked_filers:
|
||||
r["index_date"] = day # not hashed (revision uses cik/accession)
|
||||
rows.append(r)
|
||||
day += timedelta(days=1)
|
||||
return rows
|
||||
|
||||
@@ -335,6 +815,30 @@ class SecFundamentalsImporter:
|
||||
).scalars().all()
|
||||
return {int(c) for c in found}
|
||||
|
||||
async def _last_shares_outstanding(self, db, ciks: set[int]) -> dict[str, float]:
|
||||
"""Latest known shares outstanding per tracked issuer — the continuity
|
||||
reference co-registrant recovery is checked against."""
|
||||
if not ciks:
|
||||
return {}
|
||||
rows = (
|
||||
await db.execute(
|
||||
select(FundamentalSnapshot.cik, FundamentalSnapshot.shares_outstanding)
|
||||
.where(
|
||||
FundamentalSnapshot.cik.in_([cik10(c) for c in ciks]),
|
||||
FundamentalSnapshot.shares_outstanding.is_not(None),
|
||||
)
|
||||
# Last write per cik wins, so the sort must be total: an amendment
|
||||
# and its original share a period_end, and an undefined tie there
|
||||
# would make recovery non-deterministic across runs and dialects.
|
||||
.order_by(
|
||||
FundamentalSnapshot.period_end,
|
||||
FundamentalSnapshot.filed_date,
|
||||
FundamentalSnapshot.accession,
|
||||
)
|
||||
)
|
||||
).all()
|
||||
return {cik: float(shares) for cik, shares in rows}
|
||||
|
||||
async def _existing_by_accession(self, db, accessions: list[str]) -> dict[str, FundamentalSnapshot]:
|
||||
if not accessions:
|
||||
return {}
|
||||
@@ -361,6 +865,73 @@ def _companyfacts_structure_error(cf: Any) -> str | None:
|
||||
return None
|
||||
|
||||
|
||||
def _missing(
|
||||
cik: int,
|
||||
row: dict[str, Any],
|
||||
reason: str,
|
||||
today: date,
|
||||
coregistrants: list[int] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""One unresolvable index row, carrying everything needed to look the filing
|
||||
up by hand (EDGAR accession + the index date it was seen on) and to decide
|
||||
whether it is still young enough to be worth blocking on."""
|
||||
index_date = row.get("index_date")
|
||||
age_days = (
|
||||
(today - index_date).days if isinstance(index_date, date) else 0
|
||||
)
|
||||
if row.get("_retry_queue"):
|
||||
age_days = max(age_days, MISSING_XBRL_RETRY_DAYS + 1)
|
||||
return {
|
||||
"cik": cik10(cik),
|
||||
"accession": row["accession"],
|
||||
"form": row.get("form"),
|
||||
"index_date": index_date,
|
||||
# A newly observed row without a date blocks safely. A durable queue row
|
||||
# has already passed the bounded window and is forced aged-out above.
|
||||
"age_days": age_days,
|
||||
"reason": reason,
|
||||
"coregistrants": list(coregistrants or []),
|
||||
}
|
||||
|
||||
|
||||
def _within_retry_window(missing: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
return [m for m in missing if m.get("age_days", 0) <= MISSING_XBRL_RETRY_DAYS]
|
||||
|
||||
|
||||
def _past_retry_window(missing: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
return [m for m in missing if m.get("age_days", 0) > MISSING_XBRL_RETRY_DAYS]
|
||||
|
||||
|
||||
def _shares_continuous(shares: float | None, reference: float | None) -> bool:
|
||||
"""Does a co-registrant-recovered share count look like this issuer's own?
|
||||
|
||||
The failure worth preventing is storing a subsidiary's standalone facts as the
|
||||
parent's. A co-registrant shell holds a token float — Florida Power & Light
|
||||
against NextEra's 2.09bn shares — so any sane band separates them while still
|
||||
tolerating buybacks and issuance. With no history to compare against (a newly
|
||||
tracked issuer) or no share count at all, recovery is refused, not guessed.
|
||||
"""
|
||||
if not shares or not reference:
|
||||
return False
|
||||
return RECOVERY_SHARES_MIN <= shares / reference <= RECOVERY_SHARES_MAX
|
||||
|
||||
|
||||
def _reason_counts(missing: list[dict[str, Any]]) -> str:
|
||||
counts = Counter(m["reason"] for m in missing)
|
||||
return ", ".join(f"{reason}={n}" for reason, n in sorted(counts.items()))
|
||||
|
||||
|
||||
def _missing_detail(missing: list[dict[str, Any]], limit: int = 10) -> str:
|
||||
detail = ", ".join(
|
||||
f"{m['cik']}/{m['accession']} {m.get('form') or '?'} "
|
||||
f"[{m.get('index_date') or '?'}] {m['reason']}"
|
||||
for m in missing[:limit]
|
||||
)
|
||||
if len(missing) > limit:
|
||||
detail += f", +{len(missing) - limit} more"
|
||||
return detail
|
||||
|
||||
|
||||
def _filing_meta(sub: dict[str, Any]) -> tuple[dict[str, FilingMeta], set[str]]:
|
||||
"""(xbrl_meta, nonxbrl_accessions) from a submissions payload. xbrl_meta only
|
||||
includes 10-K/10-Q(/A) filings that are XBRL and have full period metadata."""
|
||||
@@ -395,5 +966,26 @@ def _row_values(row: SnapshotRow, run_id: int) -> dict[str, Any]:
|
||||
|
||||
|
||||
def _diff_fields(row: SnapshotRow, old: FundamentalSnapshot) -> list[str]:
|
||||
"""Source fields where a re-parsed row differs from the stored (immutable) row."""
|
||||
return [col for col in _COMPARE_COLS if getattr(row, col) != getattr(old, col)]
|
||||
"""Source fields where a re-parsed row differs from the stored row."""
|
||||
return [
|
||||
col for col in _COMPARE_COLS
|
||||
if not _same_value(getattr(row, col), getattr(old, col))
|
||||
]
|
||||
|
||||
|
||||
def _same_value(parsed: Any, stored: Any) -> bool:
|
||||
"""Compare a freshly parsed value against its stored round-trip.
|
||||
|
||||
Datetimes need care: every timestamp here is UTC by construction, but
|
||||
``DateTime(timezone=True)`` only preserves tzinfo on Postgres — SQLite hands
|
||||
back a naive value. Comparing representations would report an unchanged row
|
||||
as differing, which would both spam the discrepancy warning and make a
|
||||
reparse rewrite every row it touched. Compare instants instead.
|
||||
"""
|
||||
if isinstance(parsed, datetime) and isinstance(stored, datetime):
|
||||
return _as_utc(parsed) == _as_utc(stored)
|
||||
return parsed == stored
|
||||
|
||||
|
||||
def _as_utc(value: datetime) -> datetime:
|
||||
return value if value.tzinfo is not None else value.replace(tzinfo=timezone.utc)
|
||||
|
||||
@@ -16,6 +16,7 @@ changes on the framework's failure commit). The proposals are applied only in
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import logging
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Iterable
|
||||
@@ -23,11 +24,21 @@ from typing import Iterable
|
||||
from sqlalchemy import select, update
|
||||
|
||||
from app.models.ticker import Ticker
|
||||
from app.services import settings_store
|
||||
from app.services.earnings_alignment import normalise_symbol
|
||||
from app.services.sec_client import SecClient
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# JSON {symbol: cik} pinning a ticker to a specific registrant, overriding
|
||||
# company_tickers.json. Needed when SEC maps a ticker to a successor entity that
|
||||
# has not filed: XOM points at CIK 2115436 "ExxonMobil Holdings Corp" (zero XBRL
|
||||
# filings) while every 10-K/10-Q — including one filed 2026-05-04 — is still under
|
||||
# CIK 34088. Which registrant is the real filer is a judgement about a corporate
|
||||
# event, so it is pinned explicitly rather than guessed. The importer's
|
||||
# `no_xbrl_filings` warning is what tells you a pin is needed.
|
||||
CIK_OVERRIDES_KEY = "sec_cik_overrides"
|
||||
|
||||
|
||||
@dataclass
|
||||
class ResolvedUniverse:
|
||||
@@ -43,6 +54,7 @@ async def resolve_ciks(db, client: SecClient) -> ResolvedUniverse:
|
||||
"""Resolve tracked tickers to CIKs via company_tickers.json. **Read-only** —
|
||||
returns the mapping + proposed `tickers.cik` writes; mutates nothing."""
|
||||
ticker_to_cik = await client.company_tickers()
|
||||
overrides = await cik_overrides(db)
|
||||
rows = (await db.execute(select(Ticker.id, Ticker.symbol, Ticker.cik))).all()
|
||||
|
||||
result = ResolvedUniverse()
|
||||
@@ -50,7 +62,7 @@ async def resolve_ciks(db, client: SecClient) -> ResolvedUniverse:
|
||||
if not symbol:
|
||||
continue
|
||||
sym = normalise_symbol(symbol)
|
||||
cik = ticker_to_cik.get(sym)
|
||||
cik = overrides.get(sym) or ticker_to_cik.get(sym)
|
||||
if cik is None:
|
||||
continue # ADRs / non-SEC issuers — snapshots simply absent
|
||||
result.symbol_to_cik[sym] = cik
|
||||
@@ -65,6 +77,34 @@ async def resolve_ciks(db, client: SecClient) -> ResolvedUniverse:
|
||||
return result
|
||||
|
||||
|
||||
async def cik_overrides(db) -> dict[str, int]:
|
||||
"""Manual ``{symbol: cik}`` pins from ``SystemSetting[CIK_OVERRIDES_KEY]``.
|
||||
|
||||
A malformed setting must never take the importer down, so anything unparseable
|
||||
is logged and ignored — the run then falls back to company_tickers.json.
|
||||
"""
|
||||
raw = await settings_store.get_value(db, CIK_OVERRIDES_KEY)
|
||||
if not raw:
|
||||
return {}
|
||||
try:
|
||||
loaded = json.loads(raw)
|
||||
except (TypeError, ValueError):
|
||||
logger.warning("%s is not valid JSON — ignoring CIK overrides", CIK_OVERRIDES_KEY)
|
||||
return {}
|
||||
if not isinstance(loaded, dict):
|
||||
logger.warning("%s must be a {symbol: cik} object — ignoring", CIK_OVERRIDES_KEY)
|
||||
return {}
|
||||
out: dict[str, int] = {}
|
||||
for symbol, cik in loaded.items():
|
||||
try:
|
||||
out[normalise_symbol(str(symbol))] = int(cik)
|
||||
except (TypeError, ValueError):
|
||||
logger.warning("%s: bad entry %r -> %r — ignoring", CIK_OVERRIDES_KEY, symbol, cik)
|
||||
if out:
|
||||
logger.info("resolve_ciks: %d CIK override(s) applied: %s", len(out), sorted(out))
|
||||
return out
|
||||
|
||||
|
||||
async def fetch_sic_updates(
|
||||
client: SecClient, cik_to_ticker_ids: dict[int, Iterable[int]]
|
||||
) -> list[tuple[int, str | None, str | None]]:
|
||||
|
||||
@@ -13,6 +13,9 @@ APP_GROUP="${APP_GROUP:-deploy}"
|
||||
ENV_FILE="${ENV_FILE:-/opt/signalplatform/.env}"
|
||||
MIN_FREE_GB="${DOLT_MIN_FREE_DISK_GB:-5}"
|
||||
EARNINGS_DIR="${DOLT_DATA_DIR}/${DOLT_EARNINGS_SUBDIR}"
|
||||
DOLT_IDENTITY_NAME="${DOLT_IDENTITY_NAME:-Signal Platform}"
|
||||
DOLT_IDENTITY_EMAIL="${DOLT_IDENTITY_EMAIL:-signal-platform@localhost}"
|
||||
FUNDAMENTALS_PARITY_REPORT_DIR="${FUNDAMENTALS_PARITY_REPORT_DIR:-/var/lib/signal-platform/reports/fundamentals-parity}"
|
||||
|
||||
fail() {
|
||||
echo "ERROR: $*" >&2
|
||||
@@ -25,6 +28,38 @@ version_ok() {
|
||||
grep -Eq "(^|[[:space:]])v?${DOLT_VERSION}([[:space:]]|$)" <<<"$output"
|
||||
}
|
||||
|
||||
as_app_user() {
|
||||
if [[ "$(id -un)" == "$APP_USER" ]]; then
|
||||
"$@"
|
||||
else
|
||||
command -v runuser >/dev/null 2>&1 || fail "runuser is required"
|
||||
runuser -u "$APP_USER" -- "$@"
|
||||
fi
|
||||
}
|
||||
|
||||
repo_command() {
|
||||
(
|
||||
cd "$EARNINGS_DIR"
|
||||
as_app_user "$@"
|
||||
)
|
||||
}
|
||||
|
||||
repo_config_value() {
|
||||
repo_command "$DOLT_BINARY" config --get "$1"
|
||||
}
|
||||
|
||||
configure_identity() {
|
||||
local name email
|
||||
name="$(repo_config_value user.name 2>/dev/null || true)"
|
||||
email="$(repo_config_value user.email 2>/dev/null || true)"
|
||||
if [[ -z "$name" ]]; then
|
||||
repo_command "$DOLT_BINARY" config --local --add user.name "$DOLT_IDENTITY_NAME"
|
||||
fi
|
||||
if [[ -z "$email" ]]; then
|
||||
repo_command "$DOLT_BINARY" config --local --add user.email "$DOLT_IDENTITY_EMAIL"
|
||||
fi
|
||||
}
|
||||
|
||||
check_free_space() {
|
||||
local available_kb
|
||||
available_kb="$(df -Pk "$DOLT_DATA_DIR" | awk 'NR == 2 {print $4}')"
|
||||
@@ -45,9 +80,12 @@ check_env() {
|
||||
|| fail "set DOLT_EARNINGS_SUBDIR=$DOLT_EARNINGS_SUBDIR in $ENV_FILE"
|
||||
grep -Eq '^SEC_USER_AGENT=.*@.*' "$ENV_FILE" \
|
||||
|| fail "SEC_USER_AGENT in $ENV_FILE must contain a real contact email"
|
||||
grep -Fqx "FUNDAMENTALS_PARITY_REPORT_DIR=$FUNDAMENTALS_PARITY_REPORT_DIR" "$ENV_FILE" \
|
||||
|| fail "set FUNDAMENTALS_PARITY_REPORT_DIR=$FUNDAMENTALS_PARITY_REPORT_DIR in $ENV_FILE"
|
||||
}
|
||||
|
||||
check_all() {
|
||||
local identity_name identity_email
|
||||
id "$APP_USER" >/dev/null 2>&1 || fail "missing service user: $APP_USER"
|
||||
[[ -x "$DOLT_BINARY" ]] || fail "missing Dolt binary: $DOLT_BINARY"
|
||||
version_ok || fail "expected Dolt $DOLT_VERSION at $DOLT_BINARY"
|
||||
@@ -62,6 +100,19 @@ check_all() {
|
||||
else
|
||||
fail "run --check as $APP_USER (or install runuser)"
|
||||
fi
|
||||
identity_name="$(repo_config_value user.name 2>/dev/null || true)"
|
||||
identity_email="$(repo_config_value user.email 2>/dev/null || true)"
|
||||
[[ -n "$identity_name" ]] || fail "missing Dolt user.name for $EARNINGS_DIR"
|
||||
[[ -n "$identity_email" ]] || fail "missing Dolt user.email for $EARNINGS_DIR"
|
||||
[[ -d "$FUNDAMENTALS_PARITY_REPORT_DIR" ]] \
|
||||
|| fail "missing parity report directory: $FUNDAMENTALS_PARITY_REPORT_DIR"
|
||||
if [[ "$(id -un)" == "$APP_USER" ]]; then
|
||||
[[ -w "$FUNDAMENTALS_PARITY_REPORT_DIR" ]] \
|
||||
|| fail "parity report directory is not writable by $APP_USER"
|
||||
else
|
||||
runuser -u "$APP_USER" -- test -w "$FUNDAMENTALS_PARITY_REPORT_DIR" \
|
||||
|| fail "parity report directory is not writable by $APP_USER"
|
||||
fi
|
||||
check_free_space
|
||||
check_env
|
||||
echo "OK: Dolt $DOLT_VERSION and earnings clone are provisioned"
|
||||
@@ -88,6 +139,7 @@ fi
|
||||
version_ok || fail "Dolt $DOLT_VERSION installation failed"
|
||||
|
||||
install -d -o "$APP_USER" -g "$APP_GROUP" -m 0750 "$DOLT_DATA_DIR"
|
||||
install -d -o "$APP_USER" -g "$APP_GROUP" -m 0750 "$FUNDAMENTALS_PARITY_REPORT_DIR"
|
||||
check_free_space
|
||||
|
||||
if [[ ! -d "$EARNINGS_DIR/.dolt" ]]; then
|
||||
@@ -97,4 +149,5 @@ if [[ ! -d "$EARNINGS_DIR/.dolt" ]]; then
|
||||
"$DOLT_BINARY" clone post-no-preference/earnings "$EARNINGS_DIR"
|
||||
fi
|
||||
|
||||
configure_identity
|
||||
check_all
|
||||
|
||||
@@ -426,6 +426,10 @@ workstream B — Alpaca remains the price source throughout.
|
||||
API values across the tracked universe, report per-field deltas and resulting
|
||||
fundamental-score/ranking changes, require explicit approval. Definition
|
||||
changes (e.g. TTM vs provider convention) called out, not averaged away.
|
||||
**Status 2026-07-24: the gate has been exercised and the evidence supports
|
||||
approval** — see the handoff section below. Step (c) is implemented behind the
|
||||
default-off `fundamental_data_sec_dolt_cutover_enabled` SystemSetting; the
|
||||
remaining production action is flipping that switch on and observing it.
|
||||
- A6. Remove FMP/Finnhub/Alpha Vantage; keep monitoring + manual fallback.
|
||||
|
||||
**Workstream B (independent, start when wanted):**
|
||||
@@ -477,6 +481,51 @@ workstream B — Alpaca remains the price source throughout.
|
||||
- Score-parity diff reviewed and approved before `fundamental_data` cutover.
|
||||
- Scheduled imports never block the API event loop.
|
||||
|
||||
## Handoff — remaining work after the A5 parity investigation (2026-07-24)
|
||||
|
||||
The 2026-07-23 parity report surfaced coverage gaps and wrong values; a nine-pass
|
||||
investigation traced every one to parser/identity bugs (not source data), fixed them,
|
||||
and reparsed production twice. Full evidence trail:
|
||||
`reports/fundamentals-parity-20260723-findings.md` (root causes, decisions, validation)
|
||||
plus the before/after reports (`fundamentals-parity-20260723T…` / `…20260724T….json`).
|
||||
Post-fix: candidate scores 504 of 511 vs legacy's 507 (gap = PSKY/Q new registrants +
|
||||
FITB, all explained); revenue-growth agreement 0.0038 median abs delta where both exist.
|
||||
Dennis reviewed the evidence 2026-07-24 and directed proceeding to cutover.
|
||||
|
||||
**Task 1 — A5 activation (IMPLEMENTED 2026-07-24; production switch remains).** The
|
||||
post-activation local refresh of `fundamental_data` derives `pe_ratio` and
|
||||
`market_cap` from newest valid snapshots × latest PostgreSQL close, `revenue_growth`
|
||||
from snapshots, `earnings_surprise`/`next_earnings_date` from `earnings_events`; mark
|
||||
affected cached fundamental scores stale; must run identically when SEC is unreachable.
|
||||
It consumes `fundamentals_derivation.derive()` outputs, NOT raw snapshot fields —
|
||||
that path carries the split guard (`ttm_diluted_eps`
|
||||
nulls when contaminated, with `ttm_diluted_eps_caveat`) and the multi-class share
|
||||
fallback (`shares_outstanding` + `shares_outstanding_estimated`). Parity and activation
|
||||
share the same candidate builder. Activation is the explicit
|
||||
`fundamental_data_sec_dolt_cutover_enabled` SystemSetting and defaults off. It is
|
||||
managed by the **Fundamentals data source** card in Admin → Settings; while active,
|
||||
the weekly legacy collector skips itself so it cannot overwrite the SEC/Dolt cache.
|
||||
See `docs/fundamentals-deployment.md` for the production flip and rollback procedure.
|
||||
|
||||
**Task 2 — A6 decommissioning.** After a short observation window: remove
|
||||
FMP/Finnhub/Alpha Vantage providers, config and env keys; keep monitoring + manual
|
||||
fallback. Gated by the acceptance criteria above — especially forward-calendar
|
||||
timeliness from `dolt_earnings` (its `source_max_date` ran ~5 weeks ahead as of
|
||||
2026-07-23, which passes).
|
||||
|
||||
**Known caveats to carry (documented in the findings report, not bugs to fix):**
|
||||
- KLAC-class post-filing splits: P/E wrong until the next 10-Q; undetectable from
|
||||
snapshots. Workstream B's `corporate_actions` table is the natural future fix.
|
||||
- BRK-B: no share count exists anywhere in companyfacts → no market cap, correctly.
|
||||
- FITB: unscored (split guard + no taggable revenue) — the one name that lost its
|
||||
score relative to legacy; composite renormalises.
|
||||
- Share-change guard at 25% nulls P/E for stock-funded M&A too (COF, WAT…);
|
||||
revisit only if the ~3% universe hit-rate proves painful.
|
||||
- `sec_cik_overrides` SystemSetting pins XOM → 34088 (applied in prod); the
|
||||
`no_xbrl_filings` SystemEvent says when a new pin is needed.
|
||||
- After any future parser change, stored rows need `scripts/reparse_fundamentals.py`
|
||||
(dry-run default; `--apply` rewrites) — snapshots are otherwise immutable.
|
||||
|
||||
## Deferred (explicitly, until a concrete need appears)
|
||||
|
||||
- Workstream B itself is deferred relative to A and blocks nothing in A.
|
||||
|
||||
@@ -226,7 +226,12 @@ primary period (safe — a filing's own context is correct for its current perio
|
||||
|
||||
Identifying `User-Agent` with contact email on every request; well under 10 req/s
|
||||
with spacing; exponential backoff on 429; **403 → alert and stop, never
|
||||
retry-loop**. New config: `sec_user_agent`, `sec_request_spacing_seconds`,
|
||||
retry-loop** — with one carved-out exception: `www.sec.gov/Archives` is served
|
||||
from an S3 bucket without a `ListBucket` grant, so an **absent** file 403s with
|
||||
S3's `AccessDenied` XML rather than 404 (every weekend/holiday daily index does
|
||||
this). That one shape is read as "missing"; a real rejection is the WAF's
|
||||
`text/html` "Undeclared Automated Tool" page and still stops the run.
|
||||
New config: `sec_user_agent`, `sec_request_spacing_seconds`,
|
||||
`sec_max_retries`. Keep only the last ~2 fetched artifacts on disk for debugging
|
||||
(reproducibility is the normalized Postgres rows, per the plan).
|
||||
|
||||
|
||||
+135
-11
@@ -1,17 +1,25 @@
|
||||
# Fundamentals production deployment
|
||||
|
||||
This is the one-time production setup for the Dolt earnings and SEC fundamentals
|
||||
imports. Both imports remain shadow inputs until the separate A5 scoring-cutover
|
||||
approval. Do not add OS cron entries: the application scheduler owns both jobs.
|
||||
imports. The A5 scoring cutover was approved on 2026-07-24; the compat-cache write
|
||||
path is still default-off until the explicit production switch below is set. Do
|
||||
not add OS cron entries: the application scheduler owns both jobs.
|
||||
|
||||
## What the deployment adds
|
||||
|
||||
- `Dolt Earnings Import (shadow)` runs daily at 02:30 America/New_York.
|
||||
- `SEC Fundamentals Import (shadow)` runs daily at 04:00 America/New_York.
|
||||
- `SEC Fundamentals Import` runs daily at 04:00 America/New_York. Its local
|
||||
`fundamental_data` refresh runs only when the A5 switch is enabled.
|
||||
- `Fundamentals Parity Report (read-only)` runs daily at 05:30 America/New_York.
|
||||
- Both jobs are visible, toggleable, and manually triggerable in Admin → Jobs.
|
||||
- Cron expressions are editable in Admin → Schedule.
|
||||
- Every attempt is recorded in `data_import_runs`; failures also create a system
|
||||
event. A failed validation does not promote partial data.
|
||||
- An SEC filing still missing after the short publication-lag window enters
|
||||
`sec_filing_gaps`. The daily importer retries it automatically; affected
|
||||
tickers are excluded from actionable setups until a snapshot is recovered or
|
||||
a later valid 10-K/10-Q supersedes the gap. Migration `028` materializes older
|
||||
promoted gaps into this queue once, so setup reads never scan import history.
|
||||
|
||||
The systemd service uses one application worker. The import framework also holds
|
||||
a PostgreSQL advisory lock per source, so an overlapping manual/scheduled run is
|
||||
@@ -28,11 +36,14 @@ DOLT_EARNINGS_SUBDIR=earnings
|
||||
DOLT_MIN_FREE_DISK_GB=5.0
|
||||
SEC_USER_AGENT=signal-platform/1.0 (contact: real-address@example.com)
|
||||
SEC_REQUEST_SPACING_SECONDS=0.2
|
||||
FUNDAMENTALS_PARITY_REPORT_DIR=/var/lib/signal-platform/reports/fundamentals-parity
|
||||
```
|
||||
|
||||
Use a real monitored contact address. Keep at least 5 GB free at the Dolt data
|
||||
path; 8–10 GB gives comfortable growth headroom. The data directory must stay
|
||||
outside `/opt/signalplatform`, because deployments use `rsync --delete` there.
|
||||
The parity-report directory is also persistent and owned by the service user;
|
||||
its small timestamped JSON/CSV bundles form the temporary A5 review trail.
|
||||
|
||||
## One-time provisioning
|
||||
|
||||
@@ -49,8 +60,18 @@ curl -fsS http://127.0.0.1:8998/api/v1/health
|
||||
|
||||
The provisioner is idempotent. It installs the pinned Dolt version, creates the
|
||||
persistent directory as `deploy:deploy`, clones
|
||||
`post-no-preference/earnings`, verifies free space and `.env`, and refuses an
|
||||
unexpected Dolt version. It does not modify PostgreSQL or start an import.
|
||||
`post-no-preference/earnings`, configures a repository-local author identity for
|
||||
`dolt pull`, verifies free space and `.env`, and refuses an unexpected
|
||||
Dolt version. It does not modify PostgreSQL or start an import. The public clone
|
||||
does not require `dolt login`.
|
||||
|
||||
For a server provisioned before the author-identity check was added, repair the
|
||||
existing clone once with:
|
||||
|
||||
```bash
|
||||
sudo -u deploy -H /usr/local/bin/dolt config --global --add user.name "Signal Platform"
|
||||
sudo -u deploy -H /usr/local/bin/dolt config --global --add user.email "signal-platform@localhost"
|
||||
```
|
||||
|
||||
Do not replace the pinned version with `latest`. A future Dolt upgrade should be
|
||||
a reviewed change to `DOLT_VERSION`, followed by the same provision/check flow.
|
||||
@@ -61,13 +82,36 @@ In Admin → Jobs, wait until no other job is running, then:
|
||||
|
||||
1. Trigger **Dolt Earnings Import (shadow)**. Expect `completed` with import
|
||||
status `promoted`; a repeat without an upstream change should report `no_op`.
|
||||
2. Trigger **SEC Fundamentals Import (shadow)**. The first run performs the
|
||||
2. Trigger **SEC Fundamentals Import**. The first run performs the
|
||||
tracked-universe history backfill and can take materially longer than a daily
|
||||
incremental run. Expect `completed` with import status `promoted`.
|
||||
3. Check Admin → System Events. There should be no new import error.
|
||||
4. Confirm the next-run times correspond to 02:30 and 04:00 New York time.
|
||||
5. Open several ticker pages and confirm the fundamentals panel has populated
|
||||
data and still handles partial/missing issuers cleanly.
|
||||
data and still handles partial/missing issuers cleanly. A ticker held by the
|
||||
quality gate should show **New setups paused** with the specific SEC reason.
|
||||
|
||||
## A5 parity observation window
|
||||
|
||||
After both shadow imports are healthy, trigger **Fundamentals Parity Report
|
||||
(read-only)** once in Admin → Jobs. The **A5 Fundamentals Parity** card above
|
||||
the jobs shows the latest coverage/delta summary and provides authenticated JSON
|
||||
and CSV downloads. The canonical server-side bundles are archived at:
|
||||
|
||||
```text
|
||||
/var/lib/signal-platform/reports/fundamentals-parity/
|
||||
```
|
||||
|
||||
The scheduler then generates one report daily at 05:30 New York time, after the
|
||||
02:30 Dolt and 04:00 SEC jobs. Review 5–7 consecutive reports before making the
|
||||
cutover decision. A report never writes `fundamental_data`, dimension/composite
|
||||
scores, rankings, qualification state, or an approval flag. Materiality bands
|
||||
only highlight rows for review; A5 still requires explicit approval.
|
||||
|
||||
Each bundle contains legacy and candidate P/E, revenue growth, and earnings
|
||||
surprise; definition notes; source revisions and price dates; recomputed legacy
|
||||
and candidate fundamental scores; and per-universe fundamental-rank changes.
|
||||
Definition changes remain explicit even when numeric deltas are small.
|
||||
|
||||
Optional database verification:
|
||||
|
||||
@@ -118,16 +162,96 @@ Expect `OK: source lock is busy`. This is the remaining live-PostgreSQL
|
||||
mutual-exclusion check; SQLite unit tests cannot exercise PostgreSQL advisory
|
||||
locks. A second Admin trigger should independently report the job as busy.
|
||||
|
||||
## A5 production activation (approved 2026-07-24)
|
||||
|
||||
The write path is controlled by the SystemSetting
|
||||
`fundamental_data_sec_dolt_cutover_enabled`. An absent value, `false`, or any
|
||||
value other than `true` leaves `fundamental_data` untouched. Before enabling it,
|
||||
confirm the normal PostgreSQL backup containing `fundamental_data` is current.
|
||||
|
||||
In **Admin → Settings → Fundamentals data source**:
|
||||
|
||||
1. Turn on **Use SEC + Dolt for scoring inputs** and accept the confirmation.
|
||||
2. Click **Run refresh now**. The SEC import may be `promoted` or `no_op`; either
|
||||
result runs the local cache refresh.
|
||||
|
||||
The weekly legacy collector is automatically skipped while the switch is on, so
|
||||
it cannot overwrite the activated cache. The switch remains visible even before
|
||||
its SystemSetting row exists because the safe default is off.
|
||||
|
||||
If the Admin UI is unavailable, enable the cutover directly in PostgreSQL:
|
||||
|
||||
```sql
|
||||
INSERT INTO system_settings (key, value, updated_at)
|
||||
VALUES ('fundamental_data_sec_dolt_cutover_enabled', 'true', now())
|
||||
ON CONFLICT (key) DO UPDATE
|
||||
SET value = EXCLUDED.value, updated_at = now();
|
||||
```
|
||||
|
||||
Then trigger **SEC Fundamentals Import** once in Admin → Jobs. Once enabled, the
|
||||
same refresh also runs after an SEC network/validation failure or a source-lock
|
||||
skip, because it reads only PostgreSQL snapshots, earnings events, and closes.
|
||||
The job message appends the cache row count and changed score-input count when
|
||||
the import itself completed successfully.
|
||||
|
||||
Verify the switch and refreshed rows:
|
||||
|
||||
```sql
|
||||
SELECT key, value, updated_at
|
||||
FROM system_settings
|
||||
WHERE key = 'fundamental_data_sec_dolt_cutover_enabled';
|
||||
|
||||
SELECT count(*) AS rows,
|
||||
max(fetched_at) AS refreshed_at,
|
||||
count(pe_ratio) AS pe_available,
|
||||
count(revenue_growth) AS growth_available,
|
||||
count(earnings_surprise) AS surprise_available,
|
||||
count(next_earnings_date) AS next_date_available
|
||||
FROM fundamental_data;
|
||||
|
||||
SELECT dimension, is_stale, count(*)
|
||||
FROM dimension_scores
|
||||
WHERE dimension = 'fundamental'
|
||||
GROUP BY dimension, is_stale;
|
||||
|
||||
SELECT is_stale, count(*)
|
||||
FROM composite_scores
|
||||
GROUP BY is_stale;
|
||||
```
|
||||
|
||||
The first refresh intentionally marks affected fundamental and composite score
|
||||
caches stale. The normal 15:30 near-close scanner recomputes them before using
|
||||
the rankings; until then, reads truthfully expose the stale state. Observe at
|
||||
least several scheduled cycles before A6 removes the legacy providers.
|
||||
|
||||
## Failure and rollback
|
||||
|
||||
- Disable the failing shadow job in Admin → Jobs. This stops scheduled imports
|
||||
without changing existing data or the legacy scoring path.
|
||||
- To stop the A5 cache writes without stopping SEC snapshot ingestion, turn off
|
||||
**Use SEC + Dolt for scoring inputs** in Admin → Settings. If the UI is
|
||||
unavailable, set `fundamental_data_sec_dolt_cutover_enabled` back to `false`
|
||||
with the SQL above (changing only the value). This prevents the next local
|
||||
refresh but does not restore rows already replaced. Restore `fundamental_data`
|
||||
from the pre-cutover database backup, or—before A6—manually run the legacy
|
||||
Fundamental Collector if its provider keys and quota are still available.
|
||||
- Disable a failing source-import job in Admin → Jobs only when ingestion itself
|
||||
must stop. Existing promoted snapshots/events remain available.
|
||||
- Inspect the job runtime, latest `data_import_runs.validation_json`, service
|
||||
logs, and Admin → System Events before retrying.
|
||||
- `unresolved_filing` is emitted once when a filing enters automatic retry. It
|
||||
does not require a server command. If the gap is still current after 14 days,
|
||||
`filing_gap_aged` is emitted once with the CIK, accession, and parser/mapping
|
||||
reason. A later valid 10-K/10-Q retires the gap even when the original SEC
|
||||
accession never becomes usable.
|
||||
- Successful co-registrant recovery is logged without a warning. New registrants
|
||||
with no XBRL history are also logged quietly, but their ticker page explains
|
||||
that setups remain paused and that successor shells may need `sec_cik_overrides`.
|
||||
- Re-run `sudo -u deploy bash ./deploy/provision_fundamentals.sh --check` for
|
||||
binary, clone, permission, disk, or environment failures.
|
||||
- The Dolt clone is a reproducible cache and does not need a bespoke backup.
|
||||
PostgreSQL (including `earnings_events`, `fundamental_snapshots`, and import
|
||||
audit rows) must remain covered by the normal production database backup.
|
||||
- Do not proceed to A5 while either shadow feed is unhealthy or the parity gate
|
||||
has not received explicit approval.
|
||||
- Do not proceed to A6 until the activated cache has completed the observation
|
||||
window and the forward earnings calendar remains timely.
|
||||
- If report generation fails, inspect Admin → System Events and verify
|
||||
`FUNDAMENTALS_PARITY_REPORT_DIR` exists and is writable by `deploy`. Existing
|
||||
reports and all live data remain untouched.
|
||||
|
||||
+14
-2
@@ -25,7 +25,7 @@ score, Structural S/R, the Gate Target Ladder, sentiment, fundamentals) is
|
||||
| 1.5× ATR initial stop | Real exit | Cuts losers fast |
|
||||
| 3× ATR trailing stop, 30-day max hold | Real exit | Best Sharpe of every exit tested |
|
||||
| Post-stop normal gate reset | Re-entry policy | Stop always closes; a later gate failure and subsequent fresh qualification define the next signal episode. The selected study arm reached Sharpe 1.77 / CAGR 48.3% at capacity 10; live scan-before-outcome timing is stricter (Sharpe 1.68 / CAGR 44.8% analogue). [Full study](post-stop-reentry.md) |
|
||||
| Max 10 concurrent positions, 1% risk per trade | Sizing | Cap never binds in practice |
|
||||
| Max 10 concurrent positions, 1% risk per trade | Sizing | The cap binds by signal count, but the focused bracket found negligible opportunity cost: cap 15 admitted every blocked setup and added only 0.0018 R/trade in affected paths. [Findings](portfolio-capacity-bracket-findings.md) |
|
||||
| Structural S/R | Human-facing product context | Clean, capped zones for charts and alerts; not read by the scanner |
|
||||
| Gate Target Ladder | Screening machinery | Volume-free transient proposals preserve the production candidate set exactly; never an exit |
|
||||
|
||||
@@ -61,7 +61,7 @@ invites overfitting.
|
||||
|---|---|
|
||||
| ATR trail multiple {1.5–4.0} | **Keep 3.0** — ≤2.0 whipsaws out the right tail; ≥2.5 is a plateau |
|
||||
| Momentum lookback (6-1, 3-1, 12-7 Novy-Marx, composites) | **Keep residual 12-1** — the others have IC ≈ 0 or weaker t-stats |
|
||||
| Selection cutoff {70…90} × book size {10, 15, 20} | **Keep 80 × 10** — monotonically worse in both directions |
|
||||
| Selection cutoff {70…90} × book size {10, 15, 20} | **Keep 80 × 10** — the focused daily bracket found no meaningful gain from cap 15, while weekly rank replacement hurt. [Findings](portfolio-capacity-bracket-findings.md) |
|
||||
| Position sizing (equal-weight, inverse-vol, risk-% sweep) | **Keep 1% fixed-fractional** |
|
||||
| Primary-target probability floor | **Keep 20%** — pruned lottery targets, 1,428 → 1,089 qualified, lifted Sharpe |
|
||||
| Primary-target R:R selector | **Keep 1.5** — target choice is intentionally independent of the later 2.0 activation floor |
|
||||
@@ -146,6 +146,7 @@ knobs.
|
||||
| **Broader universe** | Composition changes factor signs (fip tug-of-war); vol-tilt on breadth is only a **directional hypothesis** (auth. −0.048 / t −1.36) | Any prod broaden must re-validate 80/20 tilt; offline research only; research.sqlite requires completion manifest |
|
||||
| **Forward paper-trade record** | The only true out-of-sample evidence the snapshot cannot give | Time; mark entries at actual near-close fill once ops ships |
|
||||
| **Better target model for clear-air names** | The return is demonstrably there (#2 wins on raw CAGR in *both* train and test); it's the *flat* 3× ATR target that makes it too expensive in risk | Needs a per-name model, not a constant k×ATR |
|
||||
| **Minimum effective-risk floor** | In cap-never-bound paths, the confounded 0.5% floor arm removed about 8% of fills while EV rose from 0.328 to 0.399 R and PF from 1.60 to 1.75, with exposure nearly unchanged | Run the frozen single-variable cap-10 A/B. [Specification](effective-risk-floor-ab.md) / [capacity findings](portfolio-capacity-bracket-findings.md) |
|
||||
|
||||
---
|
||||
|
||||
@@ -197,4 +198,15 @@ qualification. The [daily re-entry matrix](post-stop-reentry.md) supports this
|
||||
for the current 10-position book, but not as a universal rule for other
|
||||
portfolio capacities.
|
||||
|
||||
Capacity is now closed as a negative result. The current daily Phase A control
|
||||
does reject 519 qualified entries because the ten-slot book is full versus 472
|
||||
admitted trades, so the older weekly “cap never binds” claim was stale. But the
|
||||
clean cap-15 arm admitted every opportunity the strategy requested and added
|
||||
only 0.0018 R/trade in paths where cap 10 bound. Weekly current-rank replacement
|
||||
reduced mean EV and created substantial churn. Keep cap 10 and do not build the
|
||||
replacement policy. See the [frozen specification](portfolio-capacity-bracket.md)
|
||||
and the separate [capacity findings](portfolio-capacity-bracket-findings.md).
|
||||
The only open follow-up from that run is the
|
||||
[frozen confound-free 0.5% minimum effective-risk-floor A/B](effective-risk-floor-ab.md).
|
||||
|
||||
The next real evidence is **forward**, not backward: the live paper-trade record.
|
||||
|
||||
@@ -0,0 +1,124 @@
|
||||
# Effective initial-risk floor A/B - frozen specification
|
||||
|
||||
Date frozen: 2026-08-05
|
||||
Branch: research/portfolio-capacity-rebalancing
|
||||
Runner: scripts/run_portfolio_construction_matrix.py
|
||||
Study ID: risk-floor-ab
|
||||
|
||||
## Question
|
||||
|
||||
Does rejecting an otherwise qualified cap-10 entry when its actual initial
|
||||
stop-risk after cash and notional sizing is below 0.5% of marked equity improve
|
||||
trade selection?
|
||||
|
||||
The completed capacity bracket cannot answer this. Its cash_unbounded arm
|
||||
removed the count cap and applied the 0.5% floor simultaneously. In the 70 paths
|
||||
where the control cap never bound, that arm still raised mean EV from 0.328 to
|
||||
0.399 R and profit factor from 1.60 to 1.75 while trades fell about 8% and
|
||||
exposure stayed nearly flat. Capacity was a no-op in those paths, so the floor
|
||||
is the plausible cause, but the prior arm remains confounded.
|
||||
|
||||
This A/B changes only the floor. It has no formal promotion gate and does not
|
||||
automatically change production.
|
||||
|
||||
## Frozen arms
|
||||
|
||||
1. cap10_incumbent: current production-style cap-10 control, with no minimum
|
||||
effective-risk floor.
|
||||
2. cap10_min_risk_005: the same cap-10 strategy, rejecting an entry only when
|
||||
actual initial stop-risk after cash/notional sizing is below 0.5% of marked
|
||||
equity.
|
||||
|
||||
Both arms have max_positions=10, weekly replacement disabled, 1% target risk
|
||||
per trade, and identical admission ordering. The only differing simulator
|
||||
argument is min_initial_risk_fraction: None versus 0.005.
|
||||
|
||||
All other settings remain the frozen daily Phase A control: current production
|
||||
construction universe, full-universe residual-momentum/low-volatility 80/20
|
||||
rank, threshold 80, normal gate-reset re-entry, close fills, 3x ATR trail,
|
||||
30-session maximum hold, 20% per-position notional ceiling, no leverage, and
|
||||
costs of 0.10% and 0.20% per fill.
|
||||
|
||||
Every priced symbol contributes to the daily cross-sectional rank. Rank-only
|
||||
symbols cannot submit trades. Validation retains the 450-600-symbol production
|
||||
construction guardrail and the legacy-snapshot column-scoped loader.
|
||||
|
||||
## Frozen cohorts
|
||||
|
||||
Reuse the completed bracket's point-in-time daily candidate/rank cache and
|
||||
cohort manifest:
|
||||
|
||||
- Empty book: first eligible session of each month in 2019-2025, with 504 prior
|
||||
scoring sessions and 252 measurement sessions. This is the primary start-date
|
||||
evidence.
|
||||
- Warm book: weekly seeds 63-126 sessions before each 2019-2025 annual anchor,
|
||||
with state carried into the same 252-session measurement window. This is a
|
||||
state-carrying replication, not independent evidence.
|
||||
|
||||
The expected realization is 78 empty-book paths, 97 warm paths, seven annual
|
||||
clusters in each protocol, two costs, two arms, and 700 cells.
|
||||
|
||||
Do not use warm-seed IQR as evidence. Six of seven completed-bracket anchors
|
||||
were structurally degenerate because fractional sizing is scale invariant and
|
||||
the 30-session maximum hold washed out books before anchors. The 2023 exception
|
||||
shows that state carrying itself works.
|
||||
|
||||
## Reporting and interpretation
|
||||
|
||||
For every protocol and cost, pair identical paths. Report:
|
||||
|
||||
- mean, median, P25, and P75 paired net-EV changes in R;
|
||||
- positive-path and bit-identical-path fractions;
|
||||
- the median paired delta within each year and the median across seven years;
|
||||
- simple 90% cluster-bootstrap context for EV and Calmar, with no CI gate;
|
||||
- mean paired PF, Gain-to-Pain, Sortino, Calmar/MAR, CAGR, maximum drawdown,
|
||||
total return, and Sharpe changes;
|
||||
- trades, floor rejections, holding time, cash, gross exposure, average/peak
|
||||
positions, turnover, and costs.
|
||||
|
||||
Means and identical-path fractions must appear beside medians so inert cohorts
|
||||
cannot turn a left- or right-skewed treatment into a misleading zero headline.
|
||||
For these 252-session windows, the implementation's full-window Calmar is CAGR
|
||||
divided by maximum drawdown, the same numeric definition commonly called MAR;
|
||||
do not present the duplicate label as a second independent metric.
|
||||
|
||||
Today's production membership is projected backward. Use paired differences
|
||||
for the treatment conclusion; absolute profitability remains descriptive and
|
||||
survivorship-biased. Empty and warm protocols cover the same seven market years
|
||||
and must not be interpreted as independent replications.
|
||||
|
||||
Interpretation is deliberately simple:
|
||||
|
||||
- a positive result means the isolated floor improves the paired EV
|
||||
distribution without an economically important loss of total-return or
|
||||
drawdown quality;
|
||||
- a negative result closes the floor;
|
||||
- mixed EV/portfolio-quality results are reported as a trade-off, not forced
|
||||
through a composite score.
|
||||
|
||||
## Reproducibility and macOS execution
|
||||
|
||||
The authoritative run refuses a dirty worktree. Its fingerprint includes the
|
||||
implementation commit, this specification hash, snapshot hash, candidate-cache
|
||||
key, construction view, cohort manifest, arm definitions, costs, and study
|
||||
version. Cells checkpoint atomically and --resume verifies the fingerprint.
|
||||
|
||||
From the repository root on macOS:
|
||||
|
||||
python3 -m venv .venv
|
||||
./.venv/bin/python -m pip install -e '.[dev]'
|
||||
|
||||
Preflight, reusing the completed bracket's candidate/rank cache:
|
||||
|
||||
./.venv/bin/python scripts/run_portfolio_construction_matrix.py + backtest_snapshots/research.sqlite + --study risk-floor-ab + --run-id prod505-effective-risk-floor-ab-daily-v1 + --candidate-cache reports/.cache/prod505-capacity-bracket-daily-v1-candidates.pkl + --workers 8 + --resume + --validate-only
|
||||
|
||||
Authoritative run:
|
||||
|
||||
./.venv/bin/python scripts/run_portfolio_construction_matrix.py + backtest_snapshots/research.sqlite + --study risk-floor-ab + --run-id prod505-effective-risk-floor-ab-daily-v1 + --candidate-cache reports/.cache/prod505-capacity-bracket-daily-v1-candidates.pkl + --workers 8 + --resume
|
||||
|
||||
On an M2 Pro, eight workers is the explicit high-utilization setting. Use six
|
||||
instead on a memory-constrained machine; auto intentionally caps itself at six.
|
||||
Changing worker count does not change the fingerprint or results.
|
||||
|
||||
Commit only the compact final JSON and Markdown reports. Candidate caches,
|
||||
checkpoints, raw curves, and trade ledgers remain ignored.
|
||||
@@ -0,0 +1,74 @@
|
||||
# Fundamentals ranking-overlay research
|
||||
|
||||
Status: completed 2026-07-23. Decision: keep production scoring and qualification unchanged.
|
||||
|
||||
## Question
|
||||
|
||||
Does using point-in-time SEC fundamentals to reorder already-qualified long setups improve the production portfolio's risk-adjusted return? The experiments changed ranking only; qualification, execution, sizing, capacity, costs, ATR exits, and post-stop re-entry remained unchanged.
|
||||
|
||||
## Method
|
||||
|
||||
- The control was the production 80/20 residual-momentum / volatility rank.
|
||||
- SEC facts became visible only after `accepted_at`, using the newest visible accession per fiscal period.
|
||||
- Portfolio simulations used daily entry opportunities, close fills, the production gate-reset re-entry policy, and a 30-session horizon.
|
||||
- Train contained entries before 2024-01-01, validation covered 2024, and test began 2025-01-01.
|
||||
- Missing composite scores were neutral at 50.
|
||||
- Deflated Sharpe used the complete registered arm count for each experiment.
|
||||
|
||||
The snapshot contained 511 tracked tickers, 507 unique CIKs, 30,494 SEC snapshot rows, and prices from 2021-06-24 through 2026-07-22.
|
||||
|
||||
## Initial experiment
|
||||
|
||||
The first registered matrix tested quality, growth, and balanced composites at 10%, 20%, 30%, and 40% weights: 13 trials including control.
|
||||
|
||||
No overlay passed the train and validation requirements. The most attractive full-period result, balanced at 10%, failed validation and improved test Sharpe by only 0.06.
|
||||
|
||||
Review also found that filing-time diluted EPS and shares are not reliably comparable across stock splits. A snapshot audit found share-count changes above 25% for 90 of 461 issuers with comparable 2021+ periods, including recognizable split ratios for AMZN, GOOG, NVDA, CMG, and GE plus some obvious unit anomalies. Consequently, EPS growth and share-count change cannot be trusted for historical ranking without point-in-time split factors.
|
||||
|
||||
The complete initial result is recoverable from Git commit `7f944d7`.
|
||||
|
||||
## Split-safe follow-up
|
||||
|
||||
The follow-up excluded diluted-EPS growth and share-count change completely. It tested:
|
||||
|
||||
- Quality: operating margin, FCF margin, and low net-debt/EBITDA, requiring at least two inputs.
|
||||
- Growth: revenue growth only.
|
||||
- Balanced: equal quality and growth weights.
|
||||
- Overlay weights: 5%, 10%, and 15%.
|
||||
|
||||
This produced 10 registered trials including control. Growth coverage among qualified candidates was 88.96%; lack of data was not the limiting factor.
|
||||
|
||||
| Window | Control Sharpe | Revenue-growth 5% | Delta |
|
||||
|---|---:|---:|---:|
|
||||
| Train | 1.26 | 1.31 | +0.05 |
|
||||
| Validation | 2.52 | 2.64 | +0.12 |
|
||||
| Test | 1.99 | 1.99 | 0.00 |
|
||||
| Full | 1.86 | 1.87 | +0.01 |
|
||||
|
||||
Revenue growth at 5% mechanically passed the deliberately permissive "not worse" gate, but did not demonstrate an economically meaningful edge:
|
||||
|
||||
- Test CAGR rose from 54.4% to 56.1%, while full-period CAGR fell from 52.1% to 51.5%.
|
||||
- Full-period trade overlap was 68.53%, so roughly one-third of selections changed for essentially unchanged Sharpe.
|
||||
- Revenue-growth IC was 0.0006 in train, 0.0053 in test, and 0.0116 full-period with a full-period t-stat of 0.55.
|
||||
- Growth weights of 10% and 15% deteriorated; quality and balanced composites failed.
|
||||
- The test window had already been inspected, so this follow-up was sensitivity evidence rather than a fresh out-of-sample result.
|
||||
|
||||
The complete split-safe result is recoverable from Git commit `dba7ea7`.
|
||||
|
||||
## Decision
|
||||
|
||||
- Do not add fundamental weight to production ranking or the automated qualification gate.
|
||||
- Do not run another historical weight sweep on the same sample; it would add data-mining rather than new evidence.
|
||||
- Keep fundamentals informational and user-facing in the UI.
|
||||
- A5 source-parity and cutover work can proceed independently without changing scoring behavior.
|
||||
- Treat historical EPS growth and share-count change as non-comparable across corporate actions until a split-aware solution or a conservative UI guard exists.
|
||||
|
||||
Revisit automated weighting only with materially better data, such as point-in-time split factors and historical constituent/delisting coverage, followed by genuinely new forward paper evidence.
|
||||
|
||||
## Limitations
|
||||
|
||||
The snapshot uses today's tracked universe rather than historical membership and delisted securities, creating survivorship bias. Absolute CAGR and Sharpe must not be interpreted as unbiased live expectations. The relative comparison is useful, but the observed test window and short number of independent factor windows limit statistical power.
|
||||
|
||||
## Repository cleanup
|
||||
|
||||
The experiment-only scorer, runner, Mac launcher, caches, tests, and expanded report bundles were removed after this decision. They remain recoverable from commits `eae4d34`, `34d6dda`, `7f944d7`, and `dba7ea7`. Production fundamentals derivation and ingestion remain unchanged.
|
||||
@@ -28,6 +28,13 @@ Mechanics guards confirmed before reading results: calendar truncation asserted
|
||||
| **Validation** | **1.68** | **0.72** | **41.6%** | **20.9%** | **1.99** | **239** |
|
||||
| Full (close-fill) | 1.77 | 0.50 | 48.3% | 21.6% | 2.23 | 472 |
|
||||
|
||||
**Capacity correction (2026-08-05):** the full close-fill control also records
|
||||
skipped_book_full = 519 versus 472 admitted trades, so the ten-slot book
|
||||
refuses 52.4% of admitted+blocked qualified opportunities. The older weekly
|
||||
claim that the cap never bound is stale and does not apply to this daily
|
||||
gate-reset configuration. Capacity is now isolated in the
|
||||
[focused bracket study](portfolio-capacity-bracket.md).
|
||||
|
||||
Validation SE ≈ 0.72 — almost no arm clears a 1-SE delta.
|
||||
|
||||
---
|
||||
|
||||
@@ -0,0 +1,124 @@
|
||||
# Portfolio-capacity bracket — findings
|
||||
|
||||
Date interpreted: 2026-08-05
|
||||
|
||||
Status: **capacity and weekly replacement closed as negative results; the
|
||||
minimum effective-risk floor remains an open single-variable follow-up.**
|
||||
|
||||
This document interprets the frozen v2 run without modifying its generated
|
||||
outputs:
|
||||
|
||||
- result commit: `24482c6`;
|
||||
- simulation source commit: `6fc82ae8574de9104c83273e018391e75a5f8ac6`;
|
||||
- frozen specification SHA-256:
|
||||
`f1e37783cf6d157ecc827d48211fa45da16f0a0ac19cd23686b3902d347a1898`;
|
||||
- JSON SHA-256:
|
||||
`2435875667097db7416a0d96f412db81d2f2d09ba053748c9f2cfb8a0cba4417`;
|
||||
- Markdown SHA-256:
|
||||
`dc3f5de25eb0a156ce51d0025c90e04ac0977e9502dec47bcf1b25bdcf609c81`.
|
||||
|
||||
The run completed 78 empty-book paths, 97 warm-seed paths, seven annual
|
||||
clusters under both protocols, two cost levels, four arms, and 1,400 cells with
|
||||
no validation errors. The construction universe was 505 priced tradable
|
||||
symbols plus 4,149 priced rank-only symbols.
|
||||
|
||||
## Capacity is economically free
|
||||
|
||||
The clean capacity treatment is `cap15_incumbent`: it changes no sizing or
|
||||
admission rule. Its cap never bound in any cell (maximum observed position count
|
||||
12; zero full-book skips), so it absorbed every opportunity blocked by cap 10.
|
||||
|
||||
At 0.10% per fill, split the 175 paths by whether the paired control recorded
|
||||
any `skipped_book_full`. Values below are mean paired changes in net EV per
|
||||
trade, in R:
|
||||
|
||||
| Arm | Cap never bound (n=70) | Cap did bind (n=105) |
|
||||
|---|---:|---:|
|
||||
| `cap15_incumbent` | +0.0000 | +0.0018 |
|
||||
| `cash_unbounded` | +0.0714 | +0.0077 |
|
||||
| `cap10_weekly_top10` | -0.0246 | -0.0426 |
|
||||
|
||||
The exact zero for cap15 in the never-bound stratum is also a harness validity
|
||||
check: when the treatment cannot act, results are identical. Where it does act,
|
||||
giving the strategy every slot it requested adds only 0.0018 R/trade. The old
|
||||
519-blocked-versus-472-admitted count was true, but it did not imply that the
|
||||
blocked opportunities were economically valuable.
|
||||
|
||||
Decision: **keep the production cap at 10.** Do not remove it or raise it in the
|
||||
expectation of additional edge.
|
||||
|
||||
## The positive arm measured the risk floor
|
||||
|
||||
`cash_unbounded` combined two treatments: no count cap and a 0.5% minimum
|
||||
effective initial-risk fraction. Its EV effect is roughly nine times larger in
|
||||
the 70 paths where the control cap never bound, so capacity cannot explain the
|
||||
improvement.
|
||||
|
||||
Within that never-bound stratum:
|
||||
|
||||
| Measure | Control | `cash_unbounded` |
|
||||
|---|---:|---:|
|
||||
| Mean trades | 75.7 | 69.9 |
|
||||
| Mean cash | 27.8% | 28.2% |
|
||||
| Mean gross exposure | 72.2% | 71.8% |
|
||||
| Mean hold | 15.4 sessions | 15.6 sessions |
|
||||
| Mean EV | +0.328 R | +0.399 R |
|
||||
| Mean profit factor | 1.60 | 1.75 |
|
||||
|
||||
The floor removes about 8% of fills while leaving exposure and holding time
|
||||
nearly unchanged. This is selection, not general de-risking: candidates that
|
||||
available sizing compresses below half the intended risk are worse on average.
|
||||
The report records repeated reject attempts, not the rejected candidates'
|
||||
ranks, so whether the effect is rank-mediated remains unknown.
|
||||
|
||||
Next research: one single-variable A/B, `cap10_incumbent` versus cap 10 with
|
||||
`min_initial_risk_fraction=0.005`, with every other rule unchanged. Do not call
|
||||
the current `cash_unbounded` result causal evidence for that floor until this
|
||||
confound-free comparison is run.
|
||||
|
||||
## Weekly replacement hurts
|
||||
|
||||
Median paired deltas read zero because enough cohorts are inert. The distribution
|
||||
is not neutral:
|
||||
|
||||
| Protocol | Mean ΔEV | P25 ΔEV | Identical paths |
|
||||
|---|---:|---:|---:|
|
||||
| Empty book | -0.0360 R | -0.0817 R | 27/78 (34.6%) |
|
||||
| Warm book | -0.0348 R | -0.1582 R | 14/97 (14.4%) |
|
||||
|
||||
The arm made 2,170 replacements and 529 same-symbol re-entries within ten
|
||||
sessions, so 24% of replacements were associated with short-horizon churn.
|
||||
|
||||
Decision: **reject weekly top-10 replacement.** Future reports should show mean
|
||||
paired effects and identical-path fractions beside medians whenever treatments
|
||||
are inert in a material share of cohorts.
|
||||
|
||||
## Warm dispersion was mostly structurally degenerate
|
||||
|
||||
For six of seven anchors, control EV IQR is numerical zero (approximately
|
||||
`1e-16`) and Calmar IQR is exactly zero. The displayed ratio `1.000` is therefore
|
||||
mostly the implementation's zero-over-zero convention, not evidence of equal
|
||||
nonzero dispersion.
|
||||
|
||||
Two mechanics cause convergence: sizing and notional limits are fractions of
|
||||
equity, making R and ratio metrics scale-invariant; and the 30-session maximum
|
||||
hold is shorter than the 63-session minimum seed offset, allowing initial books
|
||||
to wash out before the anchor.
|
||||
|
||||
The exception is 2023. Control measurement-start positions vary from 6 to 9,
|
||||
EV IQR is 0.0274 R, and Calmar IQR is 0.2675. The protocol therefore carries
|
||||
state correctly, but its chosen offsets usually erase the initialization effect
|
||||
it was intended to measure.
|
||||
|
||||
Future initialization studies should use seed offsets shorter than maximum hold,
|
||||
approximately 5–25 sessions. The current empty-book cohorts remain the primary
|
||||
start-date evidence, but they necessarily mix initialization with market regime.
|
||||
|
||||
## Final decisions
|
||||
|
||||
1. Keep cap 10; its measured opportunity cost is negligible.
|
||||
2. Reject weekly rank replacement.
|
||||
3. Do not interpret the `cash_unbounded` improvement as a capacity effect.
|
||||
4. Run only the focused cap-10 effective-risk-floor A/B next.
|
||||
5. Report means, inert fractions, and absolute dispersion beside medians and
|
||||
ratios in future sparse-treatment studies.
|
||||
@@ -0,0 +1,169 @@
|
||||
# Portfolio-capacity bracket — frozen specification
|
||||
|
||||
Date frozen: 2026-08-05
|
||||
Branch: research/portfolio-capacity-rebalancing
|
||||
Runner: scripts/run_portfolio_construction_matrix.py
|
||||
|
||||
## Question and motivation
|
||||
|
||||
The daily Phase A production control (a0_control: close fill, 30-session
|
||||
maximum hold, 1% fixed-fractional risk, no correlation or volatility overlay)
|
||||
recorded 472 trades and 519 otherwise qualified entries rejected because the
|
||||
ten-position book was full. The blocked share is 519 / (519 + 472) = 52.4%.
|
||||
The book is therefore materially arrival-order constrained.
|
||||
|
||||
This supersedes the older statement that the ten-slot cap never bound. That
|
||||
statement came from a shorter, weekly, pre-gate-reset replay and is not evidence
|
||||
about the current daily strategy.
|
||||
|
||||
The study brackets the value of capacity before tuning replacement details. It
|
||||
does not contain a formal promotion rule or automatically change production.
|
||||
Because the current ~505-name production membership is projected backward,
|
||||
paired arm-versus-control differences are the primary evidence. Absolute
|
||||
profitability is descriptive and survivorship-biased.
|
||||
|
||||
Implementation correction: the first completed v1 artifact at commit `23fe39f`
|
||||
incorrectly allowed the snapshot's broad rank-only universe to submit trades.
|
||||
That artifact is invalid, is removed from the branch, and must not be used for
|
||||
strategy conclusions. Runner v2 fixes the construction/ranking partition below.
|
||||
|
||||
## Frozen arms
|
||||
|
||||
1. **cap10_incumbent:** exact production-style cap-10 control, no displacement.
|
||||
2. **cash_unbounded:** no position-count cap; cash/no leverage and the existing
|
||||
20% per-position notional ceiling remain. Reject an entry if actual initial
|
||||
stop-risk after cash/notional sizing is below 0.5% of marked equity.
|
||||
3. **cap10_weekly_top10:** on the final trading session of each ISO week, rank
|
||||
holdings plus fresh same-day qualified entrants and retain the top ten.
|
||||
4. **cap15_incumbent:** cap 15, no displacement.
|
||||
|
||||
All arms use the frozen Phase A control configuration: daily candidate replay,
|
||||
live-like full-universe residual-momentum/low-volatility 80/20 rank, activation
|
||||
threshold 80, normal gate-reset re-entry, close fill, 3×ATR trail, 30-session
|
||||
maximum hold, 1% risk, and costs of 0.10% and 0.20% per fill.
|
||||
|
||||
Every priced symbol contributes to the daily cross-sectional rank. Only symbols
|
||||
not listed in the snapshot's `research_rank_only` side table may submit trade
|
||||
setups to any arm. The resulting construction universe must contain 450-600
|
||||
symbols (expected approximately 505); validation fails outside that frozen
|
||||
guardrail or when the side table references unknown ticker symbols.
|
||||
|
||||
The daily replay uses zero outcome horizon: setup and rank observations continue
|
||||
through the snapshot's last session because portfolio simulation, unlike outcome
|
||||
grading, does not require 30 future bars.
|
||||
|
||||
Control-parity note: a direct main-versus-branch comparison found identical
|
||||
total return, CAGR, maximum drawdown, and Sharpe. The branch intentionally
|
||||
changes only the first calendar year's `yearly_returns` convention: it starts
|
||||
from initial capital rather than equity after the first session, so day-one
|
||||
entry costs are now charged to year one. Older reports can therefore show a
|
||||
different first-year contextual return without a strategy-performance
|
||||
regression. New trade-detail and measurement-start fields are additive.
|
||||
|
||||
### Weekly-selection mechanics
|
||||
|
||||
- Ordinary exits run before entries/rebalancing.
|
||||
- Open slots may still fill from daily qualified entries during the week.
|
||||
- On the final ISO-week session, current holdings and that day's fresh qualified
|
||||
entrants use the full-universe strategy_rank for that same date.
|
||||
- Stored entry-day rank is never used.
|
||||
- Holdings with missing current rank/data are protected and consume a slot;
|
||||
entrants missing rank are ineligible.
|
||||
- Incumbents win exact rank ties; symbol is the deterministic final tie-breaker.
|
||||
- Rebalance exits pay costs and bypass cooldown/post-stop state.
|
||||
- Report entrant-pool sizes, replacements, turnover, and same-symbol re-entry
|
||||
within 5/10/20 sessions.
|
||||
|
||||
## Frozen cohorts
|
||||
|
||||
research.sqlite is expected to cover 2016-01-04 through 2026-07-17. Residual
|
||||
momentum requires 252 benchmark sessions. Empty-book starts additionally require
|
||||
504 prior scoring sessions and 252 forward measurement sessions.
|
||||
|
||||
- **Empty book:** first eligible session of each month, approximately January
|
||||
2019 through July 2025; start with no positions and measure 252 sessions.
|
||||
- **Warm book:** first session of each year 2019–2025 is the measurement anchor.
|
||||
Seed the portfolio on the first session of every ISO week falling 63–126
|
||||
trading sessions before the anchor, carry all positions and gate-reset state
|
||||
forward, and measure the same 252-session anchor window.
|
||||
|
||||
Warm portfolio returns reset to marked equity immediately before the anchor
|
||||
session. P&L after the anchor from carried positions belongs to portfolio
|
||||
returns, while trade EV includes only entries on or after the anchor. Remaining
|
||||
positions liquidate at the last measurement close with costs.
|
||||
|
||||
The validate-only mode must print realized cohort counts and fail unless both
|
||||
protocols contain the seven annual clusters 2019–2025 and every warm anchor has
|
||||
at least 12 seeds. It must also print ranking, rank-only, and tradable symbol
|
||||
counts plus the raw, removed, and retained qualified-long counts.
|
||||
|
||||
## Reporting
|
||||
|
||||
Primary reported measures:
|
||||
|
||||
- net EV per trade in R, with costs and actual initial stop-risk dollars;
|
||||
- Calmar (CAGR / max drawdown);
|
||||
- profit factor on net trade R;
|
||||
- Gain-to-Pain (sum of all monthly returns / absolute sum of negative months);
|
||||
- Sortino using daily returns and zero target.
|
||||
|
||||
Also report total return/CAGR, maximum drawdown, Sharpe, win rate, time
|
||||
underwater, exposure, cash, average/peak positions, sessions at capacity,
|
||||
turnover, costs, qualified/admitted/blocked opportunities, and minimum-risk
|
||||
rejections.
|
||||
|
||||
For each arm/protocol/cost/metric, pair identical paths with cap10_incumbent,
|
||||
take the median paired delta within each start year or annual anchor, show all
|
||||
seven cluster values, and headline their median.
|
||||
|
||||
Initialization dispersion is reported separately for EV and Calmar: calculate
|
||||
the seed-path IQR within each warm anchor, divide by the paired control IQR, show
|
||||
all seven ratios, and headline their median. Do not combine them into a composite.
|
||||
|
||||
For context only, run a deterministic 10,000-replicate cluster bootstrap over
|
||||
the seven paired annual summaries and report the central 90% percentile interval
|
||||
for median EV and Calmar deltas and warm IQR ratios. These intervals are not
|
||||
promotion gates, independent-population confidence claims, or formal inference.
|
||||
|
||||
## Reproducibility and execution
|
||||
|
||||
Candidate replay/ranks cache under reports/.cache; each matrix cell checkpoints
|
||||
atomically and resume verifies a fingerprint over the implementation commit,
|
||||
this specification hash, snapshot SHA-256, cache key, arm definitions, costs,
|
||||
and cohort manifest. An authoritative run refuses a dirty worktree.
|
||||
|
||||
The existing v1 candidate/rank cache is intentionally reusable: its
|
||||
full-universe current-day ranks are correct. Runner v2 derives a fingerprinted
|
||||
construction view by removing qualified rows whose symbols are rank-only. V2
|
||||
uses a versioned checkpoint directory, so invalid v1 portfolio cells are never
|
||||
resumed and the expensive daily rank replay does not need to run again.
|
||||
|
||||
The loader reads only ticker ID/symbol and the OHLCV columns used by replay, so
|
||||
snapshots created before SEC metadata added `tickers.cik`, `tickers.sic`, and
|
||||
`tickers.sic_description` remain valid. Do not migrate or alter the research
|
||||
snapshot: its original SHA-256 is part of the run fingerprint.
|
||||
|
||||
macOS environment setup from the repository root (zsh):
|
||||
|
||||
python3 -m venv .venv
|
||||
./.venv/bin/python -m pip install -e '.[dev]'
|
||||
|
||||
Preflight:
|
||||
|
||||
./.venv/bin/python scripts/run_portfolio_construction_matrix.py \
|
||||
backtest_snapshots/research.sqlite \
|
||||
--run-id prod505-capacity-bracket-daily-v1 \
|
||||
--workers auto \
|
||||
--resume \
|
||||
--validate-only
|
||||
|
||||
Authoritative run:
|
||||
|
||||
./.venv/bin/python scripts/run_portfolio_construction_matrix.py \
|
||||
backtest_snapshots/research.sqlite \
|
||||
--run-id prod505-capacity-bracket-daily-v1 \
|
||||
--workers auto \
|
||||
--resume
|
||||
|
||||
Commit only the compact final JSON and Markdown reports. Raw curves, trades,
|
||||
candidate caches, and checkpoints remain ignored.
|
||||
@@ -1,75 +0,0 @@
|
||||
# Regime Monitor v2 methodology
|
||||
|
||||
The Regime Monitor is an observational AI/Tech risk thermometer. It does not
|
||||
gate entries, exits, position size, ranking, or alerts about individual setups.
|
||||
|
||||
## Outputs
|
||||
|
||||
**State** measures current structural stress:
|
||||
|
||||
- Price structure, 40%: `max(P1, P2, P3)`, so the correlated 200-DMA, death-cross,
|
||||
and drawdown readings receive one capped vote.
|
||||
- Fixed-basket breadth level, 25%.
|
||||
- HY option-adjusted credit spread, 20%.
|
||||
- VIX level, 15%.
|
||||
|
||||
**Warning** measures deterioration and divergence:
|
||||
|
||||
- Fixed-basket breadth divergence while SMH holds/rises, 50%.
|
||||
- 60-session SMH/SPY relative-strength deterioration, 30%.
|
||||
- Hyperscaler capex cuts, 12%.
|
||||
- Good-news-stock-down earnings reactions, 8%.
|
||||
|
||||
Combined, RSP/SPY (former F4), and the NVDA canary (former P6) do not enter v2.
|
||||
|
||||
## Scale and missing data
|
||||
|
||||
Zero means ordinary/healthy, and only stress contributes positively. Automated
|
||||
capex `raising`/`holding` and no good-news-stock-down pattern map to zero;
|
||||
`mixed`, unknown, and stale observations are unavailable rather than neutral 50.
|
||||
Manual observations use the same categories: each hyperscaler is marked
|
||||
`raising`, `holding`, `cutting`, or `unknown`, while the earnings reaction is
|
||||
`yes`, `no`, or `mixed`. F1 is derived from the share of at least three known
|
||||
hyperscalers marked `cutting`; arbitrary numeric overrides are not accepted.
|
||||
|
||||
Scores renormalize over available fixed weights, but a band is published only at
|
||||
75% or greater coverage. Trend deltas are suppressed when the participating
|
||||
pillar set changes. Bands are stable `<30`, watch `<60`, elevated `<80`, and
|
||||
breaking `>=80`.
|
||||
|
||||
Credit uses named HY OAS anchors (3.5 mild, 5.0 elevated, 7.0 stressed) for 70%
|
||||
of its score and a ten-year upper-tail percentile for 30%.
|
||||
|
||||
## Point-in-time record
|
||||
|
||||
The first v2 run rebuilds the latest 400 trading sessions with sufficient sensor
|
||||
warm-up. Routine runs thereafter insert/update only the latest trading date.
|
||||
Fundamental observations have an effective date (normally the next session after
|
||||
collection) and are never replayed backward. The history API and main chart show
|
||||
only snapshots marked `methodology: v2`.
|
||||
|
||||
Each snapshot stores the fixed basket symbols, hash, and freeze date. Reconstructed
|
||||
history before that freeze date is retrospective/exploratory; readings after it
|
||||
form the forward record.
|
||||
|
||||
The automatic 400-session rebuild is intentionally one-shot: it runs only when
|
||||
no v2 snapshot exists. If an initial seed used partial data or the wrong basket,
|
||||
the operational reseed procedure is to remove the v2 snapshot rows and run the
|
||||
Regime Monitor job again. There is no routine force-rebuild flag.
|
||||
|
||||
## Warning study
|
||||
|
||||
The study calls the outcome a **10% correction**, not a regime break. The first
|
||||
70% of sessions freezes the 80th-percentile warning threshold; alarm episodes are
|
||||
measured on the final 30%. An alarm requires an upward crossing and another alarm
|
||||
requires a reset below the threshold. The report exposes warned/missed events,
|
||||
false alarms per year, median lead, sample dates, event count, report date, and
|
||||
whether the result is exploratory or a true forward holdout. UI claims are
|
||||
generated from that report; no performance sentence is hard-coded.
|
||||
|
||||
## Operator rule
|
||||
|
||||
Quadrant alerts default off for new/reset configurations. When enabled they
|
||||
require fresh inputs, at least 75% coverage on both axes, two consecutive daily
|
||||
confirmations, hysteresis, and cooldown. Every alert states: **Risk thermometer —
|
||||
not a trade signal.**
|
||||
@@ -0,0 +1,202 @@
|
||||
# Regime Monitor v3 methodology
|
||||
|
||||
The Regime Monitor is an observational AI/Tech risk thermometer. It does not
|
||||
gate entries, exits, position size, ranking, or alerts about individual setups.
|
||||
|
||||
v3 supersedes v2. Every parameter below was calibrated against the 408 v2
|
||||
sessions ending 2026-07-24, reproduced offline from the same Alpaca and FRED
|
||||
inputs the live job uses; the reproduction matched the stored prod distribution
|
||||
exactly (State avg 22.6/22.7, p80 35.1, max 91.2, P3 pegged 39, W1 live 108).
|
||||
|
||||
## What changed and why
|
||||
|
||||
**Fundamentals left the score.** F1 (capex) and F3 (good-news-stock-down)
|
||||
carried 12 + 8 of 100 Warning points. Pegged at maximum stress they produced a
|
||||
Warning of exactly 20.0 — below the event study's 25.3 alarm threshold, and
|
||||
still inside the "stable" band. The sourced observation could not change any
|
||||
published conclusion, so refreshing it looked like it did nothing. They are now
|
||||
a qualitative overlay reported beside the scores. Capex also stopped scoring
|
||||
`raising` and `holding` identically at 0: `holding` is the deceleration case and
|
||||
now scores 50, so a boom no longer reads the same as a stall.
|
||||
|
||||
**The drawdown sensor stopped saturating.** v2 used `dd_pct * 5`, reaching 100 at
|
||||
a 20% drawdown — the 90th percentile of the observed distribution. 39 of 408
|
||||
sessions sat at exactly 100 with no resolution left, and the price pillar showed
|
||||
the top band on 13.5% of sessions. v3 uses named anchors with headroom past the
|
||||
observed 36% maximum, and blends leader/confirm 2:1 as P1 and P2 already did
|
||||
instead of taking `max()`. P3's realized share of State falls from 65% to 40%,
|
||||
matching its nominal weight.
|
||||
|
||||
**Warning gained a sensor with range.** The HY OAS *level* is pinned at zero
|
||||
below the 3.5 mild anchor (2.77 at the cutover), so credit contributed nothing
|
||||
in a calm tape. Its 20-session rate of change still does, and spread widening is
|
||||
a classic lead.
|
||||
|
||||
**The credit percentile leg was removed.** Its reference window silently shrank
|
||||
from 10 years to 3 when ICE restricted the upstream series in April 2026, after
|
||||
which it scored 20 points of stress at a spread the same sensor's anchors call
|
||||
"mild". See Calibration below.
|
||||
|
||||
**Breadth loss counts during declines.** v2's divergence gate was
|
||||
`price_ret >= 0`, so the sensor zeroed during every selloff. On 2026-07-24 the
|
||||
basket shed 10 points of participation in 20 sessions while SMH fell 11.9% and
|
||||
Warning printed exactly 0. v3 tapers to a floor instead: deterioration counts
|
||||
fully when price masks it (true divergence, the dangerous pre-top case) and at
|
||||
35% when price confirms it. Breadth *level* lives in State, but breadth
|
||||
*velocity* appears nowhere else, so this is not double counting.
|
||||
|
||||
**Bands are per axis.** v2 Warning never exceeded 64.9 in 408 sessions while
|
||||
State reached 91.2, yet both used 30/60/80 with quadrant dividers at 60. The
|
||||
upper half of the Warning axis was unreachable.
|
||||
|
||||
## Outputs
|
||||
|
||||
**State** — current structural stress:
|
||||
|
||||
- Price structure, 40%: `max(P1, P2, P3)`, one capped vote for correlated reads.
|
||||
- Fixed-basket breadth level, 25%.
|
||||
- HY option-adjusted credit spread level, 20%.
|
||||
- VIX level, 15%.
|
||||
|
||||
**Warning** — deterioration and divergence:
|
||||
|
||||
- Fixed-basket breadth divergence, 45%.
|
||||
- 60-session SMH/SPY relative-strength deterioration, 30%.
|
||||
- HY OAS 20-session widening, 25%.
|
||||
|
||||
Combined, RSP/SPY (former F4), and the NVDA canary (former P6) do not enter v3.
|
||||
|
||||
## Calibration
|
||||
|
||||
P3 drawdown anchors, as (drawdown %, score): 0→0, 4→10, 8→25, 16→50, 28→78,
|
||||
40→100, flat outside. Credit impulse is relative (+35% over 20 sessions = 100)
|
||||
rather than absolute, because +0.5pp means something very different at an OAS of
|
||||
2.7 than at 8.0.
|
||||
|
||||
Bands are round, meaning-anchored numbers, not percentile fits — percentile
|
||||
thresholds would drift on every rebuild and silently rewrite what past snapshots
|
||||
meant. Realized shares over the calibration window:
|
||||
|
||||
| Axis | stable | watch | elevated | breaking | thresholds |
|
||||
|------|--------|-------|----------|----------|------------|
|
||||
| State | 73.3% | 15.0% | 8.3% | 3.4% | 20 / 50 / 80 |
|
||||
| Warning | 69.4% | 19.6% | 7.6% | 3.4% | 20 / 40 / 60 |
|
||||
|
||||
Quadrant dividers sit at each axis's watch/elevated boundary: State 50,
|
||||
Warning 40.
|
||||
|
||||
Scores renormalize over available fixed weights, but a band is published only at
|
||||
75% or greater coverage. Trend deltas are suppressed when the participating
|
||||
pillar set changes. Zero means ordinary/healthy; only stress contributes.
|
||||
|
||||
Credit level is the named HY OAS anchors alone: 3.5 mild, 5.0 elevated, 7.0
|
||||
stressed, linear between, and nothing else. v2 blended those anchors at 70% with
|
||||
a 30% upper-tail percentile over a nominally 10-year window.
|
||||
|
||||
That leg was removed rather than repaired. ICE restricted FRED to a rolling
|
||||
3-year window for `BAMLH0A0HYM2` in April 2026 — the series metadata states it
|
||||
outright ("Starting in April 2026, this series will only include 3 years of
|
||||
observations"), and an unbounded request returns the same 795 observations as a
|
||||
30-year one. The v2 percentile therefore ranked the current spread against three
|
||||
uniformly tight years (range 2.59–4.61 over the calibration window), which made
|
||||
it fire early and saturate absurdly: at an OAS of 3.50 — the level the anchors
|
||||
call *mild*, scoring zero stress — the blended sensor read 20.1, and the
|
||||
percentile leg pegged at 100 by an OAS of 4.5. Across the 408 sessions it
|
||||
roughly tripled the credit sensor's average (2.70 vs 1.00) and more than doubled
|
||||
its nonzero days (60 vs 27).
|
||||
|
||||
The anchors already encode the long-run distribution as constants, so the
|
||||
percentile was a second, noisier estimate of the same thing. What it was
|
||||
genuinely reaching for — "unusual versus recent history" — is now W3 on the
|
||||
Warning axis, computed as a rate of change, which is where deterioration
|
||||
belongs. Removing it moved State's average by −0.4 and its maximum by −3.8, left
|
||||
Warning bit-identical, and did not shift any band threshold.
|
||||
|
||||
A long-history alternative (`BAA10Y`, Fed-published, 7,712 observations back to
|
||||
1997) was considered and rejected: ranking an HY spread against investment-grade
|
||||
history is not a coherent statistic, and it would rescue a leg that is redundant
|
||||
anyway.
|
||||
|
||||
Every snapshot now records `data_quality.credit_history_days` and
|
||||
`vix_history_days`. This defect was invisible for roughly three months because
|
||||
nothing asserted the window the code claimed; the spans make a future upstream
|
||||
truncation show up in the record instead of quietly reshaping a sensor.
|
||||
|
||||
**Survivorship caveat.** The basket was frozen 2026-07-15 but the calibration
|
||||
window reaches back to 2024, so names were partly selected for having done well.
|
||||
Every distribution above inherits that bias. It is the same bias v2 carried, so
|
||||
the v2/v3 comparison is like-for-like, but the absolute band shares are
|
||||
optimistic.
|
||||
|
||||
## Point-in-time record
|
||||
|
||||
The first run under a new `METHODOLOGY` rebuilds the latest 400 trading sessions
|
||||
with sufficient sensor warm-up; routine runs thereafter insert/update only the
|
||||
latest trading date. The history API and main chart show only snapshots matching
|
||||
the current methodology, so a bump reseeds the series rather than splicing two
|
||||
formulas into one line.
|
||||
|
||||
The fundamental overlay keeps its effective date (normally the next session after
|
||||
collection) and is never replayed backward, so a rebuild cannot stamp today's
|
||||
observation onto historical snapshots. Because the observation is stored in a
|
||||
single slot, a refresh replaces the previously effective record: the snapshot
|
||||
therefore reports the overlay as `pending` until the new effective date, and the
|
||||
live reading additionally carries `fundamental_context` so a just-collected
|
||||
observation is visible immediately rather than appearing to have done nothing.
|
||||
|
||||
Each snapshot stores the fixed basket symbols, hash, and freeze date.
|
||||
Reconstructed history before that freeze date is retrospective/exploratory.
|
||||
|
||||
## Warning study
|
||||
|
||||
The study calls the outcome a **10% correction**, not a regime break. The first
|
||||
70% of sessions freezes the 80th-percentile warning threshold; alarm episodes are
|
||||
measured on the final 30%. Because v3 dropped fundamentals from the score, the
|
||||
study now measures exactly the live Warning score rather than a technical-only
|
||||
approximation of it, and both are computed from one shared sensor definition
|
||||
(`warning_sensor_scores`) so they cannot drift apart.
|
||||
|
||||
A cached report is discarded when its methodology no longer matches, so the panel
|
||||
reverts to "not run yet" after a bump rather than showing stale numbers. **Re-run
|
||||
the Event Study job after cutting over to v3.**
|
||||
|
||||
### Reading the result
|
||||
|
||||
The report carries a `reliability` block and the UI renders its warnings, because
|
||||
the headline numbers invite over-reading in two specific ways.
|
||||
|
||||
**The holdout is thin.** The study detects 11 corrections across 5 years but the
|
||||
70/30 split leaves only 4 in the test period. Recall is therefore one event away
|
||||
from a materially different headline, and in practice the event that flips is
|
||||
decided by where the frozen threshold happens to land rather than by whether the
|
||||
score saw anything. The v3 cutover run illustrates it: v3 scored 2/4 against v2's
|
||||
3/4, but "v3 without the credit sensor" scores 3/4 at a *higher* threshold
|
||||
(35.5) than shipped v3 misses it at (32.3) — because the alarm rule needs a
|
||||
rising edge, and a lower threshold can mean the alarm already fired outside the
|
||||
20-session horizon and never reset below. Below `MIN_EVENTS_FOR_CONFIDENCE`
|
||||
holdout events the report says so explicitly.
|
||||
|
||||
Some events carry no information at all for comparison: in that run every
|
||||
variant caught 2026-03-06, every variant missed 2026-06-05, and every variant
|
||||
"caught" 2025-11-20 with a 1-session lead, which is coincident rather than a
|
||||
warning.
|
||||
|
||||
**Sensor coverage can straddle the split.** The score renormalises over available
|
||||
sensors, so a training window predating a sensor's history freezes the threshold
|
||||
on a different construct than the holdout is measured against. At the v3 cutover
|
||||
only 39% of training sessions had all three Warning sensors versus 100% of the
|
||||
test period, because credit history begins 2023-07-25.
|
||||
|
||||
Restricting the threshold to sensor-matched training sessions was tried and is
|
||||
*not* the fix: those sessions are a calm recent stretch, so the threshold drops
|
||||
from 32.3 to 22.5 and false alarms rise from 3.3 to 8.6 per year. It trades a
|
||||
coverage bias for a regime-selection bias. The honest position is that the
|
||||
threshold is hypersensitive to window choice at this sample size; the report
|
||||
states its limits rather than pretending to a precision it does not have.
|
||||
|
||||
## Operator rule
|
||||
|
||||
Quadrant alerts default off for new/reset configurations. When enabled they
|
||||
require fresh inputs, at least 75% coverage on both axes, two consecutive daily
|
||||
confirmations, hysteresis, and cooldown. Every alert states: **Risk thermometer —
|
||||
not a trade signal.**
|
||||
@@ -4,6 +4,7 @@ import type {
|
||||
AdminUser,
|
||||
AlertConfig,
|
||||
AlertTestResult,
|
||||
FundamentalsCutoverConfig,
|
||||
PipelineReadiness,
|
||||
RecommendationConfig,
|
||||
ScheduleConfig,
|
||||
@@ -56,6 +57,18 @@ export function updateSetting(key: string, value: string) {
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
export function getFundamentalsCutoverSettings() {
|
||||
return apiClient
|
||||
.get<FundamentalsCutoverConfig>('admin/settings/fundamentals-cutover')
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
export function updateFundamentalsCutoverSettings(enabled: boolean) {
|
||||
return apiClient
|
||||
.put<FundamentalsCutoverConfig>('admin/settings/fundamentals-cutover', { enabled })
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
export function getRecommendationSettings() {
|
||||
return apiClient
|
||||
.get<RecommendationConfig>('admin/settings/recommendations')
|
||||
@@ -233,6 +246,40 @@ export interface TriggerJobResponse {
|
||||
cadence?: BacktestCadence;
|
||||
}
|
||||
|
||||
export interface ParityFieldStats {
|
||||
legacy_available: number;
|
||||
candidate_available: number;
|
||||
both_available: number;
|
||||
material_differences: number;
|
||||
median_absolute_delta: number | null;
|
||||
p95_absolute_delta: number | null;
|
||||
max_absolute_delta: number | null;
|
||||
}
|
||||
|
||||
export interface FundamentalsParityReport {
|
||||
report_version: number;
|
||||
generated_at: string;
|
||||
as_of_date: string;
|
||||
approval_status: string;
|
||||
read_only: boolean;
|
||||
summary: {
|
||||
universe_count: number;
|
||||
legacy_fundamental_score_available: number;
|
||||
candidate_fundamental_score_available: number;
|
||||
fundamental_scores_compared: number;
|
||||
fundamental_score_material_changes: number;
|
||||
fundamental_rank_changes: number;
|
||||
field_stats: Record<string, ParityFieldStats>;
|
||||
};
|
||||
source_runs: Record<string, {
|
||||
run_id: number;
|
||||
status: string;
|
||||
revision: string | null;
|
||||
source_max_date: string | null;
|
||||
completed_at: string | null;
|
||||
} | null>;
|
||||
}
|
||||
|
||||
export type BacktestTargetModel = 'production_gtl' | 'structural_sr';
|
||||
export type BacktestCadence = 'weekly' | 'daily';
|
||||
|
||||
@@ -259,6 +306,24 @@ export function triggerJob(
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
export function getFundamentalsParityReport() {
|
||||
return apiClient
|
||||
.get<FundamentalsParityReport | null>('admin/fundamentals-parity')
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
export function getFundamentalsParityCsv() {
|
||||
return apiClient
|
||||
.get<{ filename: string; content: string } | null>('admin/fundamentals-parity/csv')
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
export function getFundamentalsParityJson() {
|
||||
return apiClient
|
||||
.get<{ filename: string; content: string } | null>('admin/fundamentals-parity/json')
|
||||
.then((r) => r.data);
|
||||
}
|
||||
|
||||
// System events (operational warnings / errors)
|
||||
export interface SystemEvent {
|
||||
id: number;
|
||||
|
||||
@@ -0,0 +1,176 @@
|
||||
import {
|
||||
useFundamentalsCutoverSettings,
|
||||
useJobs,
|
||||
useTriggerJob,
|
||||
useUpdateFundamentalsCutoverSettings,
|
||||
} from '../../hooks/useAdmin';
|
||||
import { SkeletonCard } from '../ui/Skeleton';
|
||||
|
||||
const SEC_JOB = 'sec_fundamentals_import';
|
||||
|
||||
function formatRun(iso: string | null | undefined): string {
|
||||
if (!iso) return 'not run in this process';
|
||||
const minutes = Math.floor((Date.now() - new Date(iso).getTime()) / 60_000);
|
||||
if (minutes < 1) return 'just now';
|
||||
if (minutes < 60) return `${minutes}m ago`;
|
||||
const hours = Math.floor(minutes / 60);
|
||||
return hours < 24 ? `${hours}h ago` : `${Math.floor(hours / 24)}d ago`;
|
||||
}
|
||||
|
||||
export function FundamentalsCutoverSettings() {
|
||||
const cutover = useFundamentalsCutoverSettings();
|
||||
const update = useUpdateFundamentalsCutoverSettings();
|
||||
const trigger = useTriggerJob();
|
||||
const { data: jobs } = useJobs();
|
||||
|
||||
if (cutover.isLoading) return <SkeletonCard />;
|
||||
if (cutover.isError || !cutover.data) {
|
||||
return (
|
||||
<p className="text-sm text-red-400">
|
||||
{(cutover.error as Error)?.message || 'Failed to load fundamentals data source'}
|
||||
</p>
|
||||
);
|
||||
}
|
||||
|
||||
const enabled = cutover.data.enabled;
|
||||
const secJob = jobs?.find((job) => job.name === SEC_JOB);
|
||||
const runningJob = jobs?.find((job) => job.running);
|
||||
const refreshBlocked = Boolean(runningJob && runningJob.name !== SEC_JOB);
|
||||
|
||||
const changeSource = () => {
|
||||
const next = !enabled;
|
||||
const confirmed = window.confirm(
|
||||
next
|
||||
? 'Activate SEC + Dolt fundamentals? The next SEC import will replace the legacy cache and mark affected scores stale.'
|
||||
: 'Pause SEC + Dolt cache refreshes? Existing cache values will stay in place; legacy values are not restored automatically.',
|
||||
);
|
||||
if (confirmed) update.mutate(next);
|
||||
};
|
||||
|
||||
return (
|
||||
<section className="glass overflow-hidden" aria-labelledby="fundamentals-source-title">
|
||||
<div className={`h-0.5 ${enabled ? 'bg-gradient-to-r from-sky-500 via-cyan-300 to-emerald-400' : 'bg-white/[0.06]'}`} />
|
||||
<div className="space-y-5 p-5">
|
||||
<div className="flex flex-wrap items-start justify-between gap-3">
|
||||
<div>
|
||||
<div className="flex items-center gap-2">
|
||||
<h3 id="fundamentals-source-title" className="text-sm font-semibold text-gray-200">
|
||||
Fundamentals data source
|
||||
</h3>
|
||||
<span
|
||||
className={`rounded-full border px-2 py-0.5 text-[10px] font-semibold uppercase tracking-[0.14em] ${
|
||||
enabled
|
||||
? 'border-cyan-400/25 bg-cyan-400/10 text-cyan-300'
|
||||
: 'border-white/10 bg-white/[0.04] text-gray-500'
|
||||
}`}
|
||||
>
|
||||
{enabled ? 'SEC + Dolt active' : 'Legacy cache'}
|
||||
</span>
|
||||
</div>
|
||||
<p className="mt-1 max-w-3xl text-xs leading-relaxed text-gray-500">
|
||||
Controls what repopulates <span className="num text-gray-400">fundamental_data</span>, the
|
||||
compatibility cache used by scoring. SEC filings supply P/E, growth and estimated market
|
||||
cap; Dolt supplies earnings dates and surprises. Everything is derived locally from PostgreSQL.
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="grid grid-cols-[minmax(0,1fr)_5rem_minmax(0,1fr)] items-center gap-3 rounded-xl border border-white/[0.06] bg-black/10 px-4 py-3">
|
||||
<div className={enabled ? 'text-gray-600' : 'text-amber-200/90'}>
|
||||
<div className="num text-[10px] uppercase tracking-[0.16em]">Legacy APIs</div>
|
||||
<div className="mt-0.5 text-[11px]">FMP / Finnhub / Alpha Vantage</div>
|
||||
</div>
|
||||
<div className="relative h-px bg-white/10" aria-hidden="true">
|
||||
<span
|
||||
className={`absolute top-1/2 h-2.5 w-2.5 -translate-y-1/2 rounded-full border-2 border-[#0e120f] transition-all duration-300 ${
|
||||
enabled
|
||||
? 'right-0 bg-cyan-300 shadow-[0_0_12px_rgba(103,232,249,0.55)]'
|
||||
: 'left-0 bg-amber-300'
|
||||
}`}
|
||||
/>
|
||||
</div>
|
||||
<div className={`text-right ${enabled ? 'text-cyan-200' : 'text-gray-600'}`}>
|
||||
<div className="num text-[10px] uppercase tracking-[0.16em]">SEC + Dolt</div>
|
||||
<div className="mt-0.5 text-[11px]">Bulk imports → PostgreSQL cache</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="grid gap-4 border-t border-white/[0.06] pt-4 md:grid-cols-2">
|
||||
<div className="flex items-start justify-between gap-4 rounded-xl bg-white/[0.025] p-3.5">
|
||||
<div>
|
||||
<div className="num text-[10px] uppercase tracking-[0.14em] text-gray-600">1 · Source</div>
|
||||
<div className="mt-1 text-sm text-gray-200">Use SEC + Dolt for scoring inputs</div>
|
||||
<p className="mt-1 text-[11px] leading-relaxed text-gray-500">
|
||||
While active, the weekly legacy collector is skipped so it cannot overwrite the new cache.
|
||||
</p>
|
||||
</div>
|
||||
<button
|
||||
type="button"
|
||||
role="switch"
|
||||
aria-checked={enabled}
|
||||
aria-label="Use SEC and Dolt fundamentals"
|
||||
onClick={changeSource}
|
||||
disabled={update.isPending}
|
||||
className={`relative mt-1 inline-flex h-6 w-11 shrink-0 rounded-full border-2 border-transparent transition-colors focus:outline-none focus:ring-2 focus:ring-cyan-400/70 focus:ring-offset-2 focus:ring-offset-[#0e120f] disabled:cursor-wait disabled:opacity-50 ${
|
||||
enabled ? 'bg-gradient-to-r from-sky-500 to-cyan-400' : 'bg-white/10'
|
||||
}`}
|
||||
>
|
||||
<span
|
||||
className={`pointer-events-none inline-block h-5 w-5 rounded-full bg-white shadow transition-transform ${
|
||||
enabled ? 'translate-x-5' : 'translate-x-0'
|
||||
}`}
|
||||
/>
|
||||
</button>
|
||||
</div>
|
||||
|
||||
<div className="rounded-xl bg-white/[0.025] p-3.5">
|
||||
<div className="num text-[10px] uppercase tracking-[0.14em] text-gray-600">2 · Refresh</div>
|
||||
<div className="mt-1 flex flex-wrap items-center justify-between gap-3">
|
||||
<div>
|
||||
<div className="text-sm text-gray-200">Apply the source now</div>
|
||||
<p className="mt-1 text-[11px] text-gray-500">
|
||||
{secJob?.running
|
||||
? 'SEC import and cache refresh are running.'
|
||||
: secJob?.runtime_message || `Last SEC run: ${formatRun(secJob?.runtime_finished_at)}`}
|
||||
</p>
|
||||
</div>
|
||||
<button
|
||||
type="button"
|
||||
onClick={() => trigger.mutate(SEC_JOB)}
|
||||
disabled={
|
||||
!enabled ||
|
||||
trigger.isPending ||
|
||||
Boolean(secJob?.running) ||
|
||||
refreshBlocked ||
|
||||
secJob?.enabled === false
|
||||
}
|
||||
className="btn-primary px-3 py-2 text-xs disabled:cursor-not-allowed disabled:opacity-40"
|
||||
>
|
||||
<span>
|
||||
{secJob?.running
|
||||
? 'Refreshing…'
|
||||
: trigger.isPending
|
||||
? 'Starting…'
|
||||
: refreshBlocked
|
||||
? 'Another job is running'
|
||||
: 'Run refresh now'}
|
||||
</span>
|
||||
</button>
|
||||
</div>
|
||||
{!enabled && (
|
||||
<p className="mt-2 text-[11px] text-amber-300/70">Activate the source before running the refresh.</p>
|
||||
)}
|
||||
{enabled && secJob?.enabled === false && (
|
||||
<p className="mt-2 text-[11px] text-amber-300/70">Enable the SEC Fundamentals job on the Jobs tab first.</p>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<p className="text-[11px] leading-relaxed text-gray-600">
|
||||
Rollback pauses future writes only. To restore pre-cutover values, use the database backup or
|
||||
pause this source and manually run the legacy collector while its provider keys remain installed.
|
||||
</p>
|
||||
</div>
|
||||
</section>
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
import { useState } from 'react';
|
||||
import {
|
||||
getFundamentalsParityCsv,
|
||||
getFundamentalsParityJson,
|
||||
} from '../../api/admin';
|
||||
import { useFundamentalsParityReport } from '../../hooks/useAdmin';
|
||||
import { SkeletonTable } from '../ui/Skeleton';
|
||||
|
||||
const FIELD_LABELS: Record<string, string> = {
|
||||
pe_ratio: 'P/E',
|
||||
revenue_growth: 'Revenue growth',
|
||||
earnings_surprise: 'Earnings surprise',
|
||||
};
|
||||
|
||||
function downloadText(filename: string, content: string, type: string) {
|
||||
const blob = new Blob([content], { type });
|
||||
const url = URL.createObjectURL(blob);
|
||||
const anchor = document.createElement('a');
|
||||
anchor.href = url;
|
||||
anchor.download = filename;
|
||||
anchor.click();
|
||||
URL.revokeObjectURL(url);
|
||||
}
|
||||
|
||||
export function FundamentalsParityPanel() {
|
||||
const { data: report, isLoading, isError, error } = useFundamentalsParityReport();
|
||||
const [downloading, setDownloading] = useState(false);
|
||||
|
||||
if (isLoading) return <SkeletonTable rows={2} cols={4} />;
|
||||
if (isError) {
|
||||
return <p className="text-sm text-red-400">{(error as Error).message}</p>;
|
||||
}
|
||||
|
||||
if (!report) {
|
||||
return (
|
||||
<div className="glass p-5">
|
||||
<h3 className="text-sm font-semibold text-gray-200">A5 Fundamentals Parity</h3>
|
||||
<p className="mt-1 text-xs text-gray-500">
|
||||
No report yet. Trigger “Fundamentals Parity Report (read-only)” below.
|
||||
</p>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
const summary = report.summary;
|
||||
const generated = new Date(report.generated_at).toLocaleString();
|
||||
|
||||
async function downloadCsv() {
|
||||
setDownloading(true);
|
||||
try {
|
||||
const artifact = await getFundamentalsParityCsv();
|
||||
if (artifact) downloadText(artifact.filename, artifact.content, 'text/csv;charset=utf-8');
|
||||
} finally {
|
||||
setDownloading(false);
|
||||
}
|
||||
}
|
||||
|
||||
async function downloadJson() {
|
||||
setDownloading(true);
|
||||
try {
|
||||
const artifact = await getFundamentalsParityJson();
|
||||
if (artifact) downloadText(artifact.filename, artifact.content, 'application/json;charset=utf-8');
|
||||
} finally {
|
||||
setDownloading(false);
|
||||
}
|
||||
}
|
||||
|
||||
return (
|
||||
<div className="glass p-5 space-y-4">
|
||||
<div className="flex flex-wrap items-start justify-between gap-3">
|
||||
<div>
|
||||
<div className="flex flex-wrap items-center gap-2">
|
||||
<h3 className="text-sm font-semibold text-gray-200">A5 Fundamentals Parity</h3>
|
||||
<span className="rounded-full border border-amber-400/20 bg-amber-400/10 px-2 py-0.5 text-[10px] uppercase tracking-wide text-amber-300">
|
||||
approval pending
|
||||
</span>
|
||||
<span className="rounded-full border border-cyan-400/20 bg-cyan-400/10 px-2 py-0.5 text-[10px] uppercase tracking-wide text-cyan-300">
|
||||
read-only
|
||||
</span>
|
||||
</div>
|
||||
<p className="mt-1 text-xs text-gray-500">
|
||||
Generated {generated} · as of {report.as_of_date} · {summary.universe_count} tracked tickers
|
||||
</p>
|
||||
</div>
|
||||
<div className="flex gap-2">
|
||||
<button
|
||||
type="button"
|
||||
className="rounded border border-white/10 px-3 py-1.5 text-xs text-gray-300 hover:text-white"
|
||||
onClick={downloadJson}
|
||||
disabled={downloading}
|
||||
>
|
||||
Download JSON
|
||||
</button>
|
||||
<button
|
||||
type="button"
|
||||
className="rounded border border-white/10 px-3 py-1.5 text-xs text-gray-300 hover:text-white disabled:opacity-50"
|
||||
onClick={downloadCsv}
|
||||
disabled={downloading}
|
||||
>
|
||||
{downloading ? 'Preparing…' : 'Download CSV'}
|
||||
</button>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="grid gap-3 sm:grid-cols-2 lg:grid-cols-4">
|
||||
<Summary label="Candidate score coverage" value={`${summary.candidate_fundamental_score_available}/${summary.universe_count}`} />
|
||||
<Summary label="Scores compared" value={summary.fundamental_scores_compared} />
|
||||
<Summary label="Material score moves" value={summary.fundamental_score_material_changes} />
|
||||
<Summary label="Fundamental rank moves" value={summary.fundamental_rank_changes} />
|
||||
</div>
|
||||
|
||||
<div className="overflow-x-auto">
|
||||
<table className="w-full text-left text-xs">
|
||||
<thead className="text-[10px] uppercase tracking-wider text-gray-500">
|
||||
<tr>
|
||||
<th className="pb-2 pr-4 font-medium">Field</th>
|
||||
<th className="pb-2 px-3 font-medium">Legacy</th>
|
||||
<th className="pb-2 px-3 font-medium">Candidate</th>
|
||||
<th className="pb-2 px-3 font-medium">Compared</th>
|
||||
<th className="pb-2 px-3 font-medium">Material</th>
|
||||
<th className="pb-2 pl-3 font-medium">Median |Δ|</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody className="divide-y divide-white/[0.06] text-gray-300">
|
||||
{Object.entries(summary.field_stats).map(([key, stats]) => (
|
||||
<tr key={key}>
|
||||
<td className="py-2.5 pr-4">{FIELD_LABELS[key] ?? key}</td>
|
||||
<td className="py-2.5 px-3 num">{stats.legacy_available}</td>
|
||||
<td className="py-2.5 px-3 num">{stats.candidate_available}</td>
|
||||
<td className="py-2.5 px-3 num">{stats.both_available}</td>
|
||||
<td className="py-2.5 px-3 num">{stats.material_differences}</td>
|
||||
<td className="py-2.5 pl-3 num">
|
||||
{stats.median_absolute_delta == null ? 'n/a' : stats.median_absolute_delta.toFixed(2)}
|
||||
</td>
|
||||
</tr>
|
||||
))}
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
|
||||
<p className="text-[11px] leading-relaxed text-gray-500">
|
||||
Materiality bands highlight review candidates only. They do not approve a cutover or write fundamentals,
|
||||
scores, rankings, or qualification state.
|
||||
</p>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
function Summary({ label, value }: { label: string; value: string | number }) {
|
||||
return (
|
||||
<div className="rounded-lg border border-white/[0.07] bg-white/[0.025] px-3 py-2.5">
|
||||
<div className="text-[10px] uppercase tracking-wider text-gray-500">{label}</div>
|
||||
<div className="mt-1 num text-lg text-gray-200">{value}</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
@@ -25,7 +25,7 @@ function formatAgo(iso: string | null | undefined): string {
|
||||
|
||||
function lastRunColor(status: string | null | undefined): string {
|
||||
if (status === 'error') return 'text-red-300';
|
||||
if (status === 'rate_limited') return 'text-amber-300';
|
||||
if (status === 'rate_limited' || status === 'deferred') return 'text-amber-300';
|
||||
return 'text-gray-500';
|
||||
}
|
||||
|
||||
@@ -127,7 +127,7 @@ export function JobControls() {
|
||||
className={`text-[11px] font-medium ${
|
||||
job.running
|
||||
? 'text-blue-300'
|
||||
: job.runtime_status === 'rate_limited'
|
||||
: job.runtime_status === 'rate_limited' || job.runtime_status === 'deferred'
|
||||
? 'text-amber-300'
|
||||
: job.runtime_status === 'error'
|
||||
? 'text-red-300'
|
||||
@@ -140,6 +140,8 @@ export function JobControls() {
|
||||
? 'Running'
|
||||
: job.runtime_status === 'rate_limited'
|
||||
? 'Paused (rate-limited)'
|
||||
: job.runtime_status === 'deferred'
|
||||
? 'Deferred (retrying)'
|
||||
: job.runtime_status === 'error'
|
||||
? 'Last run error'
|
||||
: job.enabled
|
||||
|
||||
@@ -8,6 +8,7 @@ const DEFAULTS: ScheduleConfig = {
|
||||
schedule_daily_pipeline_cron: '0 2 * * *',
|
||||
schedule_dolt_earnings_cron: '30 2 * * *',
|
||||
schedule_sec_fundamentals_cron: '0 4 * * *',
|
||||
schedule_fundamentals_parity_cron: '30 5 * * *',
|
||||
schedule_near_close_pipeline_cron: '30 15 * * mon-fri',
|
||||
schedule_after_close_pipeline_cron: '45 16 * * mon-fri',
|
||||
schedule_intraday_pipeline_cron: '0 10-15 * * mon-fri',
|
||||
@@ -28,14 +29,20 @@ const FIELDS: { key: keyof ScheduleConfig; label: string; hint: string; mono?: b
|
||||
},
|
||||
{
|
||||
key: 'schedule_dolt_earnings_cron',
|
||||
label: 'Dolt earnings (shadow)',
|
||||
hint: 'Pull and import earnings dates/results daily at 02:30 ET. Live scoring remains untouched before A5.',
|
||||
label: 'Dolt earnings',
|
||||
hint: 'Pull and import earnings dates/results daily at 02:30 ET. The activated cache refresh uses these local events.',
|
||||
mono: true,
|
||||
},
|
||||
{
|
||||
key: 'schedule_sec_fundamentals_cron',
|
||||
label: 'SEC fundamentals (shadow)',
|
||||
hint: 'Import tracked-universe SEC facts daily at 04:00 ET. Unchanged revisions become no-op runs.',
|
||||
label: 'SEC fundamentals',
|
||||
hint: 'Import tracked-universe SEC facts daily at 04:00 ET and refresh the scoring cache when the cutover is active.',
|
||||
mono: true,
|
||||
},
|
||||
{
|
||||
key: 'schedule_fundamentals_parity_cron',
|
||||
label: 'Fundamentals parity report',
|
||||
hint: 'Read-only legacy vs SEC/Dolt comparison daily at 05:30 ET, after the bulk imports.',
|
||||
mono: true,
|
||||
},
|
||||
{
|
||||
@@ -59,7 +66,7 @@ const FIELDS: { key: keyof ScheduleConfig; label: string; hint: string; mono?: b
|
||||
{
|
||||
key: 'schedule_fundamentals_cron',
|
||||
label: 'Legacy fundamentals (weekly)',
|
||||
hint: 'Existing provider chain retained until the A5 parity approval and A6 removal.',
|
||||
hint: 'Fallback provider chain. Automatically skipped while the SEC + Dolt cutover is active.',
|
||||
mono: true,
|
||||
},
|
||||
];
|
||||
|
||||
@@ -3,6 +3,8 @@ import { useSettings, useUpdateSetting } from '../../hooks/useAdmin';
|
||||
import { SkeletonTable } from '../ui/Skeleton';
|
||||
import type { SystemSetting } from '../../lib/types';
|
||||
|
||||
const MANAGED_SETTINGS = new Set(['fundamental_data_sec_dolt_cutover_enabled']);
|
||||
|
||||
export function SettingsForm() {
|
||||
const { data: settings, isLoading, isError, error } = useSettings();
|
||||
const updateSetting = useUpdateSetting();
|
||||
@@ -32,10 +34,11 @@ export function SettingsForm() {
|
||||
if (isLoading) return <SkeletonTable rows={4} cols={2} />;
|
||||
if (isError) return <p className="text-sm text-red-400">{(error as Error)?.message || 'Failed to load settings'}</p>;
|
||||
if (!settings || settings.length === 0) return <p className="text-sm text-gray-500">No settings found.</p>;
|
||||
const visibleSettings = settings.filter((setting) => !MANAGED_SETTINGS.has(setting.key));
|
||||
|
||||
return (
|
||||
<div className="space-y-4">
|
||||
{settings.map((setting) => (
|
||||
{visibleSettings.map((setting) => (
|
||||
<div key={setting.key} className="glass p-4 flex flex-wrap items-center gap-3 glass-hover">
|
||||
<label className="min-w-[140px] text-sm font-medium text-gray-300">{setting.key}</label>
|
||||
{setting.key === 'registration' ? (
|
||||
|
||||
@@ -378,16 +378,15 @@ export function TradeChart({
|
||||
// it wanders left as more post-entry bars arrive.
|
||||
const WINDOW = 21;
|
||||
const MID = 10;
|
||||
let start: number;
|
||||
let entryIdx: number;
|
||||
if (postCount <= MID + 1) {
|
||||
start = Math.max(0, entryAbs - MID);
|
||||
entryIdx = entryAbs - start;
|
||||
} else {
|
||||
const start = postCount <= MID + 1
|
||||
? Math.max(0, entryAbs - MID)
|
||||
// Enough history: keep the latest WINDOW bars; entry falls where it falls.
|
||||
start = Math.max(0, bars.length - WINDOW);
|
||||
entryIdx = entryAbs - start;
|
||||
}
|
||||
: Math.max(0, bars.length - WINDOW);
|
||||
// A trade older than the window entered before the first visible bar. Clamp to
|
||||
// the left edge — a negative index reads past the start of `series`/`stopPath`
|
||||
// and NaNs out the price and trail paths entirely.
|
||||
const entryBeforeWindow = entryAbs < start;
|
||||
const entryIdx = Math.max(0, entryAbs - start);
|
||||
const windowBars = bars.slice(start);
|
||||
const series = windowBars.map((b) => b.close);
|
||||
if (series.length < 2) return null;
|
||||
@@ -601,7 +600,11 @@ export function TradeChart({
|
||||
{entryIdx === lastIdx && (
|
||||
<circle cx={px(entryIdx)} cy={py(series[entryIdx])} r="2" fill="var(--ink-3)" />
|
||||
)}
|
||||
<circle cx={px(entryIdx)} cy={py(entry)} r="3.5" fill="var(--ink-2)" stroke="var(--surface)" strokeWidth="1.5" />
|
||||
{/* Entry marker only when the entry bar is actually in the window — for an
|
||||
older trade the entry level line carries it instead. */}
|
||||
{!entryBeforeWindow && (
|
||||
<circle cx={px(entryIdx)} cy={py(entry)} r="3.5" fill="var(--ink-2)" stroke="var(--surface)" strokeWidth="1.5" />
|
||||
)}
|
||||
<circle cx={px(lastIdx)} cy={py(series[lastIdx])} r="4" fill={nowCol} stroke="var(--surface)" strokeWidth="2" />
|
||||
</svg>
|
||||
);
|
||||
|
||||
@@ -22,6 +22,63 @@ function pnlColor(v: number): string {
|
||||
return 'text-gray-300';
|
||||
}
|
||||
|
||||
function maxHoldText(trade: PaperTrade): string | null {
|
||||
const remaining = trade.sessions_remaining;
|
||||
if (remaining == null) return null;
|
||||
const held = trade.sessions_held ?? 0;
|
||||
if (remaining < 0) return `${held} held · past max hold`;
|
||||
if (remaining === 0) return `${held} held · max hold reached`;
|
||||
return `${held} held · ${remaining} remaining`;
|
||||
}
|
||||
|
||||
function maxHoldColor(trade: PaperTrade): string {
|
||||
const remaining = trade.sessions_remaining;
|
||||
if (remaining == null) return 'text-gray-400';
|
||||
const holdDays = Math.max(1, (trade.sessions_held ?? 0) + remaining);
|
||||
const warningAt = Math.max(1, Math.ceil(holdDays * 0.2));
|
||||
return remaining <= warningAt ? 'text-amber-300' : 'text-gray-400';
|
||||
}
|
||||
|
||||
/** Quiet secondary telemetry below the R bar. Exact timing stays in the
|
||||
* expanded row; this only communicates how far through max hold the trade is. */
|
||||
function HoldProgress({ trade }: { trade: PaperTrade }) {
|
||||
const held = trade.sessions_held;
|
||||
const remaining = trade.sessions_remaining;
|
||||
if (held == null || remaining == null) return null;
|
||||
|
||||
const total = Math.max(1, held + Math.max(0, remaining));
|
||||
const elapsedPct = remaining <= 0
|
||||
? 100
|
||||
: Math.min(100, Math.max(0, (held / total) * 100));
|
||||
const warningAt = Math.max(1, Math.ceil(total * 0.2));
|
||||
const urgent = remaining <= warningAt;
|
||||
const color = urgent ? 'bg-amber-400/75' : 'bg-sky-400/40';
|
||||
|
||||
return (
|
||||
<div
|
||||
className="relative h-[3px] rounded-full bg-white/[0.06]"
|
||||
role="progressbar"
|
||||
aria-label="Holding period"
|
||||
aria-valuemin={0}
|
||||
aria-valuemax={total}
|
||||
aria-valuenow={Math.min(held, total)}
|
||||
aria-valuetext={remaining < 0
|
||||
? `${held} sessions held, past maximum hold`
|
||||
: `${held} sessions held, ${remaining} remaining`}
|
||||
title="Holding-period progress — click for the exact session count"
|
||||
>
|
||||
<span
|
||||
className={`absolute inset-y-0 left-0 rounded-full ${color}`}
|
||||
style={{ width: `${elapsedPct}%` }}
|
||||
/>
|
||||
<span
|
||||
className={`absolute top-1/2 h-[5px] w-[2px] -translate-x-1/2 -translate-y-1/2 rounded-full ${color}`}
|
||||
style={{ left: `${elapsedPct}%` }}
|
||||
/>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
function DirTag({ direction }: { direction: string }) {
|
||||
const isLong = direction === 'long';
|
||||
return (
|
||||
@@ -46,10 +103,22 @@ function Detail({ label, value, valueClass = 'text-gray-100' }: {
|
||||
);
|
||||
}
|
||||
|
||||
function Fact({ label, value, valueClass = 'text-gray-300' }: {
|
||||
label: string;
|
||||
value: ReactNode;
|
||||
valueClass?: string;
|
||||
}) {
|
||||
return (
|
||||
<span className="num inline-flex items-baseline gap-1.5 whitespace-nowrap">
|
||||
<span className="text-[9px] uppercase tracking-[0.14em] text-gray-600">{label}</span>
|
||||
<span className={`text-[11px] ${valueClass}`}>{value}</span>
|
||||
</span>
|
||||
);
|
||||
}
|
||||
|
||||
/** Expanded row: full trade detail + price chart with entry / trail path. */
|
||||
function TradeDetail({ trade, exitLabel, exitMode, atrMultiplier, trailingPct, onClose, closing }: {
|
||||
function TradeDetail({ trade, exitMode, atrMultiplier, trailingPct, onClose, closing }: {
|
||||
trade: PaperTrade;
|
||||
exitLabel: string | null;
|
||||
exitMode: 'time' | 'trailing' | 'atr_trailing' | 'target';
|
||||
atrMultiplier: number;
|
||||
trailingPct: number;
|
||||
@@ -66,30 +135,28 @@ function TradeDetail({ trade, exitLabel, exitMode, atrMultiplier, trailingPct, o
|
||||
staleTime: 5 * 60_000,
|
||||
});
|
||||
const opened = new Date(trade.opened_at).toLocaleDateString('en-US', { month: 'short', day: 'numeric' });
|
||||
const holdText = maxHoldText(trade);
|
||||
const exitRuleText = exitMode === 'atr_trailing'
|
||||
? `${atrMultiplier.toFixed(1)}× ATR trail`
|
||||
: exitMode === 'trailing'
|
||||
? `${Math.round(trailingPct)}% trailing stop`
|
||||
: exitMode === 'target'
|
||||
? 'target / stop'
|
||||
: null;
|
||||
const chartHint = trailMoved || exitMode === 'atr_trailing' || exitMode === 'trailing'
|
||||
? 'entry · now · stop · trail · gate'
|
||||
: 'entry · now · stop · gate';
|
||||
return (
|
||||
<div className="flex flex-col gap-4 px-2 pb-4 pt-1">
|
||||
<dl className="grid grid-cols-2 gap-x-8 gap-y-3 sm:grid-cols-4">
|
||||
<Detail label="opened" value={`${opened} · ${trade.shares} shares`} />
|
||||
<Detail label="entry → now" value={
|
||||
`${formatPrice(trade.entry_price)} → ${trade.current_price != null ? formatPrice(trade.current_price) : '—'}`
|
||||
} />
|
||||
<dl className="grid grid-cols-2 gap-x-8 gap-y-3 md:grid-cols-4 xl:grid-cols-2">
|
||||
<Detail
|
||||
label="P&L"
|
||||
value={p ? `${money(p.pnl)} · ${p.pct >= 0 ? '+' : ''}${p.pct.toFixed(1)}%` : '—'}
|
||||
valueClass={p ? pnlColor(p.pnl) : 'text-gray-500'}
|
||||
/>
|
||||
<Detail
|
||||
label="alpha vs SPY"
|
||||
value={
|
||||
trade.alpha_pct != null
|
||||
? `${trade.alpha_pct >= 0 ? '+' : ''}${trade.alpha_pct.toFixed(1)}%${trade.alpha_usd != null ? ` · ${money(trade.alpha_usd)}` : ''}`
|
||||
: '—'
|
||||
}
|
||||
valueClass={trade.alpha_pct != null ? pnlColor(trade.alpha_pct) : 'text-gray-500'}
|
||||
/>
|
||||
<Detail label="entry → now" value={
|
||||
`${formatPrice(trade.entry_price)} → ${trade.current_price != null ? formatPrice(trade.current_price) : '—'}`
|
||||
} />
|
||||
<Detail
|
||||
label={trailMoved ? 'trail' : 'stop'}
|
||||
value={
|
||||
@@ -105,27 +172,36 @@ function TradeDetail({ trade, exitLabel, exitMode, atrMultiplier, trailingPct, o
|
||||
}
|
||||
/>
|
||||
<Detail
|
||||
label="target"
|
||||
label="alpha vs SPY"
|
||||
value={
|
||||
trade.alpha_pct != null
|
||||
? `${trade.alpha_pct >= 0 ? '+' : ''}${trade.alpha_pct.toFixed(1)}%${trade.alpha_usd != null ? ` · ${money(trade.alpha_usd)}` : ''}`
|
||||
: '—'
|
||||
}
|
||||
valueClass={trade.alpha_pct != null ? pnlColor(trade.alpha_pct) : 'text-gray-500'}
|
||||
/>
|
||||
</dl>
|
||||
<div className="flex flex-wrap items-center gap-x-5 gap-y-2 border-t border-white/[0.06] pt-3">
|
||||
<Fact label="position" value={`${trade.shares} shares`} />
|
||||
<Fact
|
||||
label="holding"
|
||||
value={
|
||||
<>
|
||||
{formatPrice(trade.target)}
|
||||
{exitMode !== 'target' && (
|
||||
<span className="ml-1.5 text-[10px] text-gray-500">screening only</span>
|
||||
)}
|
||||
opened {opened}
|
||||
{holdText && <span className={maxHoldColor(trade)}> · {holdText}</span>}
|
||||
</>
|
||||
}
|
||||
/>
|
||||
<Detail label="exit rule" value={exitLabel ?? 'target/stop'} />
|
||||
<div className="flex items-end">
|
||||
<button
|
||||
onClick={onClose}
|
||||
disabled={closing}
|
||||
className="rounded-md border border-white/[0.1] px-3 py-1.5 text-xs text-gray-300 transition-colors hover:bg-white/[0.06] hover:text-white disabled:opacity-50"
|
||||
>
|
||||
Sell at market
|
||||
</button>
|
||||
</div>
|
||||
</dl>
|
||||
<Fact label="screening target" value={formatPrice(trade.target)} />
|
||||
{exitRuleText && <Fact label="exit" value={exitRuleText} />}
|
||||
<button
|
||||
onClick={onClose}
|
||||
disabled={closing}
|
||||
className="ml-auto rounded-md border border-white/[0.1] px-3 py-1.5 text-[11px] text-gray-300 transition-colors hover:bg-white/[0.06] hover:text-white disabled:opacity-50"
|
||||
>
|
||||
Sell at market
|
||||
</button>
|
||||
</div>
|
||||
{ohlcv.data && (
|
||||
<div>
|
||||
<p className="num text-[9.5px] uppercase tracking-[0.16em] text-gray-500">
|
||||
@@ -173,13 +249,14 @@ export function OpenTradesPanel() {
|
||||
const trailingPct = policy?.trailing_pct ?? 12;
|
||||
const exitLabel = policy
|
||||
? policy.mode === 'atr_trailing'
|
||||
? `${atrMultiplier.toFixed(1)}x ATR trailing stop / ${policy.hold_days}d max`
|
||||
? `${atrMultiplier.toFixed(1)}x ATR trailing stop / ${policy.hold_days} sessions max`
|
||||
: policy.mode === 'trailing'
|
||||
? `trailing ${Math.round(trailingPct)}%`
|
||||
: policy.mode === 'time'
|
||||
? `${policy.hold_days}d hold`
|
||||
? `${policy.hold_days}-session hold`
|
||||
: 'target/stop'
|
||||
: null;
|
||||
const hasMaxHold = exitMode === 'atr_trailing' || exitMode === 'time';
|
||||
|
||||
const rows = trades ?? [];
|
||||
|
||||
@@ -245,7 +322,10 @@ export function OpenTradesPanel() {
|
||||
<span className="num hidden text-xs text-gray-400 sm:block">
|
||||
{formatPrice(t.entry_price)} → {t.current_price != null ? formatPrice(t.current_price) : '—'}
|
||||
</span>
|
||||
<RBar r={p?.r ?? null} max={rMax} />
|
||||
<div className={`min-w-0 ${hasMaxHold ? 'space-y-1.5' : ''}`}>
|
||||
<RBar r={p?.r ?? null} max={rMax} />
|
||||
{hasMaxHold && <HoldProgress trade={t} />}
|
||||
</div>
|
||||
<span className={`num text-right text-[13px] font-semibold ${p?.r != null ? pnlColor(p.r) : 'text-gray-500'}`}>
|
||||
{p?.r != null ? `${p.r >= 0 ? '+' : ''}${p.r.toFixed(2)}R` : '—'}
|
||||
</span>
|
||||
@@ -256,7 +336,6 @@ export function OpenTradesPanel() {
|
||||
{open && (
|
||||
<TradeDetail
|
||||
trade={t}
|
||||
exitLabel={exitLabel}
|
||||
exitMode={exitMode}
|
||||
atrMultiplier={atrMultiplier}
|
||||
trailingPct={trailingPct}
|
||||
|
||||
@@ -138,7 +138,8 @@ export function FundamentalsPanel({ data }: FundamentalsPanelProps) {
|
||||
<div className="mt-2.5 space-y-3.5">
|
||||
{trendRows.map((r) => (
|
||||
<TrendRow key={r.key} label={r.label} kind={r.kind}
|
||||
metric={metrics[r.key]} read={reads[r.key]} />
|
||||
metric={metrics[r.key]} read={reads[r.key]}
|
||||
caveat={metrics[r.key]?.caveat} />
|
||||
))}
|
||||
</div>
|
||||
</div>
|
||||
@@ -188,9 +189,10 @@ function Bullet({ label, value, rail, comparison }: {
|
||||
|
||||
// ---- operating-trend row (delta vs reference, favorable = right) ------------
|
||||
|
||||
function TrendRow({ label, kind, metric, read }: {
|
||||
function TrendRow({ label, kind, metric, read, caveat }: {
|
||||
label: string; kind: 'growth' | 'margin' | 'share';
|
||||
metric: MetricItem | undefined; read: string | null | undefined;
|
||||
caveat: string | null | undefined;
|
||||
}) {
|
||||
const tone = readTone(read);
|
||||
const value = finiteOrNull(metric?.value);
|
||||
@@ -213,7 +215,9 @@ function TrendRow({ label, kind, metric, read }: {
|
||||
}
|
||||
const delta = value != null && ref != null ? value - ref : null;
|
||||
|
||||
const comparison = value == null ? (
|
||||
const comparison = caveat ? (
|
||||
<span style={{ color: HZ.muted }}>{caveat}</span>
|
||||
) : value == null ? (
|
||||
<span style={{ color: HZ.track }}>n/a</span>
|
||||
) : delta == null ? (
|
||||
<span style={{ color: HZ.track }}>history n/a</span>
|
||||
|
||||
@@ -23,11 +23,13 @@ function dateFromToday(days: number): string {
|
||||
}
|
||||
|
||||
function metric(key: string, value: number | null, hist: (number | null)[],
|
||||
industry: MetricItem['industry'] = null): MetricItem {
|
||||
industry: MetricItem['industry'] = null,
|
||||
caveat: string | null = null): MetricItem {
|
||||
return {
|
||||
key: key as MetricItem['key'], value,
|
||||
history: hist.map((v, i) => h(P[i], v)),
|
||||
industry, period_end: '2026-03-28', filed_date: '2026-05-01', source: 'sec',
|
||||
industry, period_end: '2026-03-28', filed_date: '2026-05-01', caveat,
|
||||
source: 'sec',
|
||||
};
|
||||
}
|
||||
const ind = (median: number, favorable_percentile: number) =>
|
||||
@@ -36,6 +38,7 @@ const ind = (median: number, favorable_percentile: number) =>
|
||||
const legacy = {
|
||||
pe_ratio: null, revenue_growth: null, earnings_surprise: null, market_cap: null,
|
||||
next_earnings_date: null, fetched_at: null, unavailable_fields: {},
|
||||
setup_eligible: true, setup_block_code: null, setup_block_reason: null,
|
||||
};
|
||||
|
||||
const full: FundamentalResponse = {
|
||||
@@ -78,12 +81,24 @@ const partial: FundamentalResponse = {
|
||||
earnings: { next: { date: dateFromToday(0), session: 'unknown', days_until: 0 }, recent: [] },
|
||||
metrics: [
|
||||
metric('revenue_growth_yoy', 12, [null, 8, 10, 12], null),
|
||||
metric('eps_growth_yoy', null, [null, null, null, null], null),
|
||||
metric(
|
||||
'eps_growth_yoy',
|
||||
null,
|
||||
[null, null, null, null],
|
||||
null,
|
||||
'Not comparable: share count changed at least 25%; possible split or corporate action.',
|
||||
),
|
||||
metric('operating_margin', 25, [24, 24, 25, 25], null),
|
||||
metric('fcf_margin', null, [null, null, null, null], null),
|
||||
metric('net_debt', null, [], null),
|
||||
metric('net_debt_to_ebitda', 1.9, [1.7, 1.8, 1.9, 1.9], null),
|
||||
metric('share_count_change_yoy', 2.1, [1.8, 2.0, 2.0, 2.1], null),
|
||||
metric(
|
||||
'share_count_change_yoy',
|
||||
null,
|
||||
[1.8, 2.0, 2.0, null],
|
||||
null,
|
||||
'Not comparable: share count changed at least 25%; possible split or corporate action.',
|
||||
),
|
||||
],
|
||||
valuation: {
|
||||
pe: 15.2, fcf_yield: null, market_cap_est: 5.4e8,
|
||||
|
||||
@@ -90,6 +90,36 @@ export function useUpdateSetting() {
|
||||
});
|
||||
}
|
||||
|
||||
export function useFundamentalsCutoverSettings() {
|
||||
return useQuery({
|
||||
queryKey: ['admin', 'fundamentals-cutover'],
|
||||
queryFn: () => adminApi.getFundamentalsCutoverSettings(),
|
||||
});
|
||||
}
|
||||
|
||||
export function useUpdateFundamentalsCutoverSettings() {
|
||||
const qc = useQueryClient();
|
||||
const { addToast } = useToast();
|
||||
|
||||
return useMutation({
|
||||
mutationFn: (enabled: boolean) =>
|
||||
adminApi.updateFundamentalsCutoverSettings(enabled),
|
||||
onSuccess: (config) => {
|
||||
qc.setQueryData(['admin', 'fundamentals-cutover'], config);
|
||||
qc.invalidateQueries({ queryKey: ['admin', 'settings'] });
|
||||
addToast(
|
||||
config.enabled ? 'success' : 'info',
|
||||
config.enabled
|
||||
? 'SEC + Dolt fundamentals activated'
|
||||
: 'SEC + Dolt cache refresh paused',
|
||||
);
|
||||
},
|
||||
onError: (error: Error) => {
|
||||
addToast('error', error.message || 'Failed to update fundamentals data source');
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
export function useRecommendationSettings() {
|
||||
return useQuery({
|
||||
queryKey: ['admin', 'recommendation-settings'],
|
||||
@@ -316,6 +346,14 @@ export function useJobs() {
|
||||
});
|
||||
}
|
||||
|
||||
export function useFundamentalsParityReport() {
|
||||
return useQuery({
|
||||
queryKey: ['admin', 'fundamentals-parity'],
|
||||
queryFn: () => adminApi.getFundamentalsParityReport(),
|
||||
refetchInterval: 15_000,
|
||||
});
|
||||
}
|
||||
|
||||
export function usePipelineReadiness() {
|
||||
return useQuery({
|
||||
queryKey: ['admin', 'pipeline-readiness'],
|
||||
|
||||
@@ -187,12 +187,17 @@ export interface ActivationConfig {
|
||||
exclude_neutral: boolean;
|
||||
}
|
||||
|
||||
export interface FundamentalsCutoverConfig {
|
||||
enabled: boolean;
|
||||
}
|
||||
|
||||
// Cron schedule for morning / near-close / after-close / intraday + fundamentals
|
||||
export interface ScheduleConfig {
|
||||
schedule_timezone: string;
|
||||
schedule_daily_pipeline_cron: string;
|
||||
schedule_dolt_earnings_cron: string;
|
||||
schedule_sec_fundamentals_cron: string;
|
||||
schedule_fundamentals_parity_cron: string;
|
||||
schedule_near_close_pipeline_cron: string;
|
||||
schedule_after_close_pipeline_cron: string;
|
||||
schedule_intraday_pipeline_cron: string;
|
||||
@@ -232,6 +237,8 @@ export interface PaperTrade {
|
||||
close_reason: 'time' | 'trailing' | 'stop' | 'target' | 'manual' | null;
|
||||
trailing_stop: number | null;
|
||||
trailing_distance_pct: number | null;
|
||||
sessions_held: number | null;
|
||||
sessions_remaining: number | null;
|
||||
}
|
||||
|
||||
export interface ExitPolicy {
|
||||
@@ -475,6 +482,9 @@ export interface RegimePillar {
|
||||
export interface RegimeReading {
|
||||
score: number | null;
|
||||
band: RegimeBand | null;
|
||||
// Per axis: State and Warning have different realized ranges, so they do not
|
||||
// share thresholds.
|
||||
bands?: { watch: number; elevated: number; breaking: number };
|
||||
coverage: number;
|
||||
minimum_coverage: number;
|
||||
available_pillars: string[];
|
||||
@@ -482,6 +492,23 @@ export interface RegimeReading {
|
||||
trend?: { delta_7: number | null; delta_30: number | null };
|
||||
}
|
||||
|
||||
/** Qualitative capex / earnings-reaction context. Not part of either score. */
|
||||
export interface RegimeFundamentalOverlay {
|
||||
available: boolean;
|
||||
pending: boolean;
|
||||
stale: boolean;
|
||||
effective_date: string | null;
|
||||
age_days: number | null;
|
||||
capex: Record<string, CapexState> | null;
|
||||
good_news_stock_down: GoodNewsReaction | null;
|
||||
capex_stress: number | null;
|
||||
earnings_stress: number | null;
|
||||
reasoning: string | null;
|
||||
source: string | null;
|
||||
fetched_at: string | null;
|
||||
observed_in_snapshot?: boolean;
|
||||
}
|
||||
|
||||
export interface RegimeHistoryPoint {
|
||||
date: string;
|
||||
state: number | null;
|
||||
@@ -498,6 +525,10 @@ export interface RegimeMonitor {
|
||||
date?: string;
|
||||
state?: RegimeReading;
|
||||
warning?: RegimeReading;
|
||||
/** Point-in-time overlay recorded in the snapshot. */
|
||||
fundamental_overlay?: RegimeFundamentalOverlay;
|
||||
/** Current observation, even when it is not effective until the next session. */
|
||||
fundamental_context?: RegimeFundamentalOverlay;
|
||||
inputs?: {
|
||||
vix: number | null;
|
||||
vix_date: string | null;
|
||||
@@ -529,7 +560,7 @@ export interface RegimeMonitor {
|
||||
}
|
||||
|
||||
export interface RegimeFundamentals {
|
||||
methodology: 'v2';
|
||||
methodology: 'v3';
|
||||
f1_score: number | null;
|
||||
f3_score: number | null;
|
||||
locked: boolean;
|
||||
@@ -575,6 +606,18 @@ export interface EventStudyReport {
|
||||
warn_threshold: number;
|
||||
basket_hash: string;
|
||||
basket_asof: string;
|
||||
credit_sensor_from?: string | null;
|
||||
};
|
||||
/** How far the headline metrics can be trusted. See _reliability(). */
|
||||
reliability?: {
|
||||
events_detected: number;
|
||||
events_in_holdout: number;
|
||||
minimum_events: number;
|
||||
underpowered: boolean;
|
||||
sensors_expected: number;
|
||||
train_full_sensor_share: number;
|
||||
holdout_full_sensor_share: number;
|
||||
sensor_coverage_mismatch: boolean;
|
||||
};
|
||||
sample?: {
|
||||
start: string;
|
||||
@@ -733,6 +776,7 @@ export interface MetricItem {
|
||||
industry: MetricIndustry | null;
|
||||
period_end: string | null;
|
||||
filed_date: string | null;
|
||||
caveat: string | null;
|
||||
source: string; // 'sec' | 'legacy_api'
|
||||
}
|
||||
|
||||
@@ -785,6 +829,9 @@ export interface FundamentalResponse {
|
||||
metrics: MetricItem[] | null;
|
||||
valuation: Valuation | null;
|
||||
reads: FundamentalsReads | null;
|
||||
setup_eligible: boolean;
|
||||
setup_block_code: string | null;
|
||||
setup_block_reason: string | null;
|
||||
}
|
||||
|
||||
// Indicators
|
||||
|
||||
@@ -5,6 +5,8 @@ import { AlertSettings } from '../components/admin/AlertSettings';
|
||||
import { SentimentProviderSettings } from '../components/admin/SentimentProviderSettings';
|
||||
import { DataCleanup } from '../components/admin/DataCleanup';
|
||||
import { JobControls } from '../components/admin/JobControls';
|
||||
import { FundamentalsParityPanel } from '../components/admin/FundamentalsParityPanel';
|
||||
import { FundamentalsCutoverSettings } from '../components/admin/FundamentalsCutoverSettings';
|
||||
import { PerformanceSettings } from '../components/admin/PerformanceSettings';
|
||||
import { PipelineReadinessPanel } from '../components/admin/PipelineReadinessPanel';
|
||||
import { SystemEventsPanel } from '../components/admin/SystemEventsPanel';
|
||||
@@ -35,6 +37,7 @@ export default function AdminPage() {
|
||||
{activeTab === 'Tickers' && <TickerManagement />}
|
||||
{activeTab === 'Settings' && (
|
||||
<div className="space-y-4">
|
||||
<FundamentalsCutoverSettings />
|
||||
<ActivationSettings />
|
||||
<ExitPolicySettings />
|
||||
<PerformanceSettings />
|
||||
@@ -48,6 +51,7 @@ export default function AdminPage() {
|
||||
{activeTab === 'Jobs' && (
|
||||
<div className="space-y-4">
|
||||
<ScheduleSettings />
|
||||
<FundamentalsParityPanel />
|
||||
<JobControls />
|
||||
<PipelineReadinessPanel />
|
||||
</div>
|
||||
|
||||
@@ -21,6 +21,7 @@ import type {
|
||||
GoodNewsReaction,
|
||||
RegimeBand,
|
||||
RegimeConfig,
|
||||
RegimeFundamentalOverlay,
|
||||
RegimeFundamentals,
|
||||
RegimeFundamentalsUpdate,
|
||||
RegimeReading,
|
||||
@@ -64,6 +65,8 @@ function ScoreGauge({
|
||||
const complete = reading?.band != null;
|
||||
const style = complete ? BAND_STYLES[reading.band as RegimeBand] : null;
|
||||
const position = Math.min(100, Math.max(0, score ?? 0));
|
||||
const bands = reading?.bands;
|
||||
const ticks = bands ? [bands.watch, bands.elevated, bands.breaking] : [30, 60, 80];
|
||||
return (
|
||||
<div className={`glass border p-6 ${style?.ring ?? 'border-white/[0.06]'}`}>
|
||||
<div className="flex flex-wrap items-end justify-between gap-3">
|
||||
@@ -98,8 +101,15 @@ function ScoreGauge({
|
||||
style={{ left: `${position}%` }}
|
||||
/>
|
||||
</div>
|
||||
<div className="mt-1.5 flex justify-between text-[10px] uppercase tracking-wider text-gray-600">
|
||||
<span>0</span><span>30</span><span>60</span><span>80</span><span>100</span>
|
||||
{/* Thresholds come from the reading: the two axes no longer share them. */}
|
||||
<div className="relative mt-1.5 h-4 text-[10px] uppercase tracking-wider text-gray-600">
|
||||
<span className="absolute left-0">0</span>
|
||||
{ticks.map((tick) => (
|
||||
<span key={tick} className="absolute -translate-x-1/2 num" style={{ left: `${tick}%` }}>
|
||||
{tick}
|
||||
</span>
|
||||
))}
|
||||
<span className="absolute right-0">100</span>
|
||||
</div>
|
||||
</>
|
||||
)}
|
||||
@@ -108,6 +118,77 @@ function ScoreGauge({
|
||||
);
|
||||
}
|
||||
|
||||
const CAPEX_TONE: Record<CapexState, string> = {
|
||||
raising: 'text-emerald-400',
|
||||
holding: 'text-amber-400',
|
||||
cutting: 'text-red-400',
|
||||
unknown: 'text-gray-500',
|
||||
};
|
||||
|
||||
function FundamentalOverlayCard({ overlay }: { overlay: RegimeFundamentalOverlay }) {
|
||||
const capex = overlay.capex ?? {};
|
||||
const reaction = overlay.good_news_stock_down;
|
||||
return (
|
||||
<div className="glass border border-white/[0.06] p-5">
|
||||
<div className="flex flex-wrap items-baseline justify-between gap-2">
|
||||
<div className="text-[11px] uppercase tracking-wider text-gray-500">
|
||||
Fundamental overlay · context, not scored
|
||||
</div>
|
||||
<div className="flex flex-wrap items-center gap-2 text-[11px] text-gray-500">
|
||||
{overlay.source && <span>{overlay.source}</span>}
|
||||
{overlay.effective_date && <span>· effective {overlay.effective_date}</span>}
|
||||
{overlay.pending && <Badge label="pending" variant="manual" />}
|
||||
{overlay.stale && <Badge label="stale" variant="manual" />}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{overlay.pending ? (
|
||||
<p className="mt-3 text-xs leading-relaxed text-amber-400/90">
|
||||
A newer observation was collected but is not effective until {overlay.effective_date ?? 'the next session'}.
|
||||
Observations are never backdated, so the reading below appears from that session onward.
|
||||
</p>
|
||||
) : (
|
||||
<>
|
||||
<div className="mt-4 grid gap-4 sm:grid-cols-2">
|
||||
<div>
|
||||
<div className="mb-2 flex items-baseline justify-between text-xs">
|
||||
<span className="font-medium text-gray-300">Hyperscaler capex guidance</span>
|
||||
<span className="num text-gray-500">{overlay.capex_stress ?? 'n/a'}</span>
|
||||
</div>
|
||||
<div className="space-y-1">
|
||||
{Object.entries(capex).map(([symbol, state]) => (
|
||||
<div key={symbol} className="flex items-center justify-between text-xs">
|
||||
<span className="font-mono text-gray-400">{symbol}</span>
|
||||
<span className={CAPEX_TONE[state] ?? 'text-gray-500'}>{state}</span>
|
||||
</div>
|
||||
))}
|
||||
</div>
|
||||
</div>
|
||||
<div>
|
||||
<div className="mb-2 flex items-baseline justify-between text-xs">
|
||||
<span className="font-medium text-gray-300">Good news, stock down</span>
|
||||
<span className="num text-gray-500">{overlay.earnings_stress ?? 'n/a'}</span>
|
||||
</div>
|
||||
<div className={`text-sm font-medium ${reaction === 'yes' ? 'text-red-400' : reaction === 'no' ? 'text-emerald-400' : 'text-gray-500'}`}>
|
||||
{reaction === 'yes' ? 'Yes — beats sold into' : reaction === 'no' ? 'No — ordinary reactions' : 'Mixed'}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
{overlay.reasoning && (
|
||||
<p className="mt-4 text-xs leading-relaxed text-gray-400">{overlay.reasoning}</p>
|
||||
)}
|
||||
</>
|
||||
)}
|
||||
|
||||
<p className="mt-4 text-[11px] leading-relaxed text-gray-600">
|
||||
These observations are qualitative, refreshed roughly quarterly, and deliberately excluded from State and
|
||||
Warning. In v2 they carried 20 of 100 Warning points — not enough to cross the study's alarm threshold even
|
||||
when both were pegged — so they are reported here rather than diluted into a daily score.
|
||||
</p>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
function PillarBreakdown({ title, reading }: { title: string; reading: RegimeReading }) {
|
||||
return (
|
||||
<Disclosure summary={`${title} pillars · ${Math.round(reading.coverage)}% coverage`}>
|
||||
@@ -190,6 +271,32 @@ function EventStudyBody({ report }: { report: EventStudyReport }) {
|
||||
</table>
|
||||
</div>
|
||||
)}
|
||||
{report.reliability && (report.reliability.underpowered || report.reliability.sensor_coverage_mismatch) && (
|
||||
<Callout variant="warning">
|
||||
<div className="space-y-1.5">
|
||||
{report.reliability.underpowered && (
|
||||
<p>
|
||||
<strong>Underpowered.</strong> Only {report.reliability.events_in_holdout} of{' '}
|
||||
{report.reliability.events_detected} detected corrections fall in the test period (
|
||||
{report.reliability.minimum_events}+ needed). Recall is one event away from a materially
|
||||
different headline, and which events flip is usually decided by where the frozen threshold
|
||||
lands rather than by what the score saw. Read the direction, not the ratio.
|
||||
</p>
|
||||
)}
|
||||
{report.reliability.sensor_coverage_mismatch && (
|
||||
<p>
|
||||
<strong>Sensor coverage differs across the split.</strong>{' '}
|
||||
{report.reliability.train_full_sensor_share}% of training sessions had all{' '}
|
||||
{report.reliability.sensors_expected} Warning sensors versus{' '}
|
||||
{report.reliability.holdout_full_sensor_share}% of test sessions
|
||||
{report.params?.credit_sensor_from && ` — credit history begins ${report.params.credit_sensor_from}`}
|
||||
. The score renormalises over what is available, so the threshold was frozen on a partly
|
||||
different construct than it is measured against.
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
</Callout>
|
||||
)}
|
||||
<p className="text-[11px] leading-relaxed text-gray-600">
|
||||
The threshold is frozen on the training period and measured on the chronological test period. Reconstructed
|
||||
pre-freeze basket history remains exploratory.
|
||||
@@ -235,8 +342,10 @@ function FundamentalsEditor({
|
||||
const [capex, setCapex] = useState<Record<string, CapexState>>(() => ({ ...data.capex }));
|
||||
const [reaction, setReaction] = useState<GoodNewsReaction>(data.good_news_stock_down);
|
||||
const knownCapex = Object.values(capex).filter((state) => state !== 'unknown');
|
||||
const cutting = knownCapex.filter((state) => state === 'cutting').length;
|
||||
const derivedF1 = knownCapex.length >= 3 ? Math.round((cutting / knownCapex.length) * 1000) / 10 : null;
|
||||
// Mirrors _CAPEX_STATE_SCORES: raising 0, holding 50, cutting 100. Holding is
|
||||
// the deceleration case and used to score identically to raising.
|
||||
const capexPoints = knownCapex.reduce((sum, state) => sum + (state === 'cutting' ? 100 : state === 'holding' ? 50 : 0), 0);
|
||||
const derivedF1 = knownCapex.length >= 3 ? Math.round((capexPoints / knownCapex.length) * 10) / 10 : null;
|
||||
const derivedF3 = reaction === 'yes' ? 100 : reaction === 'no' ? 0 : null;
|
||||
return (
|
||||
<div className="space-y-4">
|
||||
@@ -266,7 +375,7 @@ function FundamentalsEditor({
|
||||
</label>
|
||||
))}
|
||||
</div>
|
||||
<p className="mt-1.5 text-[11px] text-gray-600">Raising/holding = 0 stress; cutting = 100; at least three known names required.</p>
|
||||
<p className="mt-1.5 text-[11px] text-gray-600">Raising = 0, holding = 50, cutting = 100; at least three known names required. Display only — this does not enter Warning.</p>
|
||||
</div>
|
||||
<label className="flex items-center justify-between gap-3 text-xs text-gray-400">
|
||||
<span>
|
||||
@@ -369,9 +478,10 @@ export default function RegimePage() {
|
||||
label="Warning · deterioration & divergence"
|
||||
reading={data.warning}
|
||||
divider={data.quadrant_config?.warning_divider}
|
||||
footnote={<>Breadth divergence, SMH/SPY rollover, and point-in-time fundamental observations. Unknown or stale fundamentals reduce coverage; they never default to 50.</>}
|
||||
footnote={<>Breadth divergence, SMH/SPY rollover, and HY credit impulse. Breadth loss counts fully when price masks it and partially when price confirms it. Missing sensors reduce coverage; they never default to 50.</>}
|
||||
/>
|
||||
</div>
|
||||
{data.fundamental_context && <FundamentalOverlayCard overlay={data.fundamental_context} />}
|
||||
<p className="text-xs text-gray-600">
|
||||
Data quality · oldest market input:{' '}
|
||||
{data.data_quality?.oldest_market_input_age_days == null
|
||||
|
||||
@@ -64,10 +64,44 @@ function timeAgo(iso: string): string {
|
||||
return `${days}d ago`;
|
||||
}
|
||||
|
||||
function marketDate(date = new Date()): string {
|
||||
const parts = new Intl.DateTimeFormat('en-US', {
|
||||
timeZone: 'America/New_York',
|
||||
year: 'numeric',
|
||||
month: '2-digit',
|
||||
day: '2-digit',
|
||||
}).formatToParts(date);
|
||||
const value = (type: Intl.DateTimeFormatPartTypes) =>
|
||||
parts.find((part) => part.type === type)?.value ?? '';
|
||||
return value('year') + '-' + value('month') + '-' + value('day');
|
||||
}
|
||||
|
||||
function formatSessionDate(isoDate: string): string {
|
||||
const currentMarketDate = marketDate();
|
||||
if (isoDate === currentMarketDate) return 'Today';
|
||||
|
||||
// Parse date-only market sessions explicitly. Parsing YYYY-MM-DD directly as
|
||||
// a Date means midnight UTC and makes today's bar look many hours old.
|
||||
const [year, month, day] = isoDate.split('-').map(Number);
|
||||
if (!year || !month || !day) return isoDate;
|
||||
return new Intl.DateTimeFormat(undefined, {
|
||||
month: 'short',
|
||||
day: 'numeric',
|
||||
year: year === Number(currentMarketDate.slice(0, 4)) ? undefined : 'numeric',
|
||||
timeZone: 'UTC',
|
||||
}).format(new Date(Date.UTC(year, month - 1, day)));
|
||||
}
|
||||
|
||||
function formatOHLCVFreshness(sessionDate: string, updatedAt?: string | null): string {
|
||||
const session = formatSessionDate(sessionDate);
|
||||
return updatedAt ? session + ' · updated ' + timeAgo(updatedAt) : session;
|
||||
}
|
||||
|
||||
interface DataStatusItem {
|
||||
label: string;
|
||||
available: boolean;
|
||||
timestamp?: string | null;
|
||||
timestampLabel?: string | null;
|
||||
selector: FetchSelector; // what a refresh of this row fetches
|
||||
paid?: boolean; // provider call that may cost money/quota
|
||||
}
|
||||
@@ -100,7 +134,7 @@ function DataFreshnessBar({
|
||||
}`} />
|
||||
<span className="text-xs text-gray-400">{item.label}</span>
|
||||
{item.available && item.timestamp ? (
|
||||
<span className="text-[10px] text-gray-500">{timeAgo(item.timestamp)}</span>
|
||||
<span className="text-[10px] text-gray-500">{item.timestampLabel ?? timeAgo(item.timestamp)}</span>
|
||||
) : !item.available ? (
|
||||
<span className="text-[10px] text-gray-600">no data</span>
|
||||
) : null}
|
||||
@@ -171,10 +205,16 @@ export default function TickerDetailPage() {
|
||||
const dataStatus: DataStatusItem[] = useMemo(() => [
|
||||
{
|
||||
label: 'OHLCV',
|
||||
// Market age of the latest bar (session date), not DB insert time —
|
||||
// created_at stays frozen when the provider returns no new sessions.
|
||||
// Keep the market session date distinct from the last successful bar
|
||||
// write; treating YYYY-MM-DD as an instant makes today's session look old.
|
||||
available: !!ohlcv.data && ohlcv.data.length > 0,
|
||||
timestamp: ohlcv.data?.[ohlcv.data.length - 1]?.date,
|
||||
timestampLabel: ohlcv.data?.length
|
||||
? formatOHLCVFreshness(
|
||||
ohlcv.data[ohlcv.data.length - 1].date,
|
||||
ohlcv.data[ohlcv.data.length - 1].created_at,
|
||||
)
|
||||
: null,
|
||||
selector: ['ohlcv'] as FetchSelector,
|
||||
paid: true,
|
||||
},
|
||||
@@ -319,6 +359,15 @@ export default function TickerDetailPage() {
|
||||
busy={ingestion.isPending}
|
||||
/>
|
||||
</div>
|
||||
{fundamentals.data && !fundamentals.data.setup_eligible && (
|
||||
<div className="border-b border-white/[0.06] px-6 py-3 sm:px-7">
|
||||
<Callout variant="warning">
|
||||
<span className="font-medium">New setups paused.</span>{' '}
|
||||
{fundamentals.data.setup_block_reason ??
|
||||
'SEC fundamentals are incomplete for this ticker.'}
|
||||
</Callout>
|
||||
</div>
|
||||
)}
|
||||
<div className="p-6 pb-5 sm:p-7 sm:pb-5">
|
||||
<div className="flex flex-wrap items-start justify-between gap-x-8 gap-y-5">
|
||||
<div className="min-w-0">
|
||||
|
||||
@@ -0,0 +1,925 @@
|
||||
# A5 parity report — root-cause findings
|
||||
|
||||
Investigation of `fundamentals-parity-20260723T210658161480Z.json` (511 tickers,
|
||||
generated 2026-07-23). Method: replayed the production parser
|
||||
(`sec_facts_parser.parse_snapshots`) and derivation (`fundamentals_derivation.derive`)
|
||||
against **live SEC companyfacts**, using the importer's own `_filing_meta` and
|
||||
backfill accession set, then cross-checked prices against IBKR. No database was
|
||||
available locally, so every conclusion below is reproduced from source data rather
|
||||
than read out of prod.
|
||||
|
||||
Repro script: `scratchpad/diag.py` (`--history` replays the full backfill path).
|
||||
Every claim below was verified on the named issuer. Names that were *not*
|
||||
individually inspected are listed as unclassified — an earlier draft of this
|
||||
document guessed their cause from fiscal-year-end dates and was wrong for most of
|
||||
them, so the guessing is not repeated here.
|
||||
|
||||
## Verdict
|
||||
|
||||
Where both sides have a value the candidate data is good: P/E spearman 0.968,
|
||||
revenue growth agreeing to 4 decimals for most names, score spearman 0.825. Every
|
||||
defect found is a **parser/derivation bug or an identity problem** — not a data
|
||||
quality problem with SEC or Dolt. The largest cluster is period identity, which is
|
||||
exactly the risk A3 flagged as primary.
|
||||
|
||||
## 1. P/E outliers — splits corrupt TTM EPS, and the split guard doesn't cover it
|
||||
|
||||
`derive()` sets `result.ttm_diluted_eps` at `fundamentals_derivation.py:88` and only
|
||||
calls `_guard_split_sensitive_metrics()` at line 105, which annotates `result.metrics`
|
||||
(the `MetricSeries` objects). `ttm_diluted_eps` is a bare scalar and is never guarded.
|
||||
`fundamentals_parity_service._pe()` consumes it directly.
|
||||
|
||||
The cleanest evidence that the *candidate* side is the broken one: reconcile each
|
||||
P/E against the report's own price. Legacy comes out sane in both cases, candidate
|
||||
does not.
|
||||
|
||||
**BKNG — guard fired, nobody listened.** Share count jumps 31.7M → 774.9M between the
|
||||
FY2025 10-K and the 2026 Q1 10-Q (≈25:1 split). TTM EPS therefore sums three pre-split
|
||||
quarters (27.31 + 84.01 + 44.18 = 155.50) plus one post-split quarter (1.36) =
|
||||
**156.86** — mixed units. Live price $172.83 matches the price the report implies
|
||||
exactly (1.1018 × 156.86 = 172.83), so the price is correct and current. Against that
|
||||
price, legacy's P/E of 22.44 implies EPS ≈ 7.70 — a coherent post-split number, versus
|
||||
the candidate's 156.86. The derivation *did* raise `"Not comparable: share count
|
||||
changed at least 25%; possible split or corporate action."` on `eps_growth_yoy` and
|
||||
`share_count_change_yoy` — P/E never sees it.
|
||||
|
||||
**KLAC — the guard cannot fire.** The split post-dates the most recent 10-Q (period end
|
||||
2026-03-31), so no snapshot shows any share-count change (`share_count_change_yoy` =
|
||||
−1.2%). TTM EPS **35.31** is internally consistent and entirely pre-split; the price
|
||||
($223.30 live, ≈218.7 in the report) is post-split. Reconciling: legacy P/E 60.21
|
||||
against the report price implies EPS ≈ 3.63 ≈ 35.31/9.7 — i.e. legacy is consistent
|
||||
with a ~10:1 split and correct, and the candidate is off by exactly the split factor.
|
||||
(IBKR's split-adjusted `open_52w` of 89.36 corroborates 10:1.)
|
||||
|
||||
This is the important case: **a split after the latest filing is undetectable from
|
||||
snapshots alone.** No share-count test can catch it. Reconciliation needs a corporate
|
||||
actions source or a price-vs-EPS plausibility check.
|
||||
|
||||
**COF — not a bug, a definition difference.** Shares 383M → 639M in 2025 Q2 is the
|
||||
Discover acquisition. TTM GAAP EPS is genuinely $3.92 because the merger-charge quarter
|
||||
(−10.19) sits in the window. Candidate P/E 51.01 is arithmetically correct on a GAAP TTM
|
||||
basis; legacy's 11.61 is an adjusted/forward convention. Disclose, don't fix. Note this
|
||||
single row drives the report's largest change (rank 1 → 456).
|
||||
|
||||
## 2. Bank revenue growth — concept-mapping gap (confirmed)
|
||||
|
||||
`sec_facts_parser._DURATION_USD["revenue"]` is:
|
||||
|
||||
```
|
||||
RevenueFromContractWithCustomerExcludingAssessedTax, Revenues, SalesRevenueNet
|
||||
```
|
||||
|
||||
Banks tag **`RevenuesNetOfInterestExpense`** in their 10-Qs:
|
||||
|
||||
| filer | 2026 Q1 10-Q tags present | parsed `revenue` |
|
||||
|---|---|---|
|
||||
| JPM | `RevenuesNetOfInterestExpense` 49,836M, `NoninterestIncome`, `InterestIncomeExpenseNet` | **null** |
|
||||
| GS | `RevenuesNetOfInterestExpense` 17,227M, `InterestAndDividendIncomeOperating`, … | **null** |
|
||||
| WFC | `RevenuesNetOfInterestExpense` 21,436M, … | **null** |
|
||||
|
||||
JPM's FY2025 10-K *also* tags `Revenues` (182,447M — identical value), so only the annual
|
||||
row populates; GS never tags `Revenues` at all. Revenue growth needs five consecutive
|
||||
quarterly values, so it is null for the whole cluster (JPM, GS, MS, WFC, TFC, MTB, FITB,
|
||||
RF, SYF, BNY, BX, BLK, SPGI, ACGL, CBOE).
|
||||
|
||||
A second variant of the same gap: **ARE** and **KHC** tag
|
||||
`RevenueFromContractWithCustomer**Including**AssessedTax` — also absent from the list —
|
||||
so revenue is null on every row while EPS parses fine.
|
||||
|
||||
**Fix:** add `RevenuesNetOfInterestExpense` and the `IncludingAssessedTax` variant.
|
||||
|
||||
**Latent risk while you're in there:** `RevenueFromContractWithCustomerExcludingAssessedTax`
|
||||
is *first* and "first present wins". For a bank that tags it, it captures only ASC-606 fee
|
||||
revenue, not total revenue — a silently **understated** number rather than a null, which is
|
||||
worse. DVN shows the same hazard from the other side: its 2026 Q1 tags both
|
||||
`RevenueFromContractWithCustomerExcludingAssessedTax` (4,508M) and `Revenues` (3,807M),
|
||||
an 18% difference decided purely by list order.
|
||||
|
||||
## 3. Period identity — the largest cluster, three confirmed mechanisms
|
||||
|
||||
### 3a. Fiscal-year label collisions (CRM, FRT, STX)
|
||||
|
||||
`_fiscal_context` majority-votes SEC's `fy`/`fp` fields, and `_select_latest_per_period`
|
||||
keys on `(fiscal_year, fiscal_period)`. When SEC's labels disagree with the calendar, two
|
||||
distinct periods collide on one key and **one is silently discarded**:
|
||||
|
||||
- **CRM** — two rows keyed `2025 FY`, ending 2025-01-31 and 2026-01-31.
|
||||
- **FRT** — two rows keyed `2024 FY`, ending 2024-12-31 and 2025-12-31.
|
||||
- **STX** — the year ending 2025-06-27 is labelled **`2027 FY`**, so it sorts *after*
|
||||
`2026 Q3` (period end 2026-04-03) and is taken as the latest quarter.
|
||||
|
||||
The survivor's `period_end` then contradicts the fiscal ordering, Q4 derivation and the
|
||||
consecutive-quarter chain break, and TTM EPS + YoY both go null.
|
||||
|
||||
**FRT is a calendar-year (Dec) filer**, so this is *not* limited to non-calendar fiscal
|
||||
years — the earlier assumption that it was is wrong. Any filer SEC labels inconsistently
|
||||
is exposed.
|
||||
|
||||
### 3b. Amendment selection blanks a period (DVN)
|
||||
|
||||
DVN has two rows for `2025 FY` (both ending 2025-12-31): the 10-K with complete financials,
|
||||
and a **10-K/A carrying no financial facts at the report date** (`rev=None eps=None`).
|
||||
`_select_latest_per_period` takes the newest `accepted_at`, so **the empty amendment wins**
|
||||
and the FY2025 row becomes all-null, breaking the chain.
|
||||
|
||||
This is the most dangerous of the three: it is not exotic. Any issuer filing a 10-K/A —
|
||||
including routine Part III amendments that restate nothing — silently loses that period.
|
||||
The rule needs to prefer the newest accession *that actually carries the fact*, per field,
|
||||
rather than the newest accession outright.
|
||||
|
||||
### 3c. 4-4-5 retail calendar — Q3 only, misses by ~2 days (COST, PEP)
|
||||
|
||||
`_EXPECTED_YTD_DAYS["Q3"] = 273` with `_YTD_TOLERANCE_DAYS = 20` accepts 253–293 days. A
|
||||
12/12/12/16-week filer's YTD-Q3 is 36 weeks ≈ **251–252 days** — just under the floor.
|
||||
|
||||
Confirmed, facts present and rejected:
|
||||
|
||||
- COST 2026 Q3: `RevenueFromContractWithCustomerExcludingAssessedTax` span=**251d**
|
||||
val=207,431M, `EarningsPerShareDiluted` span=251d val=14.01 → row stored with
|
||||
`rev=None eps=None start=None`. Same for 2025 Q3 and 2024 Q3.
|
||||
- PEP: every Q3 row is `rev=None eps=None`; Q1/Q2/FY all populate.
|
||||
|
||||
Q1 (83d vs 91±20), Q2 (167d vs 182±20) and FY (363–364d vs 365±20) all pass — only Q3
|
||||
fails, every year. The code comment claims the tolerance "covers 52/53-week fiscal
|
||||
calendars"; it does not cover 4-4-5 ones.
|
||||
|
||||
Note this does **not** apply to ordinary 13-week 52/53-week filers (STX's Q3 YTD is 279d and
|
||||
passes) — their failures are 3a, not this.
|
||||
|
||||
**Fix:** widen the Q3 tolerance to ~25 days, or derive the expected span from the filer's own
|
||||
fiscal calendar rather than a fixed 91/182/273.
|
||||
|
||||
## 4. CIK identity (XOM)
|
||||
|
||||
SEC's `company_tickers.json` now maps **XOM → CIK 2115436 "ExxonMobil Holdings Corp", which
|
||||
has 0 filings**. All 26 XBRL 10-K/10-Qs sit under the old CIK **34088 "EXXON MOBIL CORP"**.
|
||||
XOM therefore has no snapshots at all, and nothing in the pipeline notices that a tracked
|
||||
issuer resolved to a CIK with zero filings.
|
||||
|
||||
PSKY (5 filings) and Q (3 filings) are genuinely new registrants — expected, not a bug.
|
||||
|
||||
## Status of the 25 names that lose their fundamental score
|
||||
|
||||
Production requires ≥2 metrics (`scoring_service.py:502`), the same rule the parity harness
|
||||
uses, so these genuinely drop the fundamental dimension and the composite renormalises over
|
||||
the remaining four.
|
||||
|
||||
| cause (confirmed on the named issuer) | names |
|
||||
|---|---|
|
||||
| FY label collision (3a) | CRM, FRT, STX |
|
||||
| 4-4-5 Q3 span (3c) | COST, PEP |
|
||||
| revenue concept gap (§2) | ARE, KHC |
|
||||
| amendment blanks period (3b) | DVN |
|
||||
| CIK identity (§4) | XOM |
|
||||
| new registrant — expected | PSKY, Q |
|
||||
| **not yet classified** | AZO, BXP, CRWD, FCX, HAL, MOS, MTD, NTAP, PPL, REG, SJM, SWKS, WDAY |
|
||||
|
||||
13 of 25 confirmed. The unclassified 13 have not been inspected and should not be assumed to
|
||||
share a cause — the confirmed set already spans five distinct mechanisms.
|
||||
|
||||
## Recommended order of work
|
||||
|
||||
1. **Amendment selection (3b)** — highest blast radius, affects any 10-K/A filer, and the
|
||||
current rule is wrong in principle rather than at the margin.
|
||||
2. **Revenue concept list (§2)** — add `RevenuesNetOfInterestExpense` and
|
||||
`IncludingAssessedTax`; audit the ASC-606-first priority, which can understate rather
|
||||
than null.
|
||||
3. **Q3 YTD span tolerance (3c)** — effectively one line.
|
||||
4. **XOM CIK remap (§4)** — plus a validation that flags any tracked ticker resolving to a
|
||||
CIK with zero XBRL filings.
|
||||
5. **Split safety for `ttm_diluted_eps` (§1)** — propagate the existing guard to the scalar,
|
||||
and add a price-vs-EPS plausibility check for splits that post-date the last filing.
|
||||
6. **Fiscal-period identity (3a)** — the deepest fix; consider keying period identity on
|
||||
`period_end` rather than SEC's `fy`/`fp`.
|
||||
|
||||
Re-run the parity report after these and re-classify the remaining 13 before making a
|
||||
cutover decision. The current report should not be approved as-is: its coverage gaps are
|
||||
artifacts of the above, not real absences in the source data.
|
||||
|
||||
---
|
||||
|
||||
# Fixes applied (items 1–3)
|
||||
|
||||
| # | change | file | effective |
|
||||
|---|---|---|---|
|
||||
| 1 | amendment resolution is now **per field** — newest accession that actually reports a fact wins; only rows sharing the newest `period_end` are merged, so a mislabelled filing is never blended in | `fundamentals_derivation.py` | **read time — immediately** |
|
||||
| 2 | appended `RevenueFromContractWithCustomerIncludingAssessedTax` and `RevenuesNetOfInterestExpense` to the revenue concept list | `sec_facts_parser.py` | parse time — **needs reparse** |
|
||||
| 3 | YTD span tolerance 20 → 25 days, covering 4-4-5 retail calendars | `sec_facts_parser.py` | parse time — **needs reparse** |
|
||||
|
||||
Fix 2 is deliberately **additive**: the new tags go at the end of the priority list, so
|
||||
every issuer that already resolved keeps the same concept and only issuers that resolved
|
||||
to nothing gain a value. A regression test pins that ordering.
|
||||
|
||||
Tests: 7 added across `test_sec_facts_parser.py` and `test_fundamentals_derivation.py`.
|
||||
The 5 behaviour-changing ones were confirmed to fail against the pre-fix code; the other 2
|
||||
are invariance guards that pass both ways. Full unit suite: 795 passed.
|
||||
|
||||
## Validation against live SEC data
|
||||
|
||||
Re-ran the parser + derivation on live companyfacts. Every targeted name recovers, and
|
||||
the recovered values independently agree with the legacy provider:
|
||||
|
||||
| name | cause | revenue growth before → after | legacy | TTM EPS after |
|
||||
|---|---|---|---|---|
|
||||
| COST | 4-4-5 Q3 | null → **9.2311** | 9.23 | 19.88 |
|
||||
| PEP | 4-4-5 Q3 | null → **5.6197** | 5.62 | 7.63 |
|
||||
| KHC | concept (Including) | null → **−1.7457** | −1.75 | −4.85 |
|
||||
| DVN | partial 10-K/A | null → **0.0956** | −1.51 | 3.59 |
|
||||
| ARE | concept (Including) | null → **−5.3462** | −9.53 | −6.27 |
|
||||
| JPM | concept (bank) | null → **3.3388** | 108.98 | 20.89 |
|
||||
| GS | concept (bank) | null → **11.1974** | 6.67 | 54.75 |
|
||||
| WFC | concept (bank) | null → **4.1847** | 72.75 | 6.47 |
|
||||
|
||||
COST/PEP/KHC matching legacy to two decimals is strong evidence the parse is now correct.
|
||||
The banks are the opposite case and worth noting for the cutover argument: legacy's JPM
|
||||
109% and WFC 73% "revenue growth" are not plausible for a bank, while the SEC-derived
|
||||
3.3% and 4.2% are — here the candidate is **better** than what it would replace. DVN and
|
||||
ARE still differ from legacy; DVN is the `Revenues` vs ASC-606 ambiguity noted in §2 and
|
||||
is the one open definition question.
|
||||
|
||||
Regression check on names that were already correct — IRM, KLAC, BKNG — reproduces their
|
||||
previous values exactly (IRM 15.6375, KLAC 13.3895, BKNG 14.9506; TTM EPS unchanged).
|
||||
Nothing that worked before moved.
|
||||
|
||||
### Concept consistency across the bank chains (checked, clean)
|
||||
|
||||
Because `Revenues` still outranks `RevenuesNetOfInterestExpense`, a filer could resolve the
|
||||
FY row to one concept and its quarters to the other — which would make
|
||||
`Q4 = YTD(FY) − YTD(Q3)` a subtraction across two definitions, and poison every TTM window
|
||||
containing it. Checked all 15 recovered banks (`scratchpad/concept_check.py`):
|
||||
|
||||
- **14 resolve a single concept across the whole chain** (GS, WFC, MS, TFC, MTB, FITB, RF,
|
||||
SYF, BNY, BX, BLK, SPGI, ACGL, CBOE).
|
||||
- **JPM is mixed but benign**: its FY2025 row tags both, at an *identical* 182,447M, so Q4
|
||||
subtracts like for like. No filer showed the two tags disagreeing where both appear.
|
||||
|
||||
So the "candidate beats legacy for banks" claim above is safe as stated. **Residual risk:**
|
||||
a future filer whose two tags differ would fail silently. Cheapest hardening is to treat
|
||||
the two as one logical revenue concept rather than separate priority entries; the detector
|
||||
script above turns this into a one-command check.
|
||||
|
||||
## Operational note — the parser fixes need a deliberate reparse
|
||||
|
||||
`sec_fundamentals_importer.promote()` treats snapshots as **immutable per accession**: a
|
||||
re-run skips any accession already stored and records a `snapshot_discrepancy` SystemEvent
|
||||
instead. So fixes 2 and 3 change nothing for rows already in the database — recovering
|
||||
COST/PEP/JPM/etc. requires deleting the affected snapshot rows and re-importing, or adding
|
||||
an explicit reparse path. Usefully, the discrepancy warning names exactly which stored
|
||||
accessions now reconstruct differently, so a dry run over existing data will enumerate the
|
||||
blast radius before anything is rewritten.
|
||||
|
||||
---
|
||||
|
||||
# Second pass — all 25 lost names now classified
|
||||
|
||||
Re-ran `diag.py --history` over every previously unclassified name, with fixes 1–3 in
|
||||
place. (One name, DPZ, had been dropped from the unclassified list when this document was
|
||||
rewritten; it is included here.)
|
||||
|
||||
## 11 of 25 recover
|
||||
|
||||
COST, PEP, KHC, DVN, ARE, **AZO, MOS, SJM, SWKS, HAL, DPZ** — and again the recovered
|
||||
revenue growth matches the legacy provider to two decimals on every one:
|
||||
|
||||
| name | candidate | legacy | | name | candidate | legacy |
|
||||
|---|---|---|---|---|---|---|
|
||||
| AZO | 5.7405 | 5.74 | | SWKS | 2.3303 | 2.33 |
|
||||
| MOS | 12.3388 | 12.34 | | HAL | −1.7201 | −1.72 |
|
||||
| SJM | 3.7222 | 3.72 | | DPZ | 5.1573 | 5.16 |
|
||||
|
||||
Precisely: all 11 clear the ≥2-metric floor and regain a fundamental score. P/E returns for
|
||||
AZO, MOS, SWKS, DPZ, COST, PEP and DVN. ARE, KHC and SJM have genuinely negative TTM EPS,
|
||||
so their P/E stays null correctly. **HAL's TTM EPS is still null and the cause is not yet
|
||||
established** — it scores on revenue growth + surprise. Loose end.
|
||||
|
||||
## 14 remain, in four causes
|
||||
|
||||
| cause | names | count |
|
||||
|---|---|---|
|
||||
| **fiscal-year label collisions (§3a)** | CRM, FRT, STX, BXP, CRWD, MTD, NTAP, WDAY, PPL | **9** |
|
||||
| **EPS concept gap (new — §5 below)** | FCX, REG | 2 |
|
||||
| CIK identity (§4) | XOM | 1 |
|
||||
| new registrant — expected, not a bug | PSKY, Q | 2 |
|
||||
|
||||
The label bug is now the dominant cause by a wide margin, and it is more varied than first
|
||||
described — it is not only colliding `fiscal_year` values:
|
||||
|
||||
- **BXP** — a *10-Q* for period end 2026-03-31 is labelled `2026 **FY**`. The **fiscal
|
||||
period** is wrong, not just the year, so `_select_ytd` then measures the 90-day fact
|
||||
against the 365-day FY expectation and rejects it too.
|
||||
- **NTAP, WDAY, MTD, CRWD** — two different period-ends colliding on one key (the pattern
|
||||
first seen on CRM/FRT).
|
||||
- **PPL** — the worst observed: **four** rows keyed `2022 Q3`, with period ends 2022-09-30,
|
||||
2023-03-31, 2023-06-30 and 2023-09-30.
|
||||
|
||||
## 5. New cause — EPS concept coverage
|
||||
|
||||
`_EPS_CONCEPTS = ["EarningsPerShareDiluted"]` is the only tag read. Confirmed by listing
|
||||
every `USD/shares` duration concept in the relevant filings:
|
||||
|
||||
- **REG** tags only `IncomeLossFromContinuingOperationsPerDilutedShare`, on every filing —
|
||||
EPS is null everywhere, so no TTM EPS and no P/E, ever.
|
||||
- **FCX** is the nastier shape: its **10-Qs** tag `EarningsPerShareDiluted`, but its
|
||||
**10-K** tags only `IncomeLossFromContinuingOperationsPerDilutedShare`. The FY row loses
|
||||
EPS, so `Q4 = YTD(FY) − YTD(Q3)` is undefined and TTM dies — an issuer that switches
|
||||
concept *by form type* looks like partial data rather than a mapping gap.
|
||||
|
||||
**Fix:** append `IncomeLossFromContinuingOperationsPerDilutedShare` to `_EPS_CONCEPTS`.
|
||||
Same additive shape as the revenue fix; recovers REG outright and FCX's FY row.
|
||||
|
||||
**Related decision, not a fix:** PPL's 2026 Q1 tags *no diluted variant at all* — only
|
||||
`EarningsPerShareBasic` and `IncomeLossFromContinuingOperationsPerBasicShare`. Adding the
|
||||
diluted continuing-ops tag does not help it. Falling back to basic EPS is a definition
|
||||
change (basic ≠ diluted) and should be an explicit call, not a silent one.
|
||||
|
||||
---
|
||||
|
||||
# Third pass — fixes #2 and #3 applied
|
||||
|
||||
| # | change | file | effective |
|
||||
|---|---|---|---|
|
||||
| 2a | `_guard_split_sensitive_metrics()` now returns whether the *latest* period is split-suspect, and `derive()` nulls `ttm_diluted_eps` (setting `ttm_diluted_eps_caveat`) when it is | `fundamentals_derivation.py` | read time — immediately |
|
||||
| 3 | appended `IncomeLossFromContinuingOperationsPerDilutedShare` to `_EPS_CONCEPTS` | `sec_facts_parser.py` | parse time — needs reparse |
|
||||
|
||||
5 tests added; the 3 behaviour-changing ones confirmed to fail against pre-fix code, 2 are
|
||||
invariance guards. Full unit suite: **800 passed**.
|
||||
|
||||
## Validated on live data
|
||||
|
||||
| name | before | after | |
|
||||
|---|---|---|---|
|
||||
| FCX | TTM EPS null | **1.89** | recovered |
|
||||
| REG | TTM EPS null | **2.92** | recovered |
|
||||
| BKNG | TTM EPS 156.86 → P/E **1.10** | **null** + caveat | false perfect score removed |
|
||||
| COF | TTM EPS 3.92 → P/E 51.01 | **null** + caveat | see side effect below |
|
||||
| KLAC | TTM EPS 35.31 → P/E **6.19** | unchanged | **still wrong — 2b not fixed** |
|
||||
| IRM, COST | — | unchanged | no regression |
|
||||
|
||||
FCX and REG regain a fundamental score (EPS + surprise clears the ≥2 floor). Their
|
||||
**revenue growth is still null** — both are also blocked by the label bug (REG has a
|
||||
mislabelled duplicate `2024 Q2`; FCX is missing its 2024 FY row entirely).
|
||||
|
||||
## Threshold decision — RESOLVED: keep 25%
|
||||
|
||||
Measured against the database (`scratchpad/share_change_check.sql`): **15 of 467 comparable
|
||||
issuers (3.2%)** trip the ≥25% guard on their latest period.
|
||||
|
||||
| band | names | cause |
|
||||
|---|---|---|
|
||||
| ≥200% | BKNG 23.8×, ORLY 14.5×, NFLX 9.8×, NOW 5.0×, TPL 3.0× | forward splits |
|
||||
| 50–142% | CHTR (query artifact), **AMCR −68% (1-for-5 reverse split)**, WAT, COF | split + stock-funded M&A |
|
||||
| 25–47% | OMC, BG, HBAN, FITB, COHR, RKLB | stock-funded M&A, ordinary dilution |
|
||||
|
||||
**Keep the threshold at 25%**, for three reasons — the first of which is empirical and came
|
||||
out of checking AMCR:
|
||||
|
||||
1. **A real split trips at only 68%.** AMCR's 1-for-5 reverse consolidation
|
||||
(2,308,359,941 → 462,045,690 shares, ratio 4.996, between the Nov 2025 and Feb 2026
|
||||
10-Qs) shows up as −68%. Raising the bar to 100% to spare the M&A cases would have let a
|
||||
genuine split straight through. Split magnitude and M&A magnitude overlap in practice,
|
||||
not just in theory.
|
||||
2. **The cost is milder than first described.** Losing P/E leaves revenue growth + earnings
|
||||
surprise = 2 metrics, which still clears the ≥2 floor. Affected issuers keep a
|
||||
fundamental score; they lose one of three inputs.
|
||||
3. **The severities are asymmetric.** A missed split yields a P/E off by 10–25×, clamping to
|
||||
a *perfect 100* sub-score. Over-nulling yields a missing input the scorer already handles
|
||||
by renormalising.
|
||||
|
||||
Honest caveat: the guard is blunt — it detects that a share base moved, not how much damage
|
||||
resulted. AMCR's pre-fix P/E was 28.61 against legacy's 29.47, i.e. only ~10-15% off, because
|
||||
most of its YTD figures had already been restated on the post-split basis. So the guard
|
||||
sometimes removes a roughly-usable number. That is the accepted price of a rule that cannot
|
||||
measure the split factor.
|
||||
|
||||
Two data notes from the same check:
|
||||
|
||||
- **CHTR is a query artifact, not a guard trip.** The SQL picks the newest period *with* a
|
||||
share count, while `derive()` picks the newest period and then reads shares off it. CHTR's
|
||||
recent snapshots have a null `shares_outstanding`, so the query fell back to the 2016 Time
|
||||
Warner merger. In the real path its change is None and the guard never fires — so the true
|
||||
count is ~14. But it also means **CHTR has no recent share count, which breaks its market
|
||||
cap in the API** — a separate small bug.
|
||||
- **AMCR was suspected of being a `shares_outstanding` parsing bug and is not.** It is a real
|
||||
corporate action, correctly detected. `abs()` in the guard already handles reverse splits.
|
||||
|
||||
## Side effect — COF
|
||||
|
||||
The guard fires on *any* ≥25% YoY share-count move, not only splits. COF's 383M → 639M jump
|
||||
is the Discover acquisition, so it now nulls too and **loses the P/E of 51.01** that this
|
||||
document previously called "arithmetically correct on a GAAP TTM basis".
|
||||
|
||||
I think nulling is right: TTM EPS sums four quarters whose per-share figures use different
|
||||
weighted-average denominators, and across a 67% share change that sum is not a meaningful
|
||||
per-share number regardless of whether the cause was a split or an acquisition. It follows
|
||||
the formula without being a valid result.
|
||||
|
||||
But the cost is real and worth stating plainly: **any issuer doing a large stock-funded
|
||||
acquisition loses its P/E for four quarters.** That frequency has not been measured — it
|
||||
needs a count of `|share_count_change_yoy| ≥ 25%` across the universe, which needs the
|
||||
database. If it turns out to be common, the alternative is a higher or split-shaped
|
||||
threshold, at the cost of letting more BKNG-class errors through.
|
||||
|
||||
## 2b is genuinely unfixed
|
||||
|
||||
KLAC's split post-dates its most recent 10-Q, so no snapshot carries any share-count
|
||||
evidence and no guard built on share counts can fire. Its P/E is still 6.19 — the true P/E
|
||||
divided by the split factor. I did not ship a heuristic for this: the obvious one, flagging
|
||||
implausibly low P/Es, would misfire on genuinely cheap names — CHTR (3.42) and CMCSA (4.30)
|
||||
sit below KLAC's corrupted 6.19 in this very report. Detecting it needs an actual
|
||||
corporate-actions source, or a price-vs-share-count reconciliation against an external
|
||||
market-cap reference.
|
||||
|
||||
---
|
||||
|
||||
# Fourth pass — the reparse path
|
||||
|
||||
Snapshots are immutable per accession, so the parser fixes never reached stored rows.
|
||||
`promote()` skipped them and logged a discrepancy. Reparse is the deliberate exception:
|
||||
immutability protects *SEC's* record, but the stored row is **our reconstruction** — after a
|
||||
parser fix, keeping it is preserving a stale cache, not preserving history.
|
||||
|
||||
| change | file |
|
||||
|---|---|
|
||||
| `run_import(..., force=True)` bypasses the unchanged-revision no-op. The revision tracks the *source*; a fix on our side leaves it unchanged, so the gate would skip the run | `data_import.py` |
|
||||
| `SecFundamentalsImporter(reparse=True)` — forces full-history staging, and `promote()` rewrites the accessions whose reconstruction changed, stamping `import_run_id` | `sec_fundamentals_importer.py` |
|
||||
| `scripts/reparse_fundamentals.py` — **dry run by default**, `--apply` to write | new |
|
||||
|
||||
Unchanged rows are never touched; only accessions appearing in `staged.discrepancies` are
|
||||
rewritten. The update writes the full `_SNAPSHOT_COLS` set via the same `_row_values()` the
|
||||
insert uses, so a rewritten row can never be half old-parse and half new-parse. `created_at`
|
||||
keeps its original value. Nothing is wired into the scheduler.
|
||||
|
||||
## A real bug the tests caught: false-positive discrepancies
|
||||
|
||||
`test_reparse_leaves_unchanged_rows_untouched` failed on first run — reparsing *identical*
|
||||
data reported a change. Cause: `accepted_at` is written tz-aware UTC but
|
||||
`DateTime(timezone=True)` only preserves tzinfo on Postgres; SQLite returns it naive, so
|
||||
`_diff_fields` compared representations and saw a difference.
|
||||
|
||||
Left alone this would have made the dry-run report claim **every row needs rewriting** —
|
||||
exactly the misleading signal that makes a blast-radius report worthless. `_diff_fields` now
|
||||
compares datetime *instants* via `_same_value()`. This also fixes a latent false positive in
|
||||
the pre-existing `snapshot_discrepancy` warning, which shares the same code path.
|
||||
|
||||
## Verification
|
||||
|
||||
4 reparse tests added, driven through the real import framework with the fake SEC client.
|
||||
The key one seeds the database through the **pre-fix parser** (monkeypatching
|
||||
`_YTD_TOLERANCE_DAYS` back to 20 so a 4-4-5 Q3 is rejected and stored as null), then reparses
|
||||
with the fixed parser and asserts the row is rewritten in place with new provenance — the
|
||||
production scenario end to end. Also covered: unchanged rows keep their original
|
||||
`import_run_id`; `reparse=False` still reports and refuses to mutate; `force` bypasses the
|
||||
no-op. Full suite: **804 passed**.
|
||||
|
||||
Not verifiable here: this reads and writes production Postgres, which is unreachable from
|
||||
this machine, so the SQLite harness is the limit of what could be self-tested. The UPDATE is
|
||||
plain SQLAlchemy Core with no dialect-specific constructs.
|
||||
|
||||
## Running it
|
||||
|
||||
```
|
||||
python scripts/reparse_fundamentals.py # dry run, writes nothing
|
||||
python scripts/reparse_fundamentals.py --apply # rewrite changed rows
|
||||
```
|
||||
|
||||
Two cautions for whoever runs it:
|
||||
|
||||
- **Read the dry run for *kinds* of change, not just the count.** The tolerance 20→25 change
|
||||
newly accepts facts for arbitrary filers, not only the names investigated here. Sample
|
||||
changed rows for issuers that were never on the list and confirm they are recovered nulls
|
||||
and corrected values — not something unexpected.
|
||||
- **It refetches Company Facts for every tracked issuer** under the SEC throttle, because the
|
||||
facts a fixed parser now accepts were never stored. Expect a long run; the dry run pays
|
||||
that cost too, so budget for two passes.
|
||||
|
||||
Scope: this rewrites `fundamental_snapshots` only. Those rows currently feed the fundamentals
|
||||
API/UI and the parity report — scoring still reads the legacy `fundamental_data` table, and
|
||||
nothing in the backtest path touches `FundamentalSnapshot`. So a reparse **cannot** move
|
||||
composite scores or backtests until the A5 cutover happens. The plan's "changed history
|
||||
changes backtests" caution applies to workstream B's OHLCV rewrites, not to this.
|
||||
|
||||
---
|
||||
|
||||
# Fifth pass — period identity
|
||||
|
||||
The parser's own stated rule was *"period identity comes from `end == reportDate`, never
|
||||
`fy/fp`"* — but `_fiscal_context()` derived the stored `fiscal_year`/`fiscal_period` by
|
||||
majority-voting exactly those fy/fp fields. The labelling contradicted the module's own
|
||||
principle, and SEC's labels are unreliable enough to break the quarter chain.
|
||||
|
||||
`_period_identity()` now derives both from `period_end` against the issuer's
|
||||
`submissions.fiscalYearEnd`: **the form decides FY vs quarter** (a 10-Q can no longer be
|
||||
labelled FY), and **distance to the fiscal-year end decides which quarter**. The MMDD is
|
||||
threaded through `parse_snapshots(..., fiscal_year_end=...)`; without it the old fy/fp path
|
||||
is used unchanged, so nothing regresses for issuers lacking a calendar.
|
||||
|
||||
**Rejected approach:** classifying the period by fact spans. Every 10-Q carries both a YTD
|
||||
*and* a discrete fact ending at reportDate, so "best span match" reads COST's Q2 (167d) as a
|
||||
Q1; and taking the *longest* span mislabelled IRM's Q3 2020 10-Q as FY because that filing
|
||||
carries a 12-month fact. The prototype caught this as a regression on a working name before
|
||||
any code was written. Distance-to-year-end needs no facts at all and is unambiguous — the
|
||||
quarter bands sit 91 days apart, so ±35 absorbs even a 4-4-5 filer's 16-week Q4.
|
||||
|
||||
**Labels no longer match issuer naming in one case, deliberately.** A filer whose year ends
|
||||
in early January (DPZ, `fiscalYearEnd` 0102) shifts by one. That is harmless: `fiscal_year`
|
||||
and `fiscal_period` appear nowhere in the API schemas or routers — they are internal keys the
|
||||
derivation uses for ordering, YTD differencing and YoY pairing, and the API surfaces
|
||||
`period_end`. The requirement is uniqueness, monotonicity and YoY alignment, not nomenclature.
|
||||
DPZ's derived values are byte-identical before and after the shift, which is the proof.
|
||||
|
||||
## Prototype evidence (before implementing)
|
||||
|
||||
Collisions = two period ends on one key, one silently discarded. Inversions = a period
|
||||
sorting before one that precedes it.
|
||||
|
||||
| | CRM | FRT | STX | BXP | PPL | MTD | NTAP | WDAY | CRWD | COST | PEP | IRM | DPZ | AMCR | AAPL |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| collisions before | 1 | 1 | 0 | 0 | 4 | 5 | 2 | 3 | 2 | 0 | 0 | 0 | 1 | 0 | 0 |
|
||||
| collisions after | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
|
||||
| inversions before | 1 | 1 | 1 | 0 | 2 | 10 | 2 | 3 | 5 | 1 | 0 | 0 | 2 | 0 | 0 |
|
||||
| inversions after | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
|
||||
|
||||
## Validated on live data
|
||||
|
||||
8 of the 9 recover fully, every one matching the legacy provider to two decimals:
|
||||
|
||||
| name | TTM EPS | revenue growth | legacy |
|
||||
|---|---|---|---|
|
||||
| CRM | 8.63 | 10.9818 | 10.98 |
|
||||
| FRT | 5.77 | 7.4263 | 7.43 |
|
||||
| STX | 10.54 | 28.9227 | 28.92 |
|
||||
| BXP | 1.99 | 1.6227 | 1.62 |
|
||||
| MTD | 42.57 | 6.7785 | 6.78 |
|
||||
| NTAP | 6.35 | 5.3713 | 5.37 |
|
||||
| WDAY | 3.21 | 13.3165 | 13.32 |
|
||||
| CRWD | −0.10 | 23.1667 | 23.17 |
|
||||
|
||||
**PPL is partial**: revenue growth recovers (8.3353) but TTM EPS is still null — its 2026 Q1
|
||||
tags no diluted EPS variant at all, which is the open basic-vs-diluted decision, not this bug.
|
||||
Note legacy claims −58.81% revenue growth for a utility; 8.34% is far more plausible.
|
||||
|
||||
**Two bonus recoveries**: FCX and REG had recovered EPS in the fourth pass but their revenue
|
||||
growth was still blocked by label collisions. REG now reads 7.7569 against legacy's 7.76.
|
||||
FCX reads 5.4378 against legacy's −24.23 — a genuine disagreement, likely the same
|
||||
`Revenues` vs ASC-606 ambiguity flagged for DVN in §2, and worth resolving with that decision.
|
||||
|
||||
**Regression check — all byte-identical:** IRM 0.92/15.637543, COST 19.88/9.231107,
|
||||
PEP 7.63/5.619741, DPZ 17.64/5.157289, AMCR null/64.834349, JPM 20.89/3.338823,
|
||||
DVN 3.59/0.095648, AZO 145.39/5.740494. Nothing that worked before moved.
|
||||
|
||||
7 tests added at `_period_identity` covering each production shape (10-Q-labelled-FY,
|
||||
December collision, January and mid-year ends, 4-4-5 quarters, the January-crossing shift,
|
||||
and the no-calendar fallback). Full suite: **811 passed**.
|
||||
|
||||
## Reparse note
|
||||
|
||||
This changes `fiscal_year`/`fiscal_period` for a large share of rows — every non-December
|
||||
filer, not only the broken ones. The dry-run count will be **much** larger than for the
|
||||
earlier fixes, and that is expected. Read it by field: `fiscal_year`/`fiscal_period` churn is
|
||||
the intended relabelling; changes to *value* columns are the recoveries.
|
||||
|
||||
## Where the 25 stand now
|
||||
|
||||
22 of 25 have a fundamental score again. Remaining: **XOM** (CIK identity, still unfixed) and
|
||||
**PSKY / Q**, which are new registrants without enough filing history — correct behaviour,
|
||||
not a bug.
|
||||
|
||||
---
|
||||
|
||||
# Sixth pass — CIK identity, and a much larger finding about share counts
|
||||
|
||||
## XOM: pinned, plus the validation that should have caught it
|
||||
|
||||
`company_tickers.json` maps XOM to CIK 2115436 "ExxonMobil Holdings Corp", which has **zero
|
||||
XBRL filings**, while every 10-K/10-Q — including one filed 2026-05-04 — is still under CIK
|
||||
34088. Which registrant is the real filer is a judgement about a corporate event, so it is
|
||||
**pinned explicitly** rather than guessed:
|
||||
|
||||
- `sec_universe.cik_overrides()` reads a `{symbol: cik}` JSON map from
|
||||
`SystemSetting['sec_cik_overrides']` and applies it ahead of `company_tickers.json`.
|
||||
A malformed setting is logged and ignored, never fatal.
|
||||
- **To fix XOM, set:** `sec_cik_overrides = {"XOM": 34088}`.
|
||||
|
||||
The more valuable half is that nothing noticed. A tracked issuer resolving to a registrant
|
||||
with no XBRL filings can never produce a snapshot, and is restaged on *every* run forever.
|
||||
The importer now records those in `staged.no_xbrl_filings`, reports them in the validation
|
||||
summary (`no_xbrl_filings_count`), and raises a `no_xbrl_filings` SystemEvent naming the CIKs
|
||||
and pointing at the override setting. It warns rather than fails — one misresolved ticker
|
||||
must not block the whole import.
|
||||
|
||||
3 tests added. Full suite: **814 passed**.
|
||||
|
||||
## CHTR was not a bug, and the real problem is much bigger
|
||||
|
||||
I previously called this "a separate small bug". Both halves were wrong.
|
||||
|
||||
CHTR's `dei:EntityCommonStockSharesOutstanding` facts stop at **2016-06-30** — exactly when
|
||||
the Time Warner Cable / Bright House deal closed and Charter became a multi-class issuer.
|
||||
Since then the cover page reports the count **per share class**, which is dimensional, and
|
||||
companyfacts is non-dimensional — so the facts are simply not in the API. Its recent filings
|
||||
tag no consolidated common-share concept at all, only preferred and treasury.
|
||||
|
||||
This is not specific to CHTR. Of 12 issuers checked, **7 have no share count at all**:
|
||||
|
||||
| issuer | latest `shares_outstanding` | dei fact history |
|
||||
|---|---|---|
|
||||
| META | null (4/4 recent) | **never tagged** (n=0) |
|
||||
| CMCSA | null (4/4 recent) | stops 2009-12-31 |
|
||||
| BRK-B | null (4/4 recent) | stops 2011-04-29 |
|
||||
| CHTR | null (4/4 recent) | stops 2016-06-30 |
|
||||
| FOXA, NWSA, LEN | null (4/4 recent) | — |
|
||||
| GOOGL / GOOG | 12,230,000,000 | works via the `us-gaap` fallback |
|
||||
|
||||
So **market cap is silently unavailable for a meaningful slice of the large-cap universe**,
|
||||
and it is a source limitation rather than a parser defect: the two obvious workarounds are
|
||||
both already-rejected design decisions — class sums are impossible (the per-class facts are
|
||||
not in companyfacts at all), and the weighted-average diluted count is explicitly excluded
|
||||
because market cap needs a point-in-time value.
|
||||
|
||||
**No code change made.** Substituting weighted-average diluted shares would silently
|
||||
overturn a deliberate design decision and produce a subtly wrong market cap for exactly the
|
||||
biggest, most-watched names. That is a call to make explicitly, so it is listed as a decision
|
||||
below rather than quietly implemented.
|
||||
|
||||
---
|
||||
|
||||
# Seventh pass — multi-class share counts (decision taken: weighted-average fallback)
|
||||
|
||||
## Why this fallback, and why not the alternatives
|
||||
|
||||
Two candidates existed. The one **not** taken: derive the count as
|
||||
`net_income ÷ diluted_eps` from columns already stored — no migration at all, and measured
|
||||
accurate (GOOGL +0.48%, MRNA −0.45%, AAPL +0.19%, MSFT +0.18%). Rejected because it depends
|
||||
on the derived quarter chain — the very thing these fixes have been repairing, and FOXA
|
||||
already fails it — and because the two-class EPS method makes `net_income` differ from the
|
||||
EPS numerator for exactly the multi-class issuers this targets.
|
||||
|
||||
Taken instead: store the **reported** `WeightedAverageNumberOfDilutedSharesOutstanding`.
|
||||
It is the number the filer computed, needs no chain, and covers one issuer more.
|
||||
|
||||
| control | point-in-time | wavg diluted (latest qtr) | ratio |
|
||||
|---|---|---|---|
|
||||
| GOOGL | 12,230,000,000 | 12,309,000,000 | 0.9936 |
|
||||
| MRNA | 396,786,259 | 395,000,000 | 1.0045 |
|
||||
| AAPL | 14,687,356,000 | 14,725,873,000 | 0.9974 |
|
||||
| MSFT | 7,428,434,704 | 7,445,000,000 | 0.9978 |
|
||||
|
||||
## Shape of the change
|
||||
|
||||
- **Migration 027** adds `fundamental_snapshots.weighted_avg_diluted_shares`. A separate
|
||||
column, never backfilled into `shares_outstanding`, so the point-in-time column keeps its
|
||||
strict meaning and the fallback stays a read-time decision.
|
||||
- **Parser** stores the **shortest**-span fact ending at `period_end` (the most recent
|
||||
quarter's average, closest to the current count) — deliberately not the YTD one, since an
|
||||
average is not cumulative and the YTD convention does not apply.
|
||||
- **Derivation** falls back only when the cover-page count is absent, and sets
|
||||
`shares_outstanding_estimated`.
|
||||
- **API** exposes `shares_estimated`, so `market_cap_est` and `fcf_yield` are never presented
|
||||
as exact when they rest on a period average.
|
||||
|
||||
## Validated on live data
|
||||
|
||||
| issuer | shares_outstanding | estimated |
|
||||
|---|---|---|
|
||||
| GOOGL, AAPL, MSFT, MRNA | unchanged point-in-time values | **False** |
|
||||
| META | 2,564,000,000 | True |
|
||||
| CMCSA | 3,570,000,000 | True |
|
||||
| CHTR | 126,849,271 | True |
|
||||
| FOXA | 432,000,000 | True |
|
||||
| NWSA | 555,700,000 | True |
|
||||
| LEN | 240,776,000 | True |
|
||||
| **BRK-B** | **still null** | False |
|
||||
|
||||
6 of 7 recovered, no regression on the controls. **BRK-B remains unavailable** and honestly
|
||||
so: Berkshire reports per *equivalent Class A share*, dimensionally, so it has no consolidated
|
||||
weighted-average fact either. Nothing in companyfacts can give it a share count.
|
||||
|
||||
Known caveat, accepted: for issuers using the two-class method the count is the EPS
|
||||
denominator. For CHTR that is Class A only — which is also the basis on which Charter's equity
|
||||
market cap is normally quoted, so it is the right number for this purpose, but it is not
|
||||
"all shares of all classes".
|
||||
|
||||
3 tests added. Full suite: **817 passed**. Alembic single head at 027.
|
||||
|
||||
**Needs the reparse to land:** existing rows have `weighted_avg_diluted_shares = NULL` until
|
||||
`scripts/reparse_fundamentals.py --apply` runs, so market cap stays missing for these issuers
|
||||
until then.
|
||||
|
||||
---
|
||||
|
||||
# Eighth pass — revenue basis (decision: keep ASC-606, no change)
|
||||
|
||||
The two concepts measure different things: `RevenueFromContractWithCustomerExcludingAssessedTax`
|
||||
is customer-contract revenue (an E&P's oil/gas/NGL sales), while `Revenues` is the total
|
||||
income-statement line, which for commodity producers folds in mark-to-market derivative
|
||||
gains/losses. That is why DVN's ASC-606 figure is *larger*: 4,508M of sales minus ~701M of
|
||||
hedging losses gives the 3,807M `Revenues` line.
|
||||
|
||||
Measured across 21 issuers (deliberately energy-weighted, where the gap concentrates):
|
||||
|
||||
- Both tags present and differing >1%: **5 of 21** — DVN +18.4%, COP −14.3%, OXY +6.5%,
|
||||
FCX −2.8%, PPL +1.6%. Everyone else tags one, or they are identical (COST +0.0%).
|
||||
- Concept choice **flips within an issuer's chain: 0 of 21**. Whichever tag wins, the series
|
||||
is internally consistent, so YoY never compares two definitions.
|
||||
|
||||
**Decision: keep ASC-606 first, change nothing.** Derivative gains/losses are mean-reverting
|
||||
and sign-flipping; folding them into "revenue growth" turns the sub-score into a partial
|
||||
hedging-P&L read for exactly the affected names. The consistency argument for switching is
|
||||
empirically absent (zero flips), and changing would churn every dual-tagging issuer's stored
|
||||
value — widening the reparse diff — to make ~5 names noisier.
|
||||
|
||||
**Correction to the fourth/fifth-pass note:** FCX's disagreement with legacy (+5.44% vs
|
||||
−24.23%) is **not** this ambiguity. Its two tags differ by only 2.8%, and FCX's own revenue
|
||||
rose 22,703M → 25,186M YoY, so −24% is not credible — legacy is simply wrong there, and this
|
||||
decision does not touch it. So the basis choice moves only DVN, COP, OXY.
|
||||
|
||||
The mirror hazard — an issuer where ASC-606 is only a *fragment* of revenue (a bank's fee
|
||||
income) — was checked (all 15 recovered banks resolve total revenue, not a fragment). A
|
||||
fragment-detection warning was prototyped and then **removed**: with no UI surface it would
|
||||
only have lived in the run summary, and the case it guards against is not currently present.
|
||||
Documented and closed rather than shipped as dead plumbing. If a fragment case ever appears,
|
||||
it shows up as an implausibly low revenue in the next parity report.
|
||||
|
||||
---
|
||||
|
||||
# Ninth pass — basic-EPS fallback (PPL) and HAL resolved
|
||||
|
||||
## PPL: basic-EPS fallback (decision taken)
|
||||
|
||||
PPL's 2026 Q1 tags no diluted EPS variant at all, only basic — a single-filing omission
|
||||
(its other quarters tag diluted), but that one missing period broke the quarter chain and
|
||||
nulled TTM. `EarningsPerShareBasic` / `IncomeLossFromContinuingOperationsPerBasicShare` are
|
||||
now appended to `_EPS_CONCEPTS`, last, so they only fire when no diluted variant exists.
|
||||
|
||||
Evidence (19-name scan): a basic fallback helps exactly **1 name (PPL)**. Basic-vs-diluted is
|
||||
~0.5–1.2% for most, +1.2% for PPL. The one name where it genuinely diverges (TSLA +13.3%)
|
||||
already tags diluted, so it never reaches the fallback. Basic is always ≥ diluted, so the
|
||||
result slightly overstates EPS / understates P/E — accepted, since it fires only on an
|
||||
otherwise-null period.
|
||||
|
||||
Validated: PPL TTM EPS null → **1.63** (≈$36 / 1.63 = 22.1 vs legacy P/E 22.43). AAPL, MSFT,
|
||||
DUK, HAL unchanged — diluted still wins wherever present. 2 tests added. Full suite: **819
|
||||
passed**.
|
||||
|
||||
## HAL: resolved, and it was never our bug
|
||||
|
||||
HAL's TTM EPS is now **1.81** (≈$33 / 1.81 = 18.2 vs legacy P/E 18.01) — the period-identity
|
||||
and EPS-concept work already fixed it. The "unexplained null" is closed.
|
||||
|
||||
Its 2024 EPS values are garbage (680000, 1480000, …) because **Halliburton's own 2024 XBRL
|
||||
tags `EarningsPerShareDiluted = 680000` in unit USD/shares** — a filer scale error in the
|
||||
source, faithfully stored. It only poisons TTM windows that include 2024, which the current
|
||||
point-in-time report does not use, so no code change: clamping EPS to "plausible" values would
|
||||
risk masking real ones. Documented as a known source-data quirk.
|
||||
|
||||
This does surface a latent robustness point (not acted on): a single fat-fingered per-share
|
||||
value poisons any TTM window it lands in. It is invisible in the current report and out of
|
||||
scope here, but worth a note if historical TTM series are ever surfaced.
|
||||
|
||||
## All 25 lost names accounted for
|
||||
|
||||
| status | names |
|
||||
|---|---|
|
||||
| **recovered** (22) | ARE, AZO, BXP, COST, CRM, CRWD, DPZ, DVN, FCX, FRT, HAL, KHC, MOS, MTD, NTAP, PEP, PPL, REG, SJM, STX, SWKS, WDAY |
|
||||
| **XOM** | fixed by the `sec_cik_overrides` pin (needs the setting applied) |
|
||||
| **PSKY, Q** | new registrants without enough filing history — correct behaviour, not a bug |
|
||||
|
||||
## Still outstanding
|
||||
|
||||
Revised after the second pass, in the order I would take them:
|
||||
|
||||
Everything actionable without a live database is now done. What remains is one hard
|
||||
data limitation and two operational steps that only run against production.
|
||||
|
||||
1. ~~**§3a fiscal-period identity**~~ — **done**, fifth pass.
|
||||
2. ~~**§1 split contamination, part (a)**~~ — **done**, third pass.
|
||||
3. ~~**§5 EPS concept gap**~~ — **done**, third pass.
|
||||
4. ~~**§4 XOM CIK remap** + zero-filings validation~~ — **done**, sixth pass.
|
||||
5. ~~**Reparse path**~~ — **done**, fourth pass.
|
||||
6. ~~**Multi-class share counts**~~ — **done**, seventh pass (weighted-average fallback).
|
||||
7. ~~**DVN/FCX revenue basis**~~ — **decided**, eighth pass (keep ASC-606, no change).
|
||||
8. ~~**PPL basic-EPS fallback**~~ — **done**, ninth pass.
|
||||
9. ~~**HAL null TTM EPS**~~ — **resolved**, ninth pass (already fixed; 2024 is a filer error).
|
||||
|
||||
## Review finding — two fixes on this branch silently interacted
|
||||
|
||||
Caught in review, not by me. `_merge_amendments` (the per-field amendment fix, first pass)
|
||||
builds the merged period from `_MERGED_FIELDS` + `_CARRIED_FIELDS` alone, so a column in
|
||||
neither list is **absent** from the merged row, not merely stale — and every caller reads it
|
||||
with `getattr(row, name, None)`, which quietly returns `None`.
|
||||
|
||||
`weighted_avg_diluted_shares` (the market-cap fallback, seventh pass) was never added to
|
||||
`_MERGED_FIELDS`. The failure needed both fixes to be present at once: a multi-class issuer
|
||||
*and* a partial amendment on its latest period — META with a Part-III-only 10-K/A — would
|
||||
silently lose market cap and FCF yield again, i.e. the seventh pass's fix undone by the
|
||||
first pass's mechanism. I updated `_SNAPSHOT_COLS` in the importer when adding the column
|
||||
but not `_MERGED_FIELDS` in the derivation.
|
||||
|
||||
Fixed, with a regression test for the specific case. The more useful addition is a guard —
|
||||
`test_merge_lists_cover_every_parser_field` asserts the two lists cover every `SnapshotRow`
|
||||
field, so the *next* column added fails loudly instead of losing data quietly. Verified it
|
||||
would have caught this one.
|
||||
|
||||
Lesson worth keeping: a hand-maintained field list that reconstructs an object is a silent
|
||||
data-loss footgun. `_SNAPSHOT_COLS` (importer) and `_MERGED_FIELDS` (derivation) must both
|
||||
track the parser's `SnapshotRow`, and only one of them is now enforced by a test.
|
||||
|
||||
## Genuinely unfixable from this data
|
||||
|
||||
- **§1 part (b)** — a split post-dating the last filing (KLAC). No snapshot carries
|
||||
share-count evidence, so no guard built on share counts can fire. Needs a corporate-actions
|
||||
source or an external market-cap reconciliation.
|
||||
- **BRK-B market cap** — Berkshire reports per equivalent Class A share, dimensionally, so it
|
||||
has neither a cover-page count nor a weighted-average one. Nothing in companyfacts can give
|
||||
it a share count.
|
||||
|
||||
## Operational steps (production only — cannot run from here)
|
||||
|
||||
- Apply the setting `sec_cik_overrides = {"XOM": 34088}`.
|
||||
- Run `scripts/reparse_fundamentals.py` — dry run first, then `--apply`. This is what lands
|
||||
every parser-side fix (revenue/EPS concepts, Q3 span, period identity, weighted-average
|
||||
shares via migration 027) onto existing rows. Until it runs, those fixes are inert in prod.
|
||||
|
||||
## Standing decision, revisit only if it bites
|
||||
|
||||
- **COF-class share-change threshold** — kept at 25%. Revisit only if the 3.2% universe
|
||||
hit-rate proves painful.
|
||||
|
||||
## Known source-data quirk, not acted on
|
||||
|
||||
- A single fat-fingered per-share value in a filer's XBRL (HAL 2024) poisons any TTM window
|
||||
it lands in. Invisible in the current point-in-time report; relevant only if historical TTM
|
||||
series are ever surfaced.
|
||||
|
||||
PSKY and Q need nothing — they are new registrants without enough filing history, which is
|
||||
correct behaviour.
|
||||
|
||||
---
|
||||
|
||||
# Closing — post-reparse verification (2026-07-24)
|
||||
|
||||
## Production reparse
|
||||
|
||||
Two apply runs against prod (`scripts/reparse_fundamentals.py --apply`):
|
||||
|
||||
- **Run 6** (all fixes through the seventh pass): 262 inserted, 28,664 rewritten —
|
||||
99.4% of which was backfilling the new `weighted_avg_diluted_shares` column; the
|
||||
behavioural churn matched the dry run exactly. The five duration facts clustering at
|
||||
190–234 changed rows each is the 4-4-5 Q3 recovery signature. `accepted_at` changed on
|
||||
only 76 rows (0.25%), confirming the tz-comparison fix works against real Postgres.
|
||||
- **Run 7** (after the fiscal-year-end fix below): 3 inserted, 322 rewritten — BEN, DELL,
|
||||
and boundary-year relabels for 53-week filers whose derived MMDD shifted a few days.
|
||||
|
||||
`cik_updates: 1` on run 6 was the XOM pin taking effect; XOM now has 68 snapshot rows,
|
||||
latest period end 2026-03-31.
|
||||
|
||||
## Regression caught by the collision check — and its fix
|
||||
|
||||
The before/after key-collision query (~130 rows max 6 → 44 rows all 2) surfaced one real
|
||||
regression: **BEN**. `submissions.fiscalYearEnd` declares `1231` while every Franklin
|
||||
Resources 10-K ends 09-30, so `_period_identity` — which trusted the declared value — put
|
||||
BEN's real fiscal Q1 zero days from the claimed year end (no band matched) and labelled its
|
||||
fiscal Q2 as Q1. The collision discarded a period and BEN lost TTM EPS and revenue growth it
|
||||
had before the branch. Fixed in `3d42ca7`: `resolve_fiscal_year_end()` prefers the issuer's
|
||||
own most recent 10-K reportDate (which *is* the fiscal year end by definition) and treats
|
||||
the declared field as fallback. Full-universe scan: 2 of 506 issuers mis-declare (BEN 91d,
|
||||
DELL 29d); both now derive correctly (BEN rg 3.8243 vs legacy 3.82; DELL 38.5735 vs 38.57).
|
||||
|
||||
Residual collisions after run 7: 36 rows, all count-2, **latest year 2023** — the 53-week
|
||||
drift class (AVY/CDNS/RVTY/JNJ/TDY/DPZ at 5–6-year intervals). Newest-wins degrades one
|
||||
historical FY row; no current period is affected. Left alone deliberately: eliminating them
|
||||
means modelling each filer's actual 52/53-week calendar per year, for rows feeding no
|
||||
current metric.
|
||||
|
||||
## The verdict: 2026-07-24 parity report vs the 2026-07-23 baseline
|
||||
|
||||
| metric | baseline | after | |
|
||||
|---|---|---|---|
|
||||
| candidate scored | 482 | **504** | legacy scores 507; gap = PSKY, Q (new registrants) + FITB |
|
||||
| revenue_growth candidate available | 442 | **489** | banks, REITs, 4-4-5 recovered |
|
||||
| pe_ratio candidate available | 432 | **452** | net of the split-guard nulls |
|
||||
| revenue_growth median abs delta | 0.0038 | **0.0038** | 47 names added at unchanged agreement |
|
||||
| pe_ratio median / p95 abs delta | 0.5883 / 7.73 | **0.5576 / 6.03** | corrupted outliers gone |
|
||||
|
||||
Revenue `material_differences` rose 84 → 96: the newly compared names include the cases
|
||||
where **legacy is the wrong side** (JPM 108.98% vs 3.34%, PPL −58.81% vs 8.34%, FCX −24.23%
|
||||
vs +5.44%). Material is symmetric; these flag the provider being corrected.
|
||||
|
||||
The split guard is visible in the report: BKNG (1.10), COF, TPL, AMCR, WAT all null P/E now.
|
||||
**KLAC (6.19) is the one known-wrong value left** — the post-filing split documented as
|
||||
unfixable without a corporate-actions source.
|
||||
|
||||
**Correction to the seventh pass:** the claim that guard-tripped issuers "keep their
|
||||
fundamental score, losing one of three inputs" fails for **FITB**, the one name that also
|
||||
lacks revenue growth (its recent filings tag only ASC-606 fee-income fragments, then nothing)
|
||||
— nulling its contaminated P/E (Comerica merger, 661M → 902M shares) drops it to one metric
|
||||
and no score. Accepted: the composite renormalises, and legacy's 58% "revenue growth" for a
|
||||
bank was itself junk.
|
||||
|
||||
## Recommendation
|
||||
|
||||
The A5 gate evidence now supports approving the cutover: coverage within 3 of legacy with
|
||||
every gap explained, agreement essentially exact where both sides exist, every corrupted
|
||||
value either fixed or deliberately nulled with a caveat, and the remaining score deltas are
|
||||
documented definition differences — called out, not averaged away, as the plan requires.
|
||||
Carry KLAC as the one known caveat in the approval note.
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,114 @@
|
||||
# Focused daily portfolio-capacity matrix
|
||||
|
||||
Generated: 2026-08-05T19:25:17.150472+00:00
|
||||
|
||||
## Question
|
||||
|
||||
The current daily Phase A control admitted 472 trades and rejected 519 qualified opportunities because the ten-slot book was full. This run brackets the economic cost of that binding constraint; it has no formal promotion gate.
|
||||
|
||||
> Universe caveat: today's production membership is projected backward. Use paired arm-versus-control differences, not absolute profitability, for construction conclusions.
|
||||
|
||||
## Validated universes
|
||||
|
||||
- Tradable setup symbols with prices: 505.
|
||||
- Rank-only symbols with prices: 4149.
|
||||
- Full ranking symbols with prices: 4654.
|
||||
- Tradable qualified longs: 6118.
|
||||
- Rank-only qualified rows removed: 136286.
|
||||
|
||||
## Paired annual medians
|
||||
|
||||
### Empty Book — 0.10% per fill
|
||||
|
||||
| Arm | ΔEV net R | 90% context | ΔCalmar | 90% context |
|
||||
|---|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | [0.000, 0.000] | 0.000 | [0.000, 0.000] |
|
||||
| cash_unbounded | 0.044 | [-0.011, 0.060] | 0.030 | [-0.030, 0.120] |
|
||||
| cap10_weekly_top10 | 0.000 | [-0.091, 0.000] | 0.000 | [-0.260, 0.000] |
|
||||
| cap15_incumbent | 0.000 | [0.000, 0.011] | 0.000 | [0.000, 0.130] |
|
||||
|
||||
| Arm | ΔPF | ΔGain-to-Pain | ΔSortino | ΔCAGR pp | ΔMaxDD pp |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
| cash_unbounded | 0.079 | 0.047 | -0.013 | 1.350 | 0.000 |
|
||||
| cap10_weekly_top10 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
| cap15_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
|
||||
### Warm Book — 0.10% per fill
|
||||
|
||||
| Arm | ΔEV net R | 90% context | ΔCalmar | 90% context |
|
||||
|---|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | [0.000, 0.000] | 0.000 | [0.000, 0.000] |
|
||||
| cash_unbounded | 0.034 | [-0.014, 0.100] | 0.050 | [-0.160, 0.250] |
|
||||
| cap10_weekly_top10 | 0.000 | [-0.158, 0.065] | 0.000 | [-0.200, 0.330] |
|
||||
| cap15_incumbent | 0.000 | [-0.006, 0.000] | 0.000 | [0.000, 0.180] |
|
||||
|
||||
| Arm | ΔPF | ΔGain-to-Pain | ΔSortino | ΔCAGR pp | ΔMaxDD pp |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
| cash_unbounded | 0.062 | 0.085 | -0.004 | 2.200 | 0.400 |
|
||||
| cap10_weekly_top10 | 0.000 | 0.012 | 0.018 | 0.300 | 0.000 |
|
||||
| cap15_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
|
||||
### Empty Book — 0.20% per fill
|
||||
|
||||
| Arm | ΔEV net R | 90% context | ΔCalmar | 90% context |
|
||||
|---|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | [0.000, 0.000] | 0.000 | [0.000, 0.000] |
|
||||
| cash_unbounded | 0.041 | [-0.010, 0.052] | 0.030 | [-0.015, 0.100] |
|
||||
| cap10_weekly_top10 | 0.000 | [-0.090, 0.000] | 0.000 | [-0.260, 0.000] |
|
||||
| cap15_incumbent | 0.000 | [0.000, 0.010] | 0.000 | [0.000, 0.110] |
|
||||
|
||||
| Arm | ΔPF | ΔGain-to-Pain | ΔSortino | ΔCAGR pp | ΔMaxDD pp |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
| cash_unbounded | 0.066 | 0.035 | -0.014 | 0.900 | 0.000 |
|
||||
| cap10_weekly_top10 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
| cap15_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
|
||||
### Warm Book — 0.20% per fill
|
||||
|
||||
| Arm | ΔEV net R | 90% context | ΔCalmar | 90% context |
|
||||
|---|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | [0.000, 0.000] | 0.000 | [0.000, 0.000] |
|
||||
| cash_unbounded | 0.034 | [-0.022, 0.102] | 0.040 | [-0.130, 0.230] |
|
||||
| cap10_weekly_top10 | 0.000 | [-0.158, 0.065] | 0.000 | [-0.190, 0.310] |
|
||||
| cap15_incumbent | 0.000 | [-0.006, 0.000] | 0.000 | [0.000, 0.170] |
|
||||
|
||||
| Arm | ΔPF | ΔGain-to-Pain | ΔSortino | ΔCAGR pp | ΔMaxDD pp |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
| cash_unbounded | 0.060 | 0.083 | -0.003 | 2.100 | 0.300 |
|
||||
| cap10_weekly_top10 | 0.000 | 0.017 | 0.020 | 0.300 | 0.000 |
|
||||
| cap15_incumbent | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
|
||||
|
||||
## Warm-seed initialization dispersion
|
||||
|
||||
| Arm | Cost/fill | Median EV IQR ratio | Median Calmar IQR ratio |
|
||||
|---|---:|---:|---:|
|
||||
| cap10_incumbent | 0.10% | 1.000 | 1.000 |
|
||||
| cash_unbounded | 0.10% | 1.000 | 1.000 |
|
||||
| cap10_weekly_top10 | 0.10% | 1.000 | 1.000 |
|
||||
| cap15_incumbent | 0.10% | 1.000 | 1.000 |
|
||||
| cap10_incumbent | 0.20% | 1.000 | 1.000 |
|
||||
| cash_unbounded | 0.20% | 1.000 | 1.000 |
|
||||
| cap10_weekly_top10 | 0.20% | 1.000 | 1.000 |
|
||||
| cap15_incumbent | 0.20% | 1.000 | 1.000 |
|
||||
|
||||
## Capacity and operations — 0.10% per fill
|
||||
|
||||
| Arm | Median trades | Median blocked | Median positions | Peak | Turnover | Min-risk rejects |
|
||||
|---|---:|---:|---:|---:|---:|---:|
|
||||
| cap10_incumbent | 76.0 | 21.6% | 4.98 | 10 | 26.36 | 0 |
|
||||
| cash_unbounded | 74.0 | 0.0% | 4.82 | 12 | 26.76 | 85517 |
|
||||
| cap10_weekly_top10 | 88.0 | 18.1% | 5.13 | 10 | 28.44 | 0 |
|
||||
| cap15_incumbent | 79.0 | 0.0% | 5.15 | 12 | 27.32 | 0 |
|
||||
|
||||
## Weekly-ranking opportunity set
|
||||
|
||||
- Median fresh entrant pool: 0.0.
|
||||
- Median zero-entrant fraction: 0.558.
|
||||
- Replacements across reported paths: 2170.
|
||||
- Same-symbol re-entries within 10 sessions: 529.
|
||||
|
||||
Bootstrap intervals above resample seven annual summaries and are descriptive context only. They are not gates or independent-population confidence claims.
|
||||
@@ -0,0 +1,764 @@
|
||||
'''Pure helpers for the focused daily portfolio-capacity research matrix.'''
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
import random
|
||||
import statistics
|
||||
from collections import defaultdict
|
||||
from datetime import date, timedelta
|
||||
from typing import Any, Iterable
|
||||
|
||||
|
||||
ARMS: tuple[dict[str, Any], ...] = (
|
||||
{
|
||||
'id': 'cap10_incumbent',
|
||||
'label': 'Cap 10, arrival-order incumbents',
|
||||
'max_positions': 10,
|
||||
'min_initial_risk_fraction': None,
|
||||
'weekly_top_n_rebalance': False,
|
||||
},
|
||||
{
|
||||
'id': 'cash_unbounded',
|
||||
'label': 'Cash-constrained, no count cap',
|
||||
'max_positions': None,
|
||||
'min_initial_risk_fraction': 0.005,
|
||||
'weekly_top_n_rebalance': False,
|
||||
},
|
||||
{
|
||||
'id': 'cap10_weekly_top10',
|
||||
'label': 'Cap 10, weekly current-rank top 10',
|
||||
'max_positions': 10,
|
||||
'min_initial_risk_fraction': None,
|
||||
'weekly_top_n_rebalance': True,
|
||||
},
|
||||
{
|
||||
'id': 'cap15_incumbent',
|
||||
'label': 'Cap 15, arrival-order incumbents',
|
||||
'max_positions': 15,
|
||||
'min_initial_risk_fraction': None,
|
||||
'weekly_top_n_rebalance': False,
|
||||
},
|
||||
)
|
||||
|
||||
ARM_BY_ID = {arm['id']: arm for arm in ARMS}
|
||||
RISK_FLOOR_ARMS: tuple[dict[str, Any], ...] = (
|
||||
ARMS[0],
|
||||
{
|
||||
'id': 'cap10_min_risk_005',
|
||||
'label': 'Cap 10, 0.5% minimum effective initial risk',
|
||||
'max_positions': 10,
|
||||
'min_initial_risk_fraction': 0.005,
|
||||
'weekly_top_n_rebalance': False,
|
||||
},
|
||||
)
|
||||
COSTS_PER_SIDE_PCT = (0.1, 0.2)
|
||||
ANCHOR_YEARS = tuple(range(2019, 2026))
|
||||
SCORING_SESSIONS = 504
|
||||
MEASUREMENT_SESSIONS = 252
|
||||
RESIDUAL_BENCHMARK_SESSIONS = 252
|
||||
WARM_SEED_MIN_OFFSET = 63
|
||||
WARM_SEED_MAX_OFFSET = 126
|
||||
BOOTSTRAP_REPLICATES = 10_000
|
||||
BOOTSTRAP_SEED = 20260805
|
||||
PRIMARY_METRICS = (
|
||||
'ev_net_r',
|
||||
'calmar',
|
||||
'profit_factor',
|
||||
'gain_to_pain',
|
||||
'sortino',
|
||||
)
|
||||
PAIRED_METRICS = (
|
||||
*PRIMARY_METRICS,
|
||||
'cagr_pct',
|
||||
'max_drawdown_pct',
|
||||
'total_return_pct',
|
||||
'sharpe',
|
||||
)
|
||||
|
||||
|
||||
def _end_exclusive(
|
||||
sessions: list[date], start_index: int, count: int
|
||||
) -> date:
|
||||
end_index = start_index + count
|
||||
if end_index < len(sessions):
|
||||
return sessions[end_index]
|
||||
return sessions[-1] + timedelta(days=1)
|
||||
|
||||
|
||||
def build_cohort_manifest(session_dates: Iterable[date]) -> dict[str, Any]:
|
||||
sessions = sorted(set(session_dates))
|
||||
minimum = RESIDUAL_BENCHMARK_SESSIONS + SCORING_SESSIONS
|
||||
if len(sessions) <= minimum + MEASUREMENT_SESSIONS:
|
||||
raise ValueError('Snapshot is too short for the frozen cohort design')
|
||||
|
||||
index_of = {session: index for index, session in enumerate(sessions)}
|
||||
first_eligible_index = RESIDUAL_BENCHMARK_SESSIONS - 1 + SCORING_SESSIONS
|
||||
last_eligible_index = len(sessions) - MEASUREMENT_SESSIONS
|
||||
|
||||
first_by_month: dict[tuple[int, int], date] = {}
|
||||
for session in sessions:
|
||||
first_by_month.setdefault((session.year, session.month), session)
|
||||
|
||||
empty: list[dict[str, Any]] = []
|
||||
for (year, month), session in sorted(first_by_month.items()):
|
||||
index = index_of[session]
|
||||
if year not in ANCHOR_YEARS:
|
||||
continue
|
||||
if index < first_eligible_index or index > last_eligible_index:
|
||||
continue
|
||||
empty.append({
|
||||
'protocol': 'empty_book',
|
||||
'path_id': f'empty-{year:04d}-{month:02d}',
|
||||
'cluster': year,
|
||||
'simulation_start': session.isoformat(),
|
||||
'measurement_start': session.isoformat(),
|
||||
'hard_end_exclusive': _end_exclusive(
|
||||
sessions, index, MEASUREMENT_SESSIONS
|
||||
).isoformat(),
|
||||
})
|
||||
|
||||
first_by_year: dict[int, date] = {}
|
||||
for session in sessions:
|
||||
first_by_year.setdefault(session.year, session)
|
||||
|
||||
warm: list[dict[str, Any]] = []
|
||||
warm_seed_counts: dict[str, int] = {}
|
||||
for year in ANCHOR_YEARS:
|
||||
anchor = first_by_year.get(year)
|
||||
if anchor is None:
|
||||
continue
|
||||
anchor_index = index_of[anchor]
|
||||
if (
|
||||
anchor_index < WARM_SEED_MAX_OFFSET
|
||||
or anchor_index > last_eligible_index
|
||||
):
|
||||
continue
|
||||
seed_window = sessions[
|
||||
anchor_index - WARM_SEED_MAX_OFFSET:
|
||||
anchor_index - WARM_SEED_MIN_OFFSET + 1
|
||||
]
|
||||
first_by_iso_week: dict[tuple[int, int], date] = {}
|
||||
for session in seed_window:
|
||||
iso = session.isocalendar()
|
||||
first_by_iso_week.setdefault((iso.year, iso.week), session)
|
||||
seeds = sorted(first_by_iso_week.values())
|
||||
warm_seed_counts[str(year)] = len(seeds)
|
||||
for seed_index, seed in enumerate(seeds, 1):
|
||||
warm.append({
|
||||
'protocol': 'warm_book',
|
||||
'path_id': f'warm-{year}-seed-{seed_index:02d}',
|
||||
'cluster': year,
|
||||
'simulation_start': seed.isoformat(),
|
||||
'measurement_start': anchor.isoformat(),
|
||||
'hard_end_exclusive': _end_exclusive(
|
||||
sessions, anchor_index, MEASUREMENT_SESSIONS
|
||||
).isoformat(),
|
||||
'seed_offset_sessions': anchor_index - index_of[seed],
|
||||
})
|
||||
|
||||
return {
|
||||
'snapshot_first_session': sessions[0].isoformat(),
|
||||
'snapshot_last_session': sessions[-1].isoformat(),
|
||||
'session_count': len(sessions),
|
||||
'expected_clusters': list(ANCHOR_YEARS),
|
||||
'empty_book': empty,
|
||||
'warm_book': warm,
|
||||
'empty_cluster_counts': dict(
|
||||
sorted(
|
||||
(
|
||||
str(year),
|
||||
sum(1 for row in empty if row['cluster'] == year),
|
||||
)
|
||||
for year in {row['cluster'] for row in empty}
|
||||
)
|
||||
),
|
||||
'warm_seed_counts': warm_seed_counts,
|
||||
'empty_cluster_count': len({row['cluster'] for row in empty}),
|
||||
'warm_cluster_count': len({row['cluster'] for row in warm}),
|
||||
}
|
||||
|
||||
|
||||
def validate_cohort_manifest(manifest: dict[str, Any]) -> list[str]:
|
||||
errors: list[str] = []
|
||||
expected = set(ANCHOR_YEARS)
|
||||
empty_clusters = {row['cluster'] for row in manifest['empty_book']}
|
||||
warm_clusters = {row['cluster'] for row in manifest['warm_book']}
|
||||
if empty_clusters != expected:
|
||||
errors.append(
|
||||
f'empty-book clusters {sorted(empty_clusters)} != {sorted(expected)}'
|
||||
)
|
||||
if warm_clusters != expected:
|
||||
errors.append(
|
||||
f'warm-book clusters {sorted(warm_clusters)} != {sorted(expected)}'
|
||||
)
|
||||
for year in ANCHOR_YEARS:
|
||||
seed_count = int(manifest['warm_seed_counts'].get(str(year), 0))
|
||||
if seed_count < 12:
|
||||
errors.append(f'warm anchor {year} has only {seed_count} seeds')
|
||||
return errors
|
||||
|
||||
|
||||
def build_cells(
|
||||
manifest: dict[str, Any],
|
||||
*,
|
||||
arms: tuple[dict[str, Any], ...] = ARMS,
|
||||
protocols: tuple[str, ...] = ('empty_book', 'warm_book'),
|
||||
costs: tuple[float, ...] = COSTS_PER_SIDE_PCT,
|
||||
) -> list[dict[str, Any]]:
|
||||
paths = [
|
||||
path
|
||||
for protocol in protocols
|
||||
for path in manifest[protocol]
|
||||
]
|
||||
cells: list[dict[str, Any]] = []
|
||||
for cost in costs:
|
||||
for path in paths:
|
||||
for arm in arms:
|
||||
cell_id = (
|
||||
f'{arm["id"]}|{path["protocol"]}|{path["path_id"]}'
|
||||
f'|cost={cost:.1f}'
|
||||
)
|
||||
cells.append({
|
||||
**path,
|
||||
'cell_id': cell_id,
|
||||
'arm_id': arm['id'],
|
||||
'cost_per_side_pct': cost,
|
||||
})
|
||||
return cells
|
||||
|
||||
|
||||
def percentile(values: Iterable[float], probability: float) -> float | None:
|
||||
ordered = sorted(float(value) for value in values if value is not None)
|
||||
if not ordered:
|
||||
return None
|
||||
if len(ordered) == 1:
|
||||
return ordered[0]
|
||||
location = (len(ordered) - 1) * probability
|
||||
lower = math.floor(location)
|
||||
upper = math.ceil(location)
|
||||
if lower == upper:
|
||||
return ordered[lower]
|
||||
weight = location - lower
|
||||
return ordered[lower] * (1.0 - weight) + ordered[upper] * weight
|
||||
|
||||
|
||||
def iqr(values: Iterable[float]) -> float | None:
|
||||
clean: list[float] = []
|
||||
for value in values:
|
||||
if value is None:
|
||||
continue
|
||||
parsed = float(value)
|
||||
if math.isfinite(parsed):
|
||||
clean.append(parsed)
|
||||
q25 = percentile(clean, 0.25)
|
||||
q75 = percentile(clean, 0.75)
|
||||
if q25 is None or q75 is None:
|
||||
return None
|
||||
return q75 - q25
|
||||
|
||||
|
||||
def median(values: Iterable[float | None]) -> float | None:
|
||||
clean = [float(value) for value in values if value is not None]
|
||||
return statistics.median(clean) if clean else None
|
||||
|
||||
|
||||
def _safe_ratio(numerator: float | None, denominator: float | None) -> float | None:
|
||||
if numerator is None or denominator is None:
|
||||
return None
|
||||
if abs(denominator) <= 1e-12:
|
||||
return 1.0 if abs(numerator) <= 1e-12 else None
|
||||
return numerator / denominator
|
||||
|
||||
|
||||
def _stable_seed(*parts: object) -> int:
|
||||
digest = hashlib.sha256('|'.join(map(str, parts)).encode('utf-8')).digest()
|
||||
return BOOTSTRAP_SEED + int.from_bytes(digest[:4], 'big')
|
||||
|
||||
|
||||
def bootstrap_median_interval(
|
||||
values: Iterable[float | None],
|
||||
*,
|
||||
seed_parts: tuple[object, ...],
|
||||
replicates: int = BOOTSTRAP_REPLICATES,
|
||||
) -> dict[str, float | int | None]:
|
||||
clean = [float(value) for value in values if value is not None]
|
||||
if not clean:
|
||||
return {'n': 0, 'point': None, 'p05': None, 'p95': None}
|
||||
rng = random.Random(_stable_seed(*seed_parts))
|
||||
draws = [
|
||||
statistics.median(rng.choices(clean, k=len(clean)))
|
||||
for _ in range(replicates)
|
||||
]
|
||||
return {
|
||||
'n': len(clean),
|
||||
'replicates': replicates,
|
||||
'point': statistics.median(clean),
|
||||
'p05': percentile(draws, 0.05),
|
||||
'p95': percentile(draws, 0.95),
|
||||
}
|
||||
|
||||
|
||||
def _monthly_returns(
|
||||
equity_curve: list[dict[str, Any]], base_equity: float
|
||||
) -> list[float]:
|
||||
month_ends: dict[tuple[int, int], float] = {}
|
||||
for point in equity_curve:
|
||||
point_date = date.fromisoformat(str(point['date']))
|
||||
month_ends[(point_date.year, point_date.month)] = float(point['equity'])
|
||||
previous = float(base_equity)
|
||||
returns: list[float] = []
|
||||
for month in sorted(month_ends):
|
||||
equity = month_ends[month]
|
||||
if previous > 0:
|
||||
returns.append(equity / previous - 1.0)
|
||||
previous = equity
|
||||
return returns
|
||||
|
||||
|
||||
def _time_underwater(equities: list[float]) -> tuple[int, float]:
|
||||
peak = float('-inf')
|
||||
current = 0
|
||||
longest = 0
|
||||
underwater = 0
|
||||
for equity in equities:
|
||||
peak = max(peak, equity)
|
||||
if peak > 0 and equity < peak - 1e-9:
|
||||
current += 1
|
||||
underwater += 1
|
||||
longest = max(longest, current)
|
||||
else:
|
||||
current = 0
|
||||
percentage = underwater / len(equities) * 100.0 if equities else 0.0
|
||||
return longest, percentage
|
||||
|
||||
|
||||
def summarize_simulation(sim: dict[str, Any]) -> dict[str, Any]:
|
||||
trades = list(sim.get('trade_details') or [])
|
||||
equity_curve = list(sim.get('equity_curve') or [])
|
||||
net_rs = [float(trade['net_r']) for trade in trades]
|
||||
positive_rs = [value for value in net_rs if value > 0]
|
||||
negative_rs = [value for value in net_rs if value < 0]
|
||||
ev_net_r = statistics.fmean(net_rs) if net_rs else None
|
||||
profit_factor = (
|
||||
sum(positive_rs) / abs(sum(negative_rs))
|
||||
if negative_rs
|
||||
else None
|
||||
)
|
||||
|
||||
base_equity = float(
|
||||
sim.get('measurement_start_equity') or sim.get('starting_capital') or 0.0
|
||||
)
|
||||
curve_equities = [float(point['equity']) for point in equity_curve]
|
||||
daily_equities = [base_equity, *curve_equities]
|
||||
daily_returns = [
|
||||
current / previous - 1.0
|
||||
for previous, current in zip(daily_equities, daily_equities[1:])
|
||||
if previous > 0
|
||||
]
|
||||
downside_deviation = (
|
||||
math.sqrt(
|
||||
statistics.fmean(min(value, 0.0) ** 2 for value in daily_returns)
|
||||
)
|
||||
if daily_returns
|
||||
else None
|
||||
)
|
||||
sortino = (
|
||||
statistics.fmean(daily_returns) / downside_deviation * math.sqrt(252.0)
|
||||
if downside_deviation is not None and downside_deviation > 0
|
||||
else None
|
||||
)
|
||||
monthly_returns = _monthly_returns(equity_curve, base_equity)
|
||||
negative_monthly = sum(value for value in monthly_returns if value < 0)
|
||||
gain_to_pain = (
|
||||
sum(monthly_returns) / abs(negative_monthly)
|
||||
if negative_monthly < 0
|
||||
else None
|
||||
)
|
||||
longest_underwater, underwater_pct = _time_underwater(daily_equities)
|
||||
|
||||
transaction_cost = sum(
|
||||
float(trade.get('transaction_cost') or 0.0) for trade in trades
|
||||
)
|
||||
traded_notional = sum(
|
||||
float(trade.get('shares') or 0.0)
|
||||
* (float(trade.get('entry') or 0.0) + float(trade.get('fill') or 0.0))
|
||||
for trade in trades
|
||||
)
|
||||
turnover_multiple = (
|
||||
traded_notional / base_equity if base_equity > 0 else None
|
||||
)
|
||||
|
||||
ordered_rs = sorted(net_rs, reverse=True)
|
||||
ev_without_best: dict[str, float | None] = {}
|
||||
for count in (1, 5, 10):
|
||||
remaining = ordered_rs[count:]
|
||||
ev_without_best[str(count)] = (
|
||||
statistics.fmean(remaining) if remaining else None
|
||||
)
|
||||
|
||||
events = list(sim.get('weekly_rebalance_events') or [])
|
||||
entrant_sizes = [int(event['fresh_entrant_pool']) for event in events]
|
||||
eligible_sizes = [
|
||||
int(event['rank_eligible_entrant_pool']) for event in events
|
||||
]
|
||||
replacements = [int(event['replacements']) for event in events]
|
||||
|
||||
capacity_skips = int(
|
||||
sim.get('measurement_skipped_book_full', sim.get('skipped_book_full', 0))
|
||||
)
|
||||
opened = int(sim.get('opened_positions', sim.get('trades', 0)))
|
||||
capacity_opportunities = opened + capacity_skips
|
||||
|
||||
result = {
|
||||
'start_date': sim.get('start_date'),
|
||||
'end_date': sim.get('end_date'),
|
||||
'simulation_start_date': sim.get('simulation_start_date'),
|
||||
'measurement_start_equity': base_equity,
|
||||
'measurement_start_positions': sim.get('measurement_start_positions', 0),
|
||||
'trades': len(trades),
|
||||
'ev_net_r': ev_net_r,
|
||||
'profit_factor': profit_factor,
|
||||
'gain_to_pain': gain_to_pain,
|
||||
'sortino': sortino,
|
||||
'ev_without_best': ev_without_best,
|
||||
'total_return_pct': sim.get('total_return_pct'),
|
||||
'cagr_pct': sim.get('cagr_pct'),
|
||||
'max_drawdown_pct': sim.get('max_drawdown_pct'),
|
||||
'calmar': sim.get('calmar'),
|
||||
'sharpe': sim.get('sharpe'),
|
||||
'win_rate': sim.get('win_rate'),
|
||||
'avg_hold_days': sim.get('avg_hold_days'),
|
||||
'longest_underwater_sessions': longest_underwater,
|
||||
'underwater_pct': underwater_pct,
|
||||
'transaction_cost': transaction_cost,
|
||||
'turnover_multiple': turnover_multiple,
|
||||
'skipped_book_full': capacity_skips,
|
||||
'opened_positions': opened,
|
||||
'capacity_opportunities': capacity_opportunities,
|
||||
'blocked_fraction': (
|
||||
capacity_skips / capacity_opportunities
|
||||
if capacity_opportunities
|
||||
else 0.0
|
||||
),
|
||||
'skipped_min_initial_risk': int(
|
||||
sim.get('measurement_skipped_min_initial_risk', 0)
|
||||
),
|
||||
'avg_positions': sim.get('avg_positions'),
|
||||
'peak_positions': sim.get('peak_positions'),
|
||||
'sessions_at_capacity': sim.get('sessions_at_capacity'),
|
||||
'sessions_measured': sim.get('sessions_measured'),
|
||||
'avg_cash_pct': sim.get('avg_cash_pct'),
|
||||
'avg_gross_exposure_pct': sim.get('avg_gross_exposure_pct'),
|
||||
'exit_reasons': sim.get('exit_reasons'),
|
||||
}
|
||||
if events:
|
||||
result['weekly_rebalance'] = {
|
||||
'events': len(events),
|
||||
'zero_entrant_fraction': (
|
||||
sum(1 for value in entrant_sizes if value == 0) / len(events)
|
||||
),
|
||||
'entrant_pool_mean': statistics.fmean(entrant_sizes),
|
||||
'entrant_pool_median': statistics.median(entrant_sizes),
|
||||
'entrant_pool_p90': percentile(entrant_sizes, 0.9),
|
||||
'eligible_pool_mean': statistics.fmean(eligible_sizes),
|
||||
'replacements': sum(replacements),
|
||||
'weekly_rank_rejected_entries': int(
|
||||
sim.get('weekly_rank_rejected_entries', 0)
|
||||
),
|
||||
'reentries_within_5_sessions': int(
|
||||
sim.get('rebalance_reentries_within_5_sessions', 0)
|
||||
),
|
||||
'reentries_within_10_sessions': int(
|
||||
sim.get('rebalance_reentries_within_10_sessions', 0)
|
||||
),
|
||||
'reentries_within_20_sessions': int(
|
||||
sim.get('rebalance_reentries_within_20_sessions', 0)
|
||||
),
|
||||
}
|
||||
return result
|
||||
|
||||
|
||||
def _cluster_rows(
|
||||
cells: list[dict[str, Any]],
|
||||
*,
|
||||
arm_id: str,
|
||||
protocol: str,
|
||||
cost: float,
|
||||
) -> list[dict[str, Any]]:
|
||||
treatment = {
|
||||
row['path_id']: row
|
||||
for row in cells
|
||||
if row['arm_id'] == arm_id
|
||||
and row['protocol'] == protocol
|
||||
and float(row['cost_per_side_pct']) == cost
|
||||
}
|
||||
control = {
|
||||
row['path_id']: row
|
||||
for row in cells
|
||||
if row['arm_id'] == 'cap10_incumbent'
|
||||
and row['protocol'] == protocol
|
||||
and float(row['cost_per_side_pct']) == cost
|
||||
}
|
||||
shared_paths = sorted(set(treatment) & set(control))
|
||||
by_cluster: dict[int, list[tuple[dict, dict]]] = defaultdict(list)
|
||||
for path_id in shared_paths:
|
||||
row = treatment[path_id]
|
||||
by_cluster[int(row['cluster'])].append((row, control[path_id]))
|
||||
|
||||
summaries: list[dict[str, Any]] = []
|
||||
for cluster, pairs in sorted(by_cluster.items()):
|
||||
metrics: dict[str, Any] = {}
|
||||
for metric in PAIRED_METRICS:
|
||||
arm_values = [
|
||||
pair[0]['metrics'].get(metric)
|
||||
for pair in pairs
|
||||
if pair[0]['metrics'].get(metric) is not None
|
||||
and math.isfinite(float(pair[0]['metrics'][metric]))
|
||||
]
|
||||
control_values = [
|
||||
pair[1]['metrics'].get(metric)
|
||||
for pair in pairs
|
||||
if pair[1]['metrics'].get(metric) is not None
|
||||
and math.isfinite(float(pair[1]['metrics'][metric]))
|
||||
]
|
||||
deltas = [
|
||||
float(arm['metrics'][metric])
|
||||
- float(base['metrics'][metric])
|
||||
for arm, base in pairs
|
||||
if arm['metrics'].get(metric) is not None
|
||||
and base['metrics'].get(metric) is not None
|
||||
and math.isfinite(float(arm['metrics'][metric]))
|
||||
and math.isfinite(float(base['metrics'][metric]))
|
||||
]
|
||||
arm_median = median(arm_values)
|
||||
control_median = median(control_values)
|
||||
metrics[metric] = {
|
||||
'arm_median': arm_median,
|
||||
'control_median': control_median,
|
||||
'paired_delta_median': median(deltas),
|
||||
'arm_control_ratio': _safe_ratio(
|
||||
arm_median, control_median
|
||||
),
|
||||
'paired_paths': len(deltas),
|
||||
}
|
||||
summaries.append({
|
||||
'cluster': cluster,
|
||||
'paths': len(pairs),
|
||||
'metrics': metrics,
|
||||
})
|
||||
return summaries
|
||||
|
||||
|
||||
def aggregate_results(
|
||||
cells: list[dict[str, Any]],
|
||||
*,
|
||||
arms: tuple[dict[str, Any], ...] = ARMS,
|
||||
protocols: tuple[str, ...] = ('empty_book', 'warm_book'),
|
||||
costs: tuple[float, ...] = COSTS_PER_SIDE_PCT,
|
||||
include_warm_dispersion: bool = True,
|
||||
) -> dict[str, Any]:
|
||||
paired: list[dict[str, Any]] = []
|
||||
path_distributions: list[dict[str, Any]] = []
|
||||
for cost in costs:
|
||||
for protocol in protocols:
|
||||
control_by_path = {
|
||||
row['path_id']: row
|
||||
for row in cells
|
||||
if row['arm_id'] == 'cap10_incumbent'
|
||||
and row['protocol'] == protocol
|
||||
and float(row['cost_per_side_pct']) == float(cost)
|
||||
}
|
||||
for arm in arms:
|
||||
arm_id = str(arm['id'])
|
||||
clusters = _cluster_rows(
|
||||
cells,
|
||||
arm_id=arm_id,
|
||||
protocol=protocol,
|
||||
cost=float(cost),
|
||||
)
|
||||
headline: dict[str, Any] = {}
|
||||
for metric in PAIRED_METRICS:
|
||||
deltas = [
|
||||
cluster['metrics'][metric]['paired_delta_median']
|
||||
for cluster in clusters
|
||||
]
|
||||
arm_levels = [
|
||||
cluster['metrics'][metric]['arm_median']
|
||||
for cluster in clusters
|
||||
]
|
||||
control_levels = [
|
||||
cluster['metrics'][metric]['control_median']
|
||||
for cluster in clusters
|
||||
]
|
||||
arm_level = median(arm_levels)
|
||||
control_level = median(control_levels)
|
||||
metric_summary: dict[str, Any] = {
|
||||
'paired_delta_median': median(deltas),
|
||||
'arm_median': arm_level,
|
||||
'control_median': control_level,
|
||||
'arm_control_ratio': _safe_ratio(
|
||||
arm_level, control_level
|
||||
),
|
||||
}
|
||||
if metric in ('ev_net_r', 'calmar'):
|
||||
metric_summary['bootstrap_90'] = (
|
||||
bootstrap_median_interval(
|
||||
deltas,
|
||||
seed_parts=(
|
||||
arm_id,
|
||||
protocol,
|
||||
cost,
|
||||
metric,
|
||||
'paired-delta',
|
||||
),
|
||||
)
|
||||
)
|
||||
headline[metric] = metric_summary
|
||||
paired.append({
|
||||
'arm_id': arm_id,
|
||||
'protocol': protocol,
|
||||
'cost_per_side_pct': cost,
|
||||
'clusters': clusters,
|
||||
'headline': headline,
|
||||
})
|
||||
treatment_by_path = {
|
||||
row['path_id']: row
|
||||
for row in cells
|
||||
if row['arm_id'] == arm_id
|
||||
and row['protocol'] == protocol
|
||||
and float(row['cost_per_side_pct']) == float(cost)
|
||||
}
|
||||
shared_paths = sorted(
|
||||
set(treatment_by_path) & set(control_by_path)
|
||||
)
|
||||
path_metrics: dict[str, Any] = {}
|
||||
for metric in PAIRED_METRICS:
|
||||
deltas = [
|
||||
float(treatment_by_path[path_id]['metrics'][metric])
|
||||
- float(control_by_path[path_id]['metrics'][metric])
|
||||
for path_id in shared_paths
|
||||
if treatment_by_path[path_id]['metrics'].get(metric)
|
||||
is not None
|
||||
and control_by_path[path_id]['metrics'].get(metric)
|
||||
is not None
|
||||
and math.isfinite(
|
||||
float(treatment_by_path[path_id]['metrics'][metric])
|
||||
)
|
||||
and math.isfinite(
|
||||
float(control_by_path[path_id]['metrics'][metric])
|
||||
)
|
||||
]
|
||||
path_metrics[metric] = {
|
||||
'paired_paths': len(deltas),
|
||||
'paired_delta_mean': (
|
||||
statistics.fmean(deltas) if deltas else None
|
||||
),
|
||||
'paired_delta_median': median(deltas),
|
||||
'paired_delta_p25': percentile(deltas, 0.25),
|
||||
'paired_delta_p75': percentile(deltas, 0.75),
|
||||
'positive_fraction': (
|
||||
sum(delta > 0.0 for delta in deltas) / len(deltas)
|
||||
if deltas
|
||||
else None
|
||||
),
|
||||
'identical_fraction': (
|
||||
sum(abs(delta) <= 1e-12 for delta in deltas)
|
||||
/ len(deltas)
|
||||
if deltas
|
||||
else None
|
||||
),
|
||||
}
|
||||
path_distributions.append({
|
||||
'arm_id': arm_id,
|
||||
'protocol': protocol,
|
||||
'cost_per_side_pct': cost,
|
||||
'metrics': path_metrics,
|
||||
})
|
||||
|
||||
warm_rows = [
|
||||
row for row in cells if row['protocol'] == 'warm_book'
|
||||
]
|
||||
warm_dispersion: list[dict[str, Any]] = []
|
||||
for cost in costs:
|
||||
for arm in arms:
|
||||
arm_id = str(arm['id'])
|
||||
anchor_rows: list[dict[str, Any]] = []
|
||||
for cluster in ANCHOR_YEARS:
|
||||
arm_paths = [
|
||||
row
|
||||
for row in warm_rows
|
||||
if row['arm_id'] == arm_id
|
||||
and int(row['cluster']) == cluster
|
||||
and float(row['cost_per_side_pct']) == float(cost)
|
||||
]
|
||||
control_by_path = {
|
||||
row['path_id']: row
|
||||
for row in warm_rows
|
||||
if row['arm_id'] == 'cap10_incumbent'
|
||||
and int(row['cluster']) == cluster
|
||||
and float(row['cost_per_side_pct']) == float(cost)
|
||||
}
|
||||
metric_rows: dict[str, Any] = {}
|
||||
for metric in ('ev_net_r', 'calmar'):
|
||||
arm_spread = iqr(
|
||||
row['metrics'].get(metric) for row in arm_paths
|
||||
)
|
||||
control_spread = iqr(
|
||||
control_by_path[row['path_id']]['metrics'].get(metric)
|
||||
for row in arm_paths
|
||||
if row['path_id'] in control_by_path
|
||||
)
|
||||
metric_rows[metric] = {
|
||||
'arm_iqr': arm_spread,
|
||||
'control_iqr': control_spread,
|
||||
'iqr_ratio': _safe_ratio(
|
||||
arm_spread, control_spread
|
||||
),
|
||||
}
|
||||
anchor_rows.append({
|
||||
'cluster': cluster,
|
||||
'seeds': len(arm_paths),
|
||||
'metrics': metric_rows,
|
||||
})
|
||||
|
||||
headline: dict[str, Any] = {}
|
||||
for metric in ('ev_net_r', 'calmar'):
|
||||
ratios = [
|
||||
row['metrics'][metric]['iqr_ratio']
|
||||
for row in anchor_rows
|
||||
]
|
||||
headline[metric] = {
|
||||
'median_iqr_ratio': median(ratios),
|
||||
'bootstrap_90': bootstrap_median_interval(
|
||||
ratios,
|
||||
seed_parts=(
|
||||
arm_id,
|
||||
cost,
|
||||
metric,
|
||||
'warm-iqr-ratio',
|
||||
),
|
||||
),
|
||||
}
|
||||
warm_dispersion.append({
|
||||
'arm_id': arm_id,
|
||||
'cost_per_side_pct': cost,
|
||||
'anchors': anchor_rows,
|
||||
'headline': headline,
|
||||
})
|
||||
|
||||
if not include_warm_dispersion:
|
||||
warm_dispersion = []
|
||||
|
||||
return {
|
||||
'paired_per_year': paired,
|
||||
'paired_path_distributions': path_distributions,
|
||||
'warm_seed_dispersion': warm_dispersion,
|
||||
'bootstrap': {
|
||||
'replicates': BOOTSTRAP_REPLICATES,
|
||||
'seed': BOOTSTRAP_SEED,
|
||||
'interval': 'central 90% percentile, context only',
|
||||
'resampling_unit': 'seven annual paired summaries',
|
||||
},
|
||||
}
|
||||
@@ -0,0 +1,120 @@
|
||||
"""Re-derive every stored SEC snapshot with the current parser.
|
||||
|
||||
Snapshots are immutable per accession, so a parser fix does not reach rows that
|
||||
are already stored: a normal import skips them and only logs a
|
||||
``snapshot_discrepancy``. This script is the deliberate, manual exception --
|
||||
it restages every accession from SEC Company Facts and rewrites the rows whose
|
||||
reconstruction changed.
|
||||
|
||||
**Dry run by default.** Nothing is written unless ``--apply`` is passed. The dry
|
||||
run stages and validates exactly as the real run does (both are read-only) and
|
||||
reports the full blast radius: how many rows would change, which fields, and
|
||||
per-symbol before/after samples.
|
||||
|
||||
Cost: a reparse cannot be served from the database -- the facts a fixed parser now
|
||||
accepts were never stored -- so it refetches Company Facts for every tracked issuer
|
||||
under the SEC fair-access throttle. Expect a long run and a lot of network.
|
||||
|
||||
Scope note: this rewrites ``fundamental_snapshots`` only. As of the A5 gate those
|
||||
rows feed the fundamentals API/UI and the parity report; scoring still reads the
|
||||
legacy ``fundamental_data`` table, so a reparse does not move composite scores or
|
||||
backtests until the cutover happens.
|
||||
|
||||
Examples
|
||||
--------
|
||||
# dry run: report what would change, write nothing
|
||||
python scripts/reparse_fundamentals.py
|
||||
|
||||
# dry run, showing more per-field detail
|
||||
python scripts/reparse_fundamentals.py --samples 40
|
||||
|
||||
# actually rewrite the changed rows
|
||||
python scripts/reparse_fundamentals.py --apply
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import sys
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
from app.database import async_session_factory # noqa: E402
|
||||
from app.services.data_import import run_import # noqa: E402
|
||||
from app.services.sec_fundamentals_importer import SecFundamentalsImporter # noqa: E402
|
||||
|
||||
|
||||
def _parse_args() -> argparse.Namespace:
|
||||
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
ap.add_argument("--apply", action="store_true",
|
||||
help="rewrite changed rows (default: dry run, writes nothing)")
|
||||
ap.add_argument("--samples", type=int, default=20,
|
||||
help="how many changed accessions to show in detail (default 20)")
|
||||
return ap.parse_args()
|
||||
|
||||
|
||||
async def _dry_run(samples: int) -> int:
|
||||
importer = SecFundamentalsImporter(reparse=True)
|
||||
async with async_session_factory() as db:
|
||||
print("staging every tracked issuer from SEC Company Facts (this is the slow part)...")
|
||||
revision = await importer.detect_revision(db)
|
||||
staged = await importer.stage(db)
|
||||
result = await importer.validate(db, staged)
|
||||
|
||||
print(f"\nrevision : {revision}")
|
||||
print(f"issuers fetched : {staged.issuers_fetched}")
|
||||
print(f"rows reconstructed : {len(staged.rows)}")
|
||||
print(f"already stored : {len(staged.existing_accessions)}")
|
||||
print(f"WOULD BE REWRITTEN : {len(staged.discrepancies)}")
|
||||
print(f"new inserts : {len(staged.rows) - len(staged.existing_accessions)}")
|
||||
print(f"validation ok : {result.ok}")
|
||||
if not result.ok:
|
||||
print(f"validation messages : {result.messages}")
|
||||
|
||||
if staged.discrepancies:
|
||||
field_counts = Counter(f for d in staged.discrepancies for f in d["fields"])
|
||||
print("\nchanged fields (accession count per field):")
|
||||
for name, count in field_counts.most_common():
|
||||
print(f" {name:28s} {count}")
|
||||
|
||||
by_accession = {r.accession: r for r in staged.rows}
|
||||
print(f"\nfirst {min(samples, len(staged.discrepancies))} changed accessions:")
|
||||
for d in staged.discrepancies[:samples]:
|
||||
row = by_accession.get(d["accession"])
|
||||
where = f"{row.cik} {row.fiscal_year} {row.fiscal_period}" if row else "?"
|
||||
print(f" {d['accession']} {where:28s} {', '.join(d['fields'])}")
|
||||
|
||||
print(
|
||||
"\nDRY RUN -- nothing was written."
|
||||
"\nCheck that the changes are the *kinds* you expect (recovered nulls,"
|
||||
"\ncorrected values) and sample issuers you did not anticipate before"
|
||||
"\nre-running with --apply."
|
||||
)
|
||||
return 0 if result.ok else 1
|
||||
|
||||
|
||||
async def _apply() -> int:
|
||||
# force=True: the revision tracks SEC, which has not changed — the staleness
|
||||
# is on our side, so the normal no-op gate would skip this.
|
||||
run = await run_import(SecFundamentalsImporter(reparse=True), force=True)
|
||||
if run is None:
|
||||
print("another sec_facts import holds the lock; nothing done")
|
||||
return 1
|
||||
print(f"run {run.id}: status={run.status}")
|
||||
print(f" revision : {run.revision}")
|
||||
print(f" row_counts : {run.row_counts_json}")
|
||||
if run.error_details:
|
||||
print(f" error : {run.error_details}")
|
||||
return 0 if run.status == "promoted" else 1
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = _parse_args()
|
||||
return asyncio.run(_apply() if args.apply else _dry_run(args.samples))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,84 @@
|
||||
'''Shared production-style historical ranking helpers for research runners.'''
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import date
|
||||
|
||||
|
||||
def _period_percentiles(
|
||||
observations: list[dict], value_key: str
|
||||
) -> dict[tuple[str, str], float]:
|
||||
'''Rank one deterministic ticker observation per historical period.'''
|
||||
by_period: dict[tuple, list[dict]] = {}
|
||||
seen: set[tuple[str, str]] = set()
|
||||
for row in observations:
|
||||
identity = (str(row['symbol']), str(row['date']))
|
||||
if identity in seen:
|
||||
raise ValueError(f'Duplicate universe rank observation: {identity}')
|
||||
seen.add(identity)
|
||||
if row.get(value_key) is None:
|
||||
continue
|
||||
period = tuple(row['ranking_period'])
|
||||
by_period.setdefault(period, []).append(row)
|
||||
|
||||
result: dict[tuple[str, str], float] = {}
|
||||
for group in by_period.values():
|
||||
ordered = sorted(
|
||||
group,
|
||||
key=lambda row: (float(row[value_key]), str(row['symbol'])),
|
||||
)
|
||||
denominator = len(ordered) - 1
|
||||
for rank, row in enumerate(ordered):
|
||||
result[(str(row['symbol']), str(row['date']))] = round(
|
||||
rank / denominator * 100.0 if denominator > 0 else 100.0,
|
||||
2,
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def _live_universe_rank_map(
|
||||
observations: list[dict],
|
||||
benchmark_closes: dict[date, float],
|
||||
momentum_weight: float,
|
||||
) -> dict[tuple[str, str], dict[str, float | None]]:
|
||||
'''Historical equivalent of production compute_activation_ranks.
|
||||
|
||||
Every ticker contributes at most once per session. Residual momentum starts
|
||||
only once 252 benchmark closes were point-in-time available; earlier dates
|
||||
use the same raw-momentum fallback as production.
|
||||
'''
|
||||
identities = [(str(row['symbol']), str(row['date'])) for row in observations]
|
||||
if len(identities) != len(set(identities)):
|
||||
raise ValueError('Universe ranking requires one observation per ticker/date')
|
||||
|
||||
raw_pct = _period_percentiles(observations, 'momentum')
|
||||
residual_pct = _period_percentiles(observations, 'residual_momentum')
|
||||
vol_pct = _period_percentiles(observations, 'vol_6m')
|
||||
benchmark_ords = sorted(value.toordinal() for value in benchmark_closes)
|
||||
residual_start_ord = benchmark_ords[251] if len(benchmark_ords) >= 252 else None
|
||||
|
||||
ranks: dict[tuple[str, str], dict[str, float | None]] = {}
|
||||
for row in observations:
|
||||
identity = (str(row['symbol']), str(row['date']))
|
||||
asof_ord = date.fromisoformat(identity[1]).toordinal()
|
||||
momentum_pct = (
|
||||
residual_pct.get(identity)
|
||||
if residual_start_ord is not None and asof_ord >= residual_start_ord
|
||||
else raw_pct.get(identity)
|
||||
)
|
||||
volatility_pct = vol_pct.get(identity)
|
||||
strategy_rank = (
|
||||
round(
|
||||
momentum_pct * momentum_weight
|
||||
+ volatility_pct * (1.0 - momentum_weight),
|
||||
2,
|
||||
)
|
||||
if momentum_pct is not None and volatility_pct is not None
|
||||
else momentum_pct
|
||||
)
|
||||
ranks[identity] = {
|
||||
'momentum_percentile': momentum_pct,
|
||||
'volatility_percentile': volatility_pct,
|
||||
'strategy_rank': strategy_rank,
|
||||
}
|
||||
return ranks
|
||||
@@ -29,6 +29,11 @@ ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.research_rankings import ( # noqa: E402
|
||||
_live_universe_rank_map,
|
||||
_period_percentiles,
|
||||
)
|
||||
|
||||
POLICY_NAMES = (
|
||||
"immediate",
|
||||
"next_session",
|
||||
@@ -107,85 +112,6 @@ def _default_output_path() -> Path:
|
||||
return Path("reports") / f"daily-reentry-matrix-{stamp}.json"
|
||||
|
||||
|
||||
def _period_percentiles(
|
||||
observations: list[dict], value_key: str
|
||||
) -> dict[tuple[str, str], float]:
|
||||
"""Production-style percentiles, one deterministic symbol row per period."""
|
||||
by_period: dict[tuple, list[dict]] = {}
|
||||
seen: set[tuple[str, str]] = set()
|
||||
for row in observations:
|
||||
identity = (str(row["symbol"]), str(row["date"]))
|
||||
if identity in seen:
|
||||
raise ValueError(f"Duplicate universe rank observation: {identity}")
|
||||
seen.add(identity)
|
||||
if row.get(value_key) is None:
|
||||
continue
|
||||
period = tuple(row["ranking_period"])
|
||||
by_period.setdefault(period, []).append(row)
|
||||
|
||||
result: dict[tuple[str, str], float] = {}
|
||||
for group in by_period.values():
|
||||
ordered = sorted(
|
||||
group,
|
||||
key=lambda row: (float(row[value_key]), str(row["symbol"])),
|
||||
)
|
||||
denominator = len(ordered) - 1
|
||||
for rank, row in enumerate(ordered):
|
||||
result[(str(row["symbol"]), str(row["date"]))] = round(
|
||||
rank / denominator * 100.0 if denominator > 0 else 100.0,
|
||||
2,
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def _live_universe_rank_map(
|
||||
observations: list[dict],
|
||||
benchmark_closes: dict[date, float],
|
||||
momentum_weight: float,
|
||||
) -> dict[tuple[str, str], dict[str, float | None]]:
|
||||
"""Historical equivalent of ``compute_activation_ranks``.
|
||||
|
||||
Every ticker contributes at most once per session. Residual momentum starts
|
||||
only once 252 benchmark closes were point-in-time available; earlier dates
|
||||
use the same raw-momentum fallback as production.
|
||||
"""
|
||||
identities = [(str(row["symbol"]), str(row["date"])) for row in observations]
|
||||
if len(identities) != len(set(identities)):
|
||||
raise ValueError("Universe ranking requires one observation per ticker/date")
|
||||
|
||||
raw_pct = _period_percentiles(observations, "momentum")
|
||||
residual_pct = _period_percentiles(observations, "residual_momentum")
|
||||
vol_pct = _period_percentiles(observations, "vol_6m")
|
||||
benchmark_ords = sorted(value.toordinal() for value in benchmark_closes)
|
||||
residual_start_ord = benchmark_ords[251] if len(benchmark_ords) >= 252 else None
|
||||
|
||||
ranks: dict[tuple[str, str], dict[str, float | None]] = {}
|
||||
for row in observations:
|
||||
identity = (str(row["symbol"]), str(row["date"]))
|
||||
asof_ord = date.fromisoformat(identity[1]).toordinal()
|
||||
momentum_pct = (
|
||||
residual_pct.get(identity)
|
||||
if residual_start_ord is not None and asof_ord >= residual_start_ord
|
||||
else raw_pct.get(identity)
|
||||
)
|
||||
volatility_pct = vol_pct.get(identity)
|
||||
strategy_rank = (
|
||||
round(
|
||||
momentum_pct * momentum_weight
|
||||
+ volatility_pct * (1.0 - momentum_weight),
|
||||
2,
|
||||
)
|
||||
if momentum_pct is not None and volatility_pct is not None
|
||||
else momentum_pct
|
||||
)
|
||||
ranks[identity] = {
|
||||
"momentum_percentile": momentum_pct,
|
||||
"volatility_percentile": volatility_pct,
|
||||
"strategy_rank": strategy_rank,
|
||||
}
|
||||
return ranks
|
||||
|
||||
|
||||
class PrecomputedDailyEngine:
|
||||
"""Exact date/symbol lookup over the already-ranked production gate."""
|
||||
|
||||
|
||||
@@ -55,6 +55,11 @@ ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.research_rankings import ( # noqa: E402
|
||||
_live_universe_rank_map,
|
||||
_period_percentiles,
|
||||
)
|
||||
|
||||
# Must match Phase A cache when reusing research-cands.pkl
|
||||
CACHE_VERSION = "research-matrix-v1-daily-prod"
|
||||
|
||||
@@ -104,66 +109,6 @@ def _parse_args() -> argparse.Namespace:
|
||||
return p.parse_args()
|
||||
|
||||
|
||||
def _period_percentiles(
|
||||
observations: list[dict], value_key: str
|
||||
) -> dict[tuple[str, str], float]:
|
||||
by_period: dict[tuple, list[dict]] = {}
|
||||
for row in observations:
|
||||
if row.get(value_key) is None:
|
||||
continue
|
||||
period = tuple(row["ranking_period"])
|
||||
by_period.setdefault(period, []).append(row)
|
||||
result: dict[tuple[str, str], float] = {}
|
||||
for group in by_period.values():
|
||||
ordered = sorted(
|
||||
group, key=lambda row: (float(row[value_key]), str(row["symbol"]))
|
||||
)
|
||||
denominator = len(ordered) - 1
|
||||
for rank, row in enumerate(ordered):
|
||||
result[(str(row["symbol"]), str(row["date"]))] = round(
|
||||
rank / denominator * 100.0 if denominator > 0 else 100.0,
|
||||
2,
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def _live_universe_rank_map(
|
||||
observations: list[dict],
|
||||
benchmark_closes: dict[date, float],
|
||||
momentum_weight: float,
|
||||
) -> dict[tuple[str, str], dict[str, float | None]]:
|
||||
raw_pct = _period_percentiles(observations, "momentum")
|
||||
residual_pct = _period_percentiles(observations, "residual_momentum")
|
||||
vol_pct = _period_percentiles(observations, "vol_6m")
|
||||
benchmark_ords = sorted(value.toordinal() for value in benchmark_closes)
|
||||
residual_start_ord = benchmark_ords[251] if len(benchmark_ords) >= 252 else None
|
||||
ranks: dict[tuple[str, str], dict[str, float | None]] = {}
|
||||
for row in observations:
|
||||
identity = (str(row["symbol"]), str(row["date"]))
|
||||
asof_ord = date.fromisoformat(identity[1]).toordinal()
|
||||
momentum_pct = (
|
||||
residual_pct.get(identity)
|
||||
if residual_start_ord is not None and asof_ord >= residual_start_ord
|
||||
else raw_pct.get(identity)
|
||||
)
|
||||
volatility_pct = vol_pct.get(identity)
|
||||
strategy_rank = (
|
||||
round(
|
||||
momentum_pct * momentum_weight
|
||||
+ volatility_pct * (1.0 - momentum_weight),
|
||||
2,
|
||||
)
|
||||
if momentum_pct is not None and volatility_pct is not None
|
||||
else momentum_pct
|
||||
)
|
||||
ranks[identity] = {
|
||||
"momentum_percentile": momentum_pct,
|
||||
"volatility_percentile": volatility_pct,
|
||||
"strategy_rank": strategy_rank,
|
||||
}
|
||||
return ranks
|
||||
|
||||
|
||||
def _window(arm: dict, name: str) -> dict | None:
|
||||
for row in arm.get("windows") or []:
|
||||
if row.get("window") == name:
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -68,6 +68,11 @@ ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.research_rankings import ( # noqa: E402
|
||||
_live_universe_rank_map,
|
||||
_period_percentiles,
|
||||
)
|
||||
|
||||
CACHE_VERSION = "research-matrix-v1-daily-prod"
|
||||
|
||||
# Pre-registered arm catalogue (order is report order). Control is a0.
|
||||
@@ -210,66 +215,6 @@ def _sqlite_url(path: Path) -> str:
|
||||
return f"sqlite+aiosqlite:///{path.resolve().as_posix()}"
|
||||
|
||||
|
||||
def _period_percentiles(
|
||||
observations: list[dict], value_key: str
|
||||
) -> dict[tuple[str, str], float]:
|
||||
by_period: dict[tuple, list[dict]] = {}
|
||||
for row in observations:
|
||||
if row.get(value_key) is None:
|
||||
continue
|
||||
period = tuple(row["ranking_period"])
|
||||
by_period.setdefault(period, []).append(row)
|
||||
result: dict[tuple[str, str], float] = {}
|
||||
for group in by_period.values():
|
||||
ordered = sorted(
|
||||
group, key=lambda row: (float(row[value_key]), str(row["symbol"]))
|
||||
)
|
||||
denominator = len(ordered) - 1
|
||||
for rank, row in enumerate(ordered):
|
||||
result[(str(row["symbol"]), str(row["date"]))] = round(
|
||||
rank / denominator * 100.0 if denominator > 0 else 100.0,
|
||||
2,
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def _live_universe_rank_map(
|
||||
observations: list[dict],
|
||||
benchmark_closes: dict[date, float],
|
||||
momentum_weight: float,
|
||||
) -> dict[tuple[str, str], dict[str, float | None]]:
|
||||
raw_pct = _period_percentiles(observations, "momentum")
|
||||
residual_pct = _period_percentiles(observations, "residual_momentum")
|
||||
vol_pct = _period_percentiles(observations, "vol_6m")
|
||||
benchmark_ords = sorted(value.toordinal() for value in benchmark_closes)
|
||||
residual_start_ord = benchmark_ords[251] if len(benchmark_ords) >= 252 else None
|
||||
ranks: dict[tuple[str, str], dict[str, float | None]] = {}
|
||||
for row in observations:
|
||||
identity = (str(row["symbol"]), str(row["date"]))
|
||||
asof_ord = date.fromisoformat(identity[1]).toordinal()
|
||||
momentum_pct = (
|
||||
residual_pct.get(identity)
|
||||
if residual_start_ord is not None and asof_ord >= residual_start_ord
|
||||
else raw_pct.get(identity)
|
||||
)
|
||||
volatility_pct = vol_pct.get(identity)
|
||||
strategy_rank = (
|
||||
round(
|
||||
momentum_pct * momentum_weight
|
||||
+ volatility_pct * (1.0 - momentum_weight),
|
||||
2,
|
||||
)
|
||||
if momentum_pct is not None and volatility_pct is not None
|
||||
else momentum_pct
|
||||
)
|
||||
ranks[identity] = {
|
||||
"momentum_percentile": momentum_pct,
|
||||
"volatility_percentile": volatility_pct,
|
||||
"strategy_rank": strategy_rank,
|
||||
}
|
||||
return ranks
|
||||
|
||||
|
||||
def _parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(
|
||||
description=__doc__,
|
||||
|
||||
+21
-4
@@ -15,6 +15,8 @@ from sqlalchemy.ext.asyncio import (
|
||||
create_async_engine,
|
||||
)
|
||||
|
||||
from sqlalchemy import delete
|
||||
|
||||
from app.database import Base
|
||||
from app.providers.protocol import OHLCVData
|
||||
|
||||
@@ -32,14 +34,29 @@ _test_session_factory = async_sessionmaker(
|
||||
)
|
||||
|
||||
|
||||
_schema_created = False
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
async def _setup_db():
|
||||
"""Create all tables before each test and drop them after."""
|
||||
"""Hand every test an empty database.
|
||||
|
||||
The schema is built once and then truncated per test rather than dropped and
|
||||
recreated. A create_all/drop_all cycle costs ~49ms against these 22 tables and
|
||||
ran for every test in the suite — including the many that never open a session
|
||||
— where deleting every row costs ~6ms for the same guarantee. No model sets
|
||||
``sqlite_autoincrement``, so SQLite reuses rowids after a full delete and
|
||||
generated ids still restart at 1.
|
||||
"""
|
||||
global _schema_created
|
||||
async with _test_engine.begin() as conn:
|
||||
await conn.run_sync(Base.metadata.create_all)
|
||||
if not _schema_created:
|
||||
await conn.run_sync(Base.metadata.create_all)
|
||||
_schema_created = True
|
||||
else:
|
||||
for table in reversed(Base.metadata.sorted_tables):
|
||||
await conn.execute(delete(table))
|
||||
yield
|
||||
async with _test_engine.begin() as conn:
|
||||
await conn.run_sync(Base.metadata.drop_all)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
|
||||
@@ -8,7 +8,9 @@ from sqlalchemy.ext.asyncio import AsyncSession
|
||||
from app.exceptions import ValidationError
|
||||
from app.services.admin_service import (
|
||||
get_activation_config,
|
||||
get_fundamentals_cutover_config,
|
||||
update_activation_config,
|
||||
update_fundamentals_cutover_config,
|
||||
)
|
||||
|
||||
|
||||
@@ -76,3 +78,18 @@ class TestActivationConfig:
|
||||
async def test_rejects_out_of_range_confidence(self, session: AsyncSession):
|
||||
with pytest.raises(ValidationError):
|
||||
await update_activation_config(session, {"min_confidence": 120.0})
|
||||
|
||||
|
||||
class TestFundamentalsCutoverConfig:
|
||||
async def test_defaults_off_when_unset(self, session: AsyncSession):
|
||||
assert await get_fundamentals_cutover_config(session) == {"enabled": False}
|
||||
|
||||
async def test_round_trips_explicit_switch(self, session: AsyncSession):
|
||||
assert await update_fundamentals_cutover_config(session, True) == {
|
||||
"enabled": True
|
||||
}
|
||||
assert await get_fundamentals_cutover_config(session) == {"enabled": True}
|
||||
|
||||
assert await update_fundamentals_cutover_config(session, False) == {
|
||||
"enabled": False
|
||||
}
|
||||
|
||||
@@ -17,7 +17,7 @@ from __future__ import annotations
|
||||
import asyncio
|
||||
import os
|
||||
import tempfile
|
||||
from datetime import date, datetime, timezone
|
||||
from datetime import date, datetime, timedelta, timezone
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import func, select
|
||||
@@ -29,6 +29,7 @@ from app.models.data_import_run import DataImportRun
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.system_event import SystemEvent
|
||||
from app.services.data_import import (
|
||||
STATUS_DEFERRED,
|
||||
STATUS_FAILED,
|
||||
STATUS_NO_OP,
|
||||
STATUS_PROMOTED,
|
||||
@@ -67,9 +68,16 @@ class FakeImporter:
|
||||
|
||||
source = "sec_facts"
|
||||
|
||||
def __init__(self, revision, *, ok=True, n_rows=3, raise_in="none"):
|
||||
def __init__(
|
||||
self, revision, *, ok=True, retryable=False, alert_days=None,
|
||||
n_rows=3, raise_in="none",
|
||||
alert_messages=None,
|
||||
):
|
||||
self.revision = revision
|
||||
self.ok = ok
|
||||
self.retryable = retryable
|
||||
self.alert_days = alert_days
|
||||
self.alert_messages = alert_messages or []
|
||||
self.n_rows = n_rows
|
||||
self.raise_in = raise_in
|
||||
self.staged_called = False
|
||||
@@ -94,6 +102,9 @@ class FakeImporter:
|
||||
summary={"staged_rows": len(staged)},
|
||||
source_max_date=date(2026, 7, 21),
|
||||
messages=[] if self.ok else ["coverage below threshold"],
|
||||
retryable=self.retryable,
|
||||
deferred_alert_after_days=self.alert_days,
|
||||
deferred_alert_messages=self.alert_messages,
|
||||
)
|
||||
|
||||
async def promote(self, db, staged, run_id):
|
||||
@@ -200,6 +211,60 @@ async def test_failed_validation_leaves_data_untouched(engine):
|
||||
assert await _count(factory, SystemEvent) == 1 # alerted
|
||||
|
||||
|
||||
async def test_retryable_validation_defers_without_alerting(engine):
|
||||
factory = _factory(engine)
|
||||
await run_import(FakeImporter("rev1", n_rows=3), engine=engine)
|
||||
|
||||
run = await run_import(
|
||||
FakeImporter("rev2", ok=False, retryable=True, n_rows=5), engine=engine
|
||||
)
|
||||
|
||||
assert run is not None and run.status == STATUS_DEFERRED
|
||||
assert "coverage" in (run.error_details or "")
|
||||
assert await _count(factory, FundamentalSnapshot) == 3 # untouched
|
||||
assert await _count(factory, SystemEvent) == 0 # expected retry does not alert
|
||||
|
||||
|
||||
async def test_stale_deferred_validation_emits_deduplicated_warning(engine):
|
||||
factory = _factory(engine)
|
||||
promoted = await run_import(FakeImporter("rev1", n_rows=3), engine=engine)
|
||||
async with factory() as s:
|
||||
promoted.started_at = datetime.now(timezone.utc) - timedelta(days=4)
|
||||
await s.merge(promoted)
|
||||
await s.commit()
|
||||
|
||||
importer = FakeImporter(
|
||||
"rev2", ok=False, retryable=True, alert_days=3,
|
||||
alert_messages=["source detail names OLD-ACCESSION"],
|
||||
n_rows=5,
|
||||
)
|
||||
first = await run_import(importer, engine=engine)
|
||||
second = await run_import(importer, engine=engine)
|
||||
|
||||
assert first is not None and first.status == STATUS_DEFERRED
|
||||
assert second is not None and second.status == STATUS_DEFERRED
|
||||
async with factory() as s:
|
||||
events = (await s.execute(select(SystemEvent))).scalars().all()
|
||||
assert len(events) == 1
|
||||
assert events[0].severity == "warning"
|
||||
assert events[0].code == "sec_facts_deferred_stale"
|
||||
|
||||
assert "OLD-ACCESSION" in events[0].message
|
||||
assert "aged-out" not in events[0].message
|
||||
|
||||
|
||||
async def test_never_promoted_deferred_warning_says_never(engine):
|
||||
factory = _factory(engine)
|
||||
run = await run_import(
|
||||
FakeImporter("rev1", ok=False, retryable=True, alert_days=3),
|
||||
engine=engine,
|
||||
)
|
||||
|
||||
assert run is not None and run.status == STATUS_DEFERRED
|
||||
async with factory() as s:
|
||||
event = (await s.execute(select(SystemEvent))).scalar_one()
|
||||
assert "has never promoted successfully" in event.message
|
||||
|
||||
async def test_exception_in_promote_rolls_back(engine):
|
||||
factory = _factory(engine)
|
||||
await run_import(FakeImporter("rev1", n_rows=3), engine=engine) # baseline
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
"""Tests for v2 correction events and warning alarm episodes."""
|
||||
"""Tests for v3 correction events, warning alarm episodes, and report caveats."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -6,7 +6,9 @@ from datetime import date, timedelta
|
||||
|
||||
from app.services.breadth_service import _breadth_from_closes, compute_divergence_series
|
||||
from app.services.event_study_service import (
|
||||
MIN_EVENTS_FOR_CONFIDENCE,
|
||||
_percentile,
|
||||
_reliability,
|
||||
alarm_episodes,
|
||||
detect_events,
|
||||
evaluate_alarms,
|
||||
@@ -23,6 +25,40 @@ def test_detect_events_uses_rising_edge_and_cooldown():
|
||||
assert [event["index"] for event in events] == [300, 355]
|
||||
|
||||
|
||||
def test_reliability_flags_a_thin_holdout():
|
||||
"""2/4 must not read like a property of the score."""
|
||||
dates = _days(100)
|
||||
backing = dict.fromkeys(dates, 3)
|
||||
|
||||
thin = _reliability(dates, 70, backing, events_detected=11, events_in_holdout=4)
|
||||
assert thin["underpowered"] is True
|
||||
assert thin["events_detected"] == 11
|
||||
assert thin["events_in_holdout"] == 4
|
||||
assert thin["minimum_events"] == MIN_EVENTS_FOR_CONFIDENCE
|
||||
|
||||
ample = _reliability(dates, 70, backing, events_detected=20, events_in_holdout=12)
|
||||
assert ample["underpowered"] is False
|
||||
|
||||
|
||||
def test_reliability_flags_a_sensor_coverage_split():
|
||||
"""The threshold must not be frozen on a different construct than it is tested on.
|
||||
|
||||
Credit history starts partway through the training window, so the score
|
||||
renormalises over two sensors early and three later.
|
||||
"""
|
||||
dates = _days(100)
|
||||
matched = dict.fromkeys(dates, 3)
|
||||
assert _reliability(dates, 70, matched, 20, 12)["sensor_coverage_mismatch"] is False
|
||||
|
||||
# Training is 40% three-sensor; the holdout is entirely three-sensor.
|
||||
split_backing = {d: (3 if index >= 42 else 2) for index, d in enumerate(dates)}
|
||||
mismatched = _reliability(dates, 70, split_backing, 20, 12)
|
||||
assert mismatched["sensor_coverage_mismatch"] is True
|
||||
assert mismatched["train_full_sensor_share"] == 40.0
|
||||
assert mismatched["holdout_full_sensor_share"] == 100.0
|
||||
assert mismatched["sensors_expected"] == 3
|
||||
|
||||
|
||||
def test_percentile_is_fixed_from_supplied_values():
|
||||
values = [float(value) for value in range(0, 101, 10)]
|
||||
assert _percentile(values, 50) == 50.0
|
||||
@@ -52,7 +88,7 @@ def test_evaluate_alarms_counts_episodes_not_alarm_days():
|
||||
assert result["median_lead_days"] == 17.5
|
||||
|
||||
|
||||
def test_breadth_from_fixed_closes_and_pure_divergence():
|
||||
def test_breadth_from_fixed_closes_and_tapered_divergence():
|
||||
dates = _days(10)
|
||||
closes_by_symbol = {
|
||||
"A": list(zip(dates, [1.0 + index for index in range(10)])),
|
||||
@@ -67,6 +103,13 @@ def test_breadth_from_fixed_closes_and_pure_divergence():
|
||||
divergence = compute_divergence_series(falling_breadth, rising_benchmark, lookback=3)
|
||||
assert divergence[dates[-1]] > 0
|
||||
|
||||
# v3: breadth loss with price confirming it is still deterioration, scored at
|
||||
# DIVERGENCE_CONFIRMED_FLOOR of the masked case rather than discarded. v2's
|
||||
# hard gate zeroed this and left Warning at 0 through every selloff.
|
||||
falling_benchmark = list(zip(dates, [100.0 - index for index in range(10)]))
|
||||
no_divergence = compute_divergence_series(falling_breadth, falling_benchmark, lookback=3)
|
||||
assert no_divergence[dates[-1]] == 0
|
||||
confirmed = compute_divergence_series(falling_breadth, falling_benchmark, lookback=3)
|
||||
assert 0 < confirmed[dates[-1]] < divergence[dates[-1]]
|
||||
|
||||
# Flat breadth is not deterioration regardless of price direction.
|
||||
flat_breadth = {day: 60.0 for day in dates}
|
||||
assert compute_divergence_series(flat_breadth, falling_benchmark, lookback=3)[dates[-1]] == 0
|
||||
|
||||
@@ -0,0 +1,293 @@
|
||||
"""A5 activation: local candidate derivation and compat-cache refresh."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import date, datetime, timedelta, timezone
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_async_engine
|
||||
|
||||
from app.database import Base
|
||||
from app.models.earnings_event import EarningsEvent
|
||||
from app.models.fundamental import FundamentalData
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.score import CompositeScore, DimensionScore
|
||||
from app.models.settings import SystemSetting
|
||||
from app.models.ticker import Ticker
|
||||
from app.services import fundamentals_candidate_service as candidates
|
||||
from app.services import fundamentals_derivation as deriv
|
||||
from app.services import fundamental_data_refresh_service as refresh_service
|
||||
|
||||
|
||||
UTC = timezone.utc
|
||||
NOW = datetime(2026, 7, 24, 10, 0, tzinfo=UTC)
|
||||
TODAY = date(2026, 7, 24)
|
||||
|
||||
_engine = create_async_engine("sqlite+aiosqlite://", echo=False)
|
||||
_session_factory = async_sessionmaker(
|
||||
_engine, class_=AsyncSession, expire_on_commit=False
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
async def _setup_tables():
|
||||
async with _engine.begin() as connection:
|
||||
await connection.run_sync(Base.metadata.create_all)
|
||||
yield
|
||||
async with _engine.begin() as connection:
|
||||
await connection.run_sync(Base.metadata.drop_all)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
async def session() -> AsyncSession:
|
||||
async with _session_factory() as db:
|
||||
yield db
|
||||
|
||||
|
||||
def _snapshot_rows(cik: str) -> list[FundamentalSnapshot]:
|
||||
rows: list[FundamentalSnapshot] = []
|
||||
periods = ("Q1", "Q2", "Q3", "FY")
|
||||
months = (3, 6, 9, 12)
|
||||
for fiscal_year, multiplier in ((2025, 1.0), (2026, 1.1)):
|
||||
revenue = [100 * multiplier, 110 * multiplier, 120 * multiplier, 130 * multiplier]
|
||||
eps = [1.0 * multiplier, 1.1 * multiplier, 1.2 * multiplier, 1.3 * multiplier]
|
||||
for index, fiscal_period in enumerate(periods):
|
||||
period_end = date(fiscal_year, months[index], 28)
|
||||
rows.append(
|
||||
FundamentalSnapshot(
|
||||
cik=cik,
|
||||
accession=f"{cik}-{fiscal_year}-{fiscal_period}",
|
||||
form="10-K" if fiscal_period == "FY" else "10-Q",
|
||||
filed_date=period_end,
|
||||
accepted_at=datetime(
|
||||
fiscal_year, months[index], 28, tzinfo=UTC
|
||||
),
|
||||
period_end=period_end,
|
||||
fiscal_year=fiscal_year,
|
||||
fiscal_period=fiscal_period,
|
||||
revenue=sum(revenue[: index + 1]),
|
||||
diluted_eps=sum(eps[: index + 1]),
|
||||
shares_outstanding=1_000,
|
||||
)
|
||||
)
|
||||
return rows
|
||||
|
||||
|
||||
async def test_default_off_performs_no_candidate_read_or_write(
|
||||
session: AsyncSession, monkeypatch
|
||||
):
|
||||
ticker = Ticker(symbol="AAA")
|
||||
session.add(ticker)
|
||||
await session.flush()
|
||||
session.add(
|
||||
FundamentalData(
|
||||
ticker_id=ticker.id,
|
||||
pe_ratio=12,
|
||||
revenue_growth=3,
|
||||
earnings_surprise=1,
|
||||
market_cap=100,
|
||||
fetched_at=NOW,
|
||||
)
|
||||
)
|
||||
await session.commit()
|
||||
|
||||
async def should_not_read(*args, **kwargs):
|
||||
raise AssertionError("default-off refresh derived candidates")
|
||||
|
||||
monkeypatch.setattr(candidates, "build_candidates", should_not_read)
|
||||
summary = await refresh_service.refresh_if_enabled(session, today=TODAY)
|
||||
|
||||
stored = await session.scalar(
|
||||
select(FundamentalData).where(FundamentalData.ticker_id == ticker.id)
|
||||
)
|
||||
assert summary == {
|
||||
"enabled": False,
|
||||
"refreshed": 0,
|
||||
"score_inputs_changed": 0,
|
||||
"dimension_scores_staled": 0,
|
||||
"composite_scores_staled": 0,
|
||||
}
|
||||
assert stored.pe_ratio == 12
|
||||
|
||||
|
||||
async def test_activated_refresh_updates_all_fields_and_invalidates_scores(
|
||||
session: AsyncSession,
|
||||
):
|
||||
session.add(
|
||||
SystemSetting(key=refresh_service.ACTIVATION_KEY, value="true")
|
||||
)
|
||||
first = Ticker(symbol="AAA", cik="0000000001")
|
||||
second = Ticker(symbol="AAB", cik="0000000001")
|
||||
session.add_all([first, second])
|
||||
await session.flush()
|
||||
session.add_all(_snapshot_rows(first.cik))
|
||||
session.add_all(
|
||||
[
|
||||
OHLCVRecord(
|
||||
ticker_id=first.id,
|
||||
date=TODAY - timedelta(days=1),
|
||||
open=100,
|
||||
high=100,
|
||||
low=100,
|
||||
close=100,
|
||||
volume=100,
|
||||
),
|
||||
OHLCVRecord(
|
||||
ticker_id=second.id,
|
||||
date=TODAY - timedelta(days=1),
|
||||
open=200,
|
||||
high=200,
|
||||
low=200,
|
||||
close=200,
|
||||
volume=100,
|
||||
),
|
||||
EarningsEvent(
|
||||
ticker_id=first.id,
|
||||
announce_date=TODAY - timedelta(days=10),
|
||||
session="amc",
|
||||
eps_estimate=2,
|
||||
eps_actual=2.2,
|
||||
source="dolt_earnings",
|
||||
),
|
||||
EarningsEvent(
|
||||
ticker_id=first.id,
|
||||
announce_date=TODAY,
|
||||
session="amc",
|
||||
source="dolt_earnings",
|
||||
),
|
||||
]
|
||||
)
|
||||
for ticker in (first, second):
|
||||
session.add(
|
||||
FundamentalData(
|
||||
ticker_id=ticker.id,
|
||||
pe_ratio=1,
|
||||
revenue_growth=1,
|
||||
earnings_surprise=1,
|
||||
market_cap=1,
|
||||
fetched_at=NOW - timedelta(days=1),
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
DimensionScore(
|
||||
ticker_id=ticker.id,
|
||||
dimension="fundamental",
|
||||
score=50,
|
||||
is_stale=False,
|
||||
computed_at=NOW,
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
CompositeScore(
|
||||
ticker_id=ticker.id,
|
||||
score=50,
|
||||
is_stale=False,
|
||||
weights_json="{}",
|
||||
computed_at=NOW,
|
||||
)
|
||||
)
|
||||
await session.commit()
|
||||
|
||||
summary = await refresh_service.refresh_if_enabled(
|
||||
session, now=NOW, today=TODAY
|
||||
)
|
||||
|
||||
stored = {
|
||||
row.ticker_id: row
|
||||
for row in (
|
||||
await session.execute(select(FundamentalData))
|
||||
).scalars()
|
||||
}
|
||||
assert summary["refreshed"] == 2
|
||||
assert summary["score_inputs_changed"] == 2
|
||||
assert stored[first.id].pe_ratio == pytest.approx(100 / 5.06)
|
||||
assert stored[second.id].pe_ratio == pytest.approx(200 / 5.06)
|
||||
assert stored[first.id].revenue_growth == pytest.approx(10)
|
||||
assert stored[first.id].earnings_surprise == pytest.approx(10)
|
||||
assert stored[first.id].market_cap == pytest.approx(100_000)
|
||||
assert stored[first.id].next_earnings_date == TODAY
|
||||
metadata = json.loads(stored[first.id].unavailable_fields_json)
|
||||
assert metadata["source_pe_ratio"] == "sec_facts+ohlcv_records"
|
||||
assert metadata["source_next_earnings_date"] == "dolt_earnings"
|
||||
|
||||
dimensions = (
|
||||
await session.execute(select(DimensionScore))
|
||||
).scalars().all()
|
||||
composites = (
|
||||
await session.execute(select(CompositeScore))
|
||||
).scalars().all()
|
||||
assert all(row.is_stale for row in dimensions)
|
||||
assert all(row.is_stale for row in composites)
|
||||
|
||||
for row in (*dimensions, *composites):
|
||||
row.is_stale = False
|
||||
await session.commit()
|
||||
unchanged = await refresh_service.refresh_if_enabled(
|
||||
session, now=NOW + timedelta(hours=1), today=TODAY
|
||||
)
|
||||
assert unchanged["score_inputs_changed"] == 0
|
||||
assert not any(
|
||||
(await session.execute(select(DimensionScore.is_stale))).scalars()
|
||||
)
|
||||
assert not any(
|
||||
(await session.execute(select(CompositeScore.is_stale))).scalars()
|
||||
)
|
||||
|
||||
|
||||
async def test_candidate_uses_guarded_derive_outputs_and_share_fallback(
|
||||
session: AsyncSession, monkeypatch
|
||||
):
|
||||
ticker = Ticker(symbol="GUARD", cik="0000000002")
|
||||
session.add(ticker)
|
||||
await session.flush()
|
||||
session.add(
|
||||
FundamentalSnapshot(
|
||||
cik=ticker.cik,
|
||||
accession="raw-accession",
|
||||
form="10-Q",
|
||||
filed_date=TODAY,
|
||||
accepted_at=NOW,
|
||||
period_end=TODAY,
|
||||
fiscal_year=2026,
|
||||
fiscal_period="Q2",
|
||||
diluted_eps=99,
|
||||
shares_outstanding=999,
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
OHLCVRecord(
|
||||
ticker_id=ticker.id,
|
||||
date=TODAY,
|
||||
open=50,
|
||||
high=50,
|
||||
low=50,
|
||||
close=50,
|
||||
volume=100,
|
||||
)
|
||||
)
|
||||
await session.commit()
|
||||
|
||||
def guarded(_rows):
|
||||
return deriv.DerivedFundamentals(
|
||||
metrics={
|
||||
"revenue_growth_yoy": deriv.MetricSeries(value=7)
|
||||
},
|
||||
ttm_diluted_eps=None,
|
||||
ttm_diluted_eps_caveat="split guard applied",
|
||||
shares_outstanding=123,
|
||||
shares_outstanding_estimated=True,
|
||||
latest_period_end=TODAY,
|
||||
latest_filed_date=TODAY,
|
||||
)
|
||||
|
||||
monkeypatch.setattr(candidates.deriv, "derive", guarded)
|
||||
candidate = (await candidates.build_candidates(session, today=TODAY))[0]
|
||||
|
||||
assert candidate.pe_ratio is None
|
||||
assert candidate.market_cap == 50 * 123
|
||||
assert candidate.revenue_growth == 7
|
||||
assert candidate.unavailable_fields["pe_ratio"] == "split guard applied"
|
||||
assert "weighted-average" in candidate.unavailable_fields["market_cap_estimated"]
|
||||
@@ -29,6 +29,7 @@ class Snap:
|
||||
cash_and_st_investments: float | None = None
|
||||
total_debt: float | None = None
|
||||
shares_outstanding: float | None = None
|
||||
weighted_avg_diluted_shares: float | None = None
|
||||
|
||||
|
||||
_FP = ["Q1", "Q2", "Q3", "FY"]
|
||||
@@ -104,6 +105,22 @@ def test_net_debt_leverage_and_share_dilution():
|
||||
assert d.metrics["share_count_change_yoy"].value == pytest.approx(-10.0, abs=1e-6)
|
||||
|
||||
|
||||
def test_split_suspect_share_move_suppresses_share_and_eps_comparisons():
|
||||
rows = _two_years()
|
||||
for row in rows:
|
||||
if row.fiscal_year == 2026:
|
||||
row.shares_outstanding = 2000 # +100% resembles an unadjusted 2-for-1 split
|
||||
|
||||
d = fd.derive(rows)
|
||||
|
||||
for key in ("share_count_change_yoy", "eps_growth_yoy"):
|
||||
series = d.metrics[key]
|
||||
assert series.value is None
|
||||
assert series.history[-1].value is None
|
||||
assert "possible split" in series.caveat
|
||||
assert d.metrics["revenue_growth_yoy"].value == pytest.approx(10.0)
|
||||
|
||||
|
||||
def test_valuation_inputs():
|
||||
d = fd.derive(_two_years())
|
||||
# TTM diluted EPS FY2026 = 1.1+1.21+1.32+1.43 = 5.06
|
||||
@@ -172,3 +189,145 @@ def test_amendment_selection_newest_accepted_wins():
|
||||
# Q4 revenue discrete now uses the amended YTD(FY)=999999 minus YTD(Q3)=363
|
||||
# so TTM/growth reflects the amendment, proving newest accepted_at won.
|
||||
assert d.metrics["revenue_growth_yoy"].value != pytest.approx(10.0, abs=1e-6)
|
||||
|
||||
|
||||
# -- partial amendments (A5 parity findings) ---------------------------------
|
||||
|
||||
def test_partial_amendment_does_not_blank_the_period():
|
||||
# DVN's FY2025 10-K/A carries no financial facts at the report date. Taking
|
||||
# the newest accession wholesale nulled the period, and with it the quarter
|
||||
# chain, TTM and YoY.
|
||||
rows = _two_years()
|
||||
part_iii_only = Snap(2026, "FY", date(2026, 9, 30), date(2026, 11, 1),
|
||||
datetime(2027, 1, 1, tzinfo=UTC))
|
||||
baseline = fd.derive(rows)
|
||||
d = fd.derive(rows + [part_iii_only])
|
||||
assert d.ttm_diluted_eps == pytest.approx(baseline.ttm_diluted_eps)
|
||||
assert d.metrics["revenue_growth_yoy"].value == pytest.approx(
|
||||
baseline.metrics["revenue_growth_yoy"].value
|
||||
)
|
||||
|
||||
|
||||
def test_amendment_restating_one_field_leaves_the_others_intact():
|
||||
rows = _two_years()
|
||||
revenue_only = Snap(2026, "FY", date(2026, 9, 30), date(2026, 11, 1),
|
||||
datetime(2027, 1, 1, tzinfo=UTC), revenue=999999)
|
||||
baseline = fd.derive(rows)
|
||||
d = fd.derive(rows + [revenue_only])
|
||||
assert d.metrics["revenue_growth_yoy"].value != pytest.approx(
|
||||
baseline.metrics["revenue_growth_yoy"].value
|
||||
)
|
||||
assert d.ttm_diluted_eps == pytest.approx(baseline.ttm_diluted_eps) # fell back
|
||||
|
||||
|
||||
def test_same_key_row_for_a_different_period_is_never_merged():
|
||||
# SEC labels two different year-ends with one fiscal_year for some filers
|
||||
# (FRT, CRM). That is a mislabelled filing, not an amendment -- merging the
|
||||
# two would silently blend fiscal years.
|
||||
rows = _two_years()
|
||||
mislabelled = Snap(2026, "FY", date(2027, 9, 30), date(2027, 11, 1),
|
||||
datetime(2027, 12, 1, tzinfo=UTC), revenue=999999)
|
||||
selected = fd._select_latest_per_period(rows + [mislabelled])
|
||||
assert selected[(2026, "FY")] is mislabelled
|
||||
|
||||
|
||||
# -- split safety for the TTM EPS scalar (A5 parity findings) ----------------
|
||||
|
||||
def _split_rows():
|
||||
"""Two years where the share count jumps ~25x at the latest quarter, as
|
||||
BKNG's did (31.7M -> 774.9M) when its split landed mid-window."""
|
||||
rows = _two_years()
|
||||
for row in rows:
|
||||
if (row.fiscal_year, row.fiscal_period) == (2026, "FY"):
|
||||
row.shares_outstanding = 25000.0 # vs 1000 a year earlier
|
||||
return rows
|
||||
|
||||
|
||||
def test_split_suppresses_ttm_diluted_eps():
|
||||
# TTM sums four quarters of per-share values; a split inside the window
|
||||
# mixes units. Unguarded this produced BKNG's P/E of 1.10, which clamps to a
|
||||
# *perfect* fundamental sub-score -- worse than having no value at all.
|
||||
d = fd.derive(_split_rows())
|
||||
assert d.ttm_diluted_eps is None
|
||||
assert d.ttm_diluted_eps_caveat == fd.SPLIT_SENSITIVE_CAVEAT
|
||||
|
||||
|
||||
def test_ttm_diluted_eps_survives_when_no_split_is_suspected():
|
||||
d = fd.derive(_two_years())
|
||||
assert d.ttm_diluted_eps is not None
|
||||
assert d.ttm_diluted_eps_caveat is None
|
||||
|
||||
|
||||
def test_split_guard_leaves_dollar_scalars_alone():
|
||||
# Only per-share values are split-sensitive; FCF is in dollars.
|
||||
baseline = fd.derive(_two_years())
|
||||
d = fd.derive(_split_rows())
|
||||
assert d.ttm_fcf == pytest.approx(baseline.ttm_fcf)
|
||||
|
||||
|
||||
# -- multi-class share-count fallback (A5 parity findings) -------------------
|
||||
|
||||
def test_shares_fall_back_to_weighted_average_when_cover_page_count_is_absent():
|
||||
# META/CMCSA/BRK-B/CHTR report the cover-page count per share class, which is
|
||||
# dimensional and therefore absent from companyfacts -- silently removing
|
||||
# market cap and FCF yield for some of the largest issuers.
|
||||
rows = _two_years()
|
||||
for row in rows:
|
||||
row.shares_outstanding = None
|
||||
row.weighted_avg_diluted_shares = 2_564_000_000.0
|
||||
d = fd.derive(rows)
|
||||
assert d.shares_outstanding == 2_564_000_000.0
|
||||
assert d.shares_outstanding_estimated is True
|
||||
|
||||
|
||||
def test_point_in_time_share_count_is_preferred_and_not_flagged():
|
||||
baseline = fd.derive(_two_years()).shares_outstanding
|
||||
assert baseline is not None, "fixture should carry a cover-page count"
|
||||
rows = _two_years()
|
||||
for row in rows:
|
||||
row.weighted_avg_diluted_shares = 1.0 # must lose to the real count
|
||||
d = fd.derive(rows)
|
||||
assert d.shares_outstanding == baseline
|
||||
assert d.shares_outstanding_estimated is False
|
||||
|
||||
|
||||
def test_weighted_average_fallback_survives_a_partial_amendment():
|
||||
# A Part-III-only 10-K/A on a multi-class issuer's latest period: the merged
|
||||
# row must keep the weighted-average count, or market cap silently vanishes.
|
||||
rows = _two_years()
|
||||
for row in rows:
|
||||
row.shares_outstanding = None
|
||||
row.weighted_avg_diluted_shares = 2_564_000_000.0
|
||||
part_iii_only = Snap(2026, "FY", date(2026, 9, 30), date(2026, 11, 1),
|
||||
datetime(2027, 1, 1, tzinfo=UTC))
|
||||
d = fd.derive(rows + [part_iii_only])
|
||||
assert d.shares_outstanding == 2_564_000_000.0
|
||||
assert d.shares_outstanding_estimated is True
|
||||
|
||||
|
||||
def test_no_share_count_at_all_stays_none_and_unflagged():
|
||||
rows = _two_years()
|
||||
for row in rows:
|
||||
row.shares_outstanding = None
|
||||
d = fd.derive(rows)
|
||||
assert d.shares_outstanding is None
|
||||
assert d.shares_outstanding_estimated is False
|
||||
|
||||
|
||||
def test_merge_lists_cover_every_parser_field():
|
||||
"""_MERGED_FIELDS/_CARRIED_FIELDS are hand-maintained, and _merge_amendments
|
||||
builds the merged row from them alone — so a parser field missing from both
|
||||
is not merely stale on a merged period, it is *absent*, and callers using
|
||||
getattr(row, name, None) read None. That is how weighted_avg_diluted_shares
|
||||
silently lost market cap for multi-class issuers with a partial amendment.
|
||||
Adding a column to SnapshotRow must fail here rather than lose data quietly.
|
||||
"""
|
||||
import dataclasses
|
||||
|
||||
from app.services.sec_facts_parser import SnapshotRow
|
||||
|
||||
parser_fields = {f.name for f in dataclasses.fields(SnapshotRow)}
|
||||
covered = set(fd._MERGED_FIELDS) | set(fd._CARRIED_FIELDS)
|
||||
assert not parser_fields - covered, (
|
||||
f"parser fields not merged or carried: {sorted(parser_fields - covered)}"
|
||||
)
|
||||
|
||||
@@ -0,0 +1,209 @@
|
||||
"""A5 fundamentals parity report: read-only comparison + artifact archive."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import date, datetime, timezone
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import func, select
|
||||
|
||||
from app.models.data_import_run import DataImportRun
|
||||
from app.models.earnings_event import EarningsEvent
|
||||
from app.models.fundamental import FundamentalData
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.ticker import Ticker
|
||||
from app.services.fundamentals_parity_service import (
|
||||
build_report,
|
||||
fundamental_score,
|
||||
load_latest,
|
||||
load_latest_csv,
|
||||
load_latest_json,
|
||||
store_report,
|
||||
)
|
||||
|
||||
UTC = timezone.utc
|
||||
GENERATED = datetime(2026, 7, 23, 10, 30, tzinfo=UTC)
|
||||
|
||||
|
||||
def _snapshot_rows(cik: str) -> list[FundamentalSnapshot]:
|
||||
rows = []
|
||||
periods = ("Q1", "Q2", "Q3", "FY")
|
||||
months = (3, 6, 9, 12)
|
||||
for fy, multiplier in ((2025, 1.0), (2026, 1.1)):
|
||||
revenues = [100 * multiplier, 110 * multiplier, 120 * multiplier, 130 * multiplier]
|
||||
eps = [1.0 * multiplier, 1.1 * multiplier, 1.2 * multiplier, 1.3 * multiplier]
|
||||
for index, period in enumerate(periods):
|
||||
period_end = date(fy, months[index], 28)
|
||||
rows.append(
|
||||
FundamentalSnapshot(
|
||||
cik=cik,
|
||||
accession=f"{cik}-{fy}-{period}",
|
||||
form="10-K" if period == "FY" else "10-Q",
|
||||
filed_date=period_end,
|
||||
accepted_at=datetime(fy, months[index], 28, tzinfo=UTC),
|
||||
period_end=period_end,
|
||||
fiscal_year=fy,
|
||||
fiscal_period=period,
|
||||
revenue=sum(revenues[: index + 1]),
|
||||
operating_income=sum(revenues[: index + 1]) * 0.2,
|
||||
diluted_eps=sum(eps[: index + 1]),
|
||||
cfo=sum(revenues[: index + 1]) * 0.25,
|
||||
capex=sum(revenues[: index + 1]) * 0.05,
|
||||
depreciation_amortization=sum(revenues[: index + 1]) * 0.05,
|
||||
cash_and_st_investments=40,
|
||||
total_debt=100,
|
||||
shares_outstanding=1000,
|
||||
)
|
||||
)
|
||||
return rows
|
||||
|
||||
|
||||
async def _seed(db_session):
|
||||
first = Ticker(symbol="AAA", cik="0000000001", sic="3571")
|
||||
second = Ticker(symbol="BBB", cik=None, sic=None)
|
||||
db_session.add_all([first, second])
|
||||
await db_session.flush()
|
||||
db_session.add_all(_snapshot_rows(first.cik))
|
||||
db_session.add_all(
|
||||
[
|
||||
FundamentalData(
|
||||
ticker_id=first.id,
|
||||
pe_ratio=25,
|
||||
revenue_growth=5,
|
||||
earnings_surprise=0,
|
||||
fetched_at=GENERATED,
|
||||
),
|
||||
FundamentalData(
|
||||
ticker_id=second.id,
|
||||
pe_ratio=12,
|
||||
revenue_growth=3,
|
||||
earnings_surprise=None,
|
||||
fetched_at=GENERATED,
|
||||
),
|
||||
OHLCVRecord(
|
||||
ticker_id=first.id,
|
||||
date=date(2026, 7, 22),
|
||||
open=100,
|
||||
high=100,
|
||||
low=100,
|
||||
close=100,
|
||||
volume=100,
|
||||
),
|
||||
EarningsEvent(
|
||||
ticker_id=first.id,
|
||||
announce_date=date(2026, 7, 1),
|
||||
session="amc",
|
||||
eps_estimate=2,
|
||||
eps_actual=2.2,
|
||||
source="dolt_earnings",
|
||||
),
|
||||
DataImportRun(
|
||||
source="sec_facts",
|
||||
revision="sec-rev",
|
||||
status="promoted",
|
||||
source_max_date=date(2026, 7, 22),
|
||||
started_at=GENERATED,
|
||||
completed_at=GENERATED,
|
||||
),
|
||||
DataImportRun(
|
||||
source="dolt_earnings",
|
||||
revision="dolt-rev",
|
||||
status="no_op",
|
||||
source_max_date=date(2026, 7, 22),
|
||||
started_at=GENERATED,
|
||||
completed_at=GENERATED,
|
||||
),
|
||||
]
|
||||
)
|
||||
await db_session.flush()
|
||||
|
||||
|
||||
def test_score_formula_matches_production_rules():
|
||||
score = fundamental_score(pe_ratio=15, revenue_growth=0, earnings_surprise=0)
|
||||
assert score == pytest.approx((100 + 50 + 50) / 3)
|
||||
assert fundamental_score(pe_ratio=15, revenue_growth=None, earnings_surprise=None) is None
|
||||
|
||||
async def test_report_compares_sources_and_leaves_database_untouched(db_session):
|
||||
await _seed(db_session)
|
||||
before = await db_session.scalar(select(func.count()).select_from(FundamentalData))
|
||||
|
||||
report = await build_report(
|
||||
db_session,
|
||||
generated_at=GENERATED,
|
||||
today=date(2026, 7, 23),
|
||||
)
|
||||
|
||||
after = await db_session.scalar(select(func.count()).select_from(FundamentalData))
|
||||
assert before == after == 2
|
||||
assert not db_session.new and not db_session.dirty and not db_session.deleted
|
||||
assert report["read_only"] is True
|
||||
assert report["approval_status"] == "pending_explicit_approval"
|
||||
assert report["source_runs"]["sec_facts"]["revision"] == "sec-rev"
|
||||
assert report["source_runs"]["dolt_earnings"]["revision"] == "dolt-rev"
|
||||
|
||||
first = next(row for row in report["rows"] if row["symbol"] == "AAA")
|
||||
assert first["fields"]["pe_ratio"]["candidate"] == pytest.approx(
|
||||
100 / 5.06, abs=1e-4
|
||||
)
|
||||
assert first["fields"]["revenue_growth"]["candidate"] == pytest.approx(10)
|
||||
assert first["fields"]["earnings_surprise"]["candidate"] == pytest.approx(10)
|
||||
assert first["scores"]["candidate_fundamental"] is not None
|
||||
assert report["summary"]["universe_count"] == 2
|
||||
assert report["summary"]["field_stats"]["pe_ratio"]["both_available"] == 1
|
||||
|
||||
|
||||
async def test_artifacts_archive_and_latest_manifest(db_session, tmp_path):
|
||||
await _seed(db_session)
|
||||
report = await build_report(
|
||||
db_session,
|
||||
generated_at=GENERATED,
|
||||
today=date(2026, 7, 23),
|
||||
)
|
||||
|
||||
paths = store_report(report, tmp_path)
|
||||
|
||||
assert tmp_path.joinpath("latest.json").exists()
|
||||
assert paths["json"].endswith(".json") and paths["csv"].endswith(".csv")
|
||||
assert load_latest(tmp_path)["generated_at"] == GENERATED.isoformat()
|
||||
csv_artifact = load_latest_csv(tmp_path)
|
||||
assert csv_artifact is not None
|
||||
assert csv_artifact[0].endswith(".csv")
|
||||
assert "legacy_fundamental,candidate_fundamental" in csv_artifact[1]
|
||||
assert "AAA" in csv_artifact[1]
|
||||
json_artifact = load_latest_json(tmp_path)
|
||||
assert json_artifact is not None and '"rows"' in json_artifact[1]
|
||||
|
||||
|
||||
async def test_admin_endpoints_return_compact_summary_and_downloads(
|
||||
client, db_session, tmp_path, monkeypatch
|
||||
):
|
||||
from app.config import settings
|
||||
from app.dependencies import require_admin
|
||||
from app.main import app
|
||||
|
||||
await _seed(db_session)
|
||||
report = await build_report(
|
||||
db_session,
|
||||
generated_at=GENERATED,
|
||||
today=date(2026, 7, 23),
|
||||
)
|
||||
store_report(report, tmp_path)
|
||||
monkeypatch.setattr(settings, "fundamentals_parity_report_dir", str(tmp_path))
|
||||
app.dependency_overrides[require_admin] = lambda: None
|
||||
try:
|
||||
summary_response = await client.get("/api/v1/admin/fundamentals-parity")
|
||||
assert summary_response.status_code == 200
|
||||
summary = summary_response.json()["data"]
|
||||
assert summary["summary"]["universe_count"] == 2
|
||||
assert "rows" not in summary
|
||||
|
||||
csv_response = await client.get("/api/v1/admin/fundamentals-parity/csv")
|
||||
assert csv_response.status_code == 200
|
||||
assert "AAA" in csv_response.json()["data"]["content"]
|
||||
|
||||
json_response = await client.get("/api/v1/admin/fundamentals-parity/json")
|
||||
assert json_response.status_code == 200
|
||||
assert '"rows"' in json_response.json()["data"]["content"]
|
||||
finally:
|
||||
app.dependency_overrides.pop(require_admin, None)
|
||||
@@ -0,0 +1,183 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import date, datetime, timezone
|
||||
|
||||
from app.models.data_import_run import DataImportRun
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.sec_filing_gap import SecFilingGap
|
||||
from app.models.settings import SystemSetting
|
||||
from app.models.ticker import Ticker
|
||||
from app.services import fundamentals_quality_service
|
||||
|
||||
|
||||
async def test_latest_sec_validation_blocks_deferred_and_no_history_ciks(
|
||||
db_session,
|
||||
):
|
||||
missing = Ticker(symbol="MISSING", cik="0000000001")
|
||||
no_history = Ticker(symbol="NEWREG", cik="0000000002")
|
||||
healthy = Ticker(symbol="HEALTHY", cik="0000000003")
|
||||
db_session.add_all([missing, no_history, healthy])
|
||||
await db_session.flush()
|
||||
db_session.add(
|
||||
SystemSetting(
|
||||
key="fundamental_data_sec_dolt_cutover_enabled",
|
||||
value="true",
|
||||
)
|
||||
)
|
||||
db_session.add(
|
||||
DataImportRun(
|
||||
source="sec_facts",
|
||||
status="deferred",
|
||||
validation_json=json.dumps({
|
||||
"missing_xbrl": [{"cik": missing.cik, "accession": "MISSING-Q"}],
|
||||
"no_xbrl_filings": [{"cik": no_history.cik}],
|
||||
}),
|
||||
started_at=datetime.now(timezone.utc),
|
||||
)
|
||||
)
|
||||
await db_session.flush()
|
||||
|
||||
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
|
||||
missing.id,
|
||||
no_history.id,
|
||||
}
|
||||
|
||||
|
||||
async def test_sec_quality_gate_is_inactive_before_cutover(db_session):
|
||||
ticker = Ticker(symbol="SHADOW", cik="0000000042")
|
||||
db_session.add(ticker)
|
||||
await db_session.flush()
|
||||
now = datetime.now(timezone.utc)
|
||||
db_session.add(
|
||||
SecFilingGap(
|
||||
cik=ticker.cik,
|
||||
accession="SHADOW-Q",
|
||||
form="10-Q",
|
||||
index_date=date.today(),
|
||||
reason="not_in_companyfacts",
|
||||
first_seen_at=now,
|
||||
last_attempted_at=now,
|
||||
)
|
||||
)
|
||||
await db_session.flush()
|
||||
|
||||
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == set()
|
||||
|
||||
|
||||
|
||||
|
||||
async def test_active_gap_is_blocked_until_a_later_filing_supersedes_it(db_session):
|
||||
ticker = Ticker(symbol="HIST", cik="0000000043")
|
||||
now = datetime.now(timezone.utc)
|
||||
db_session.add_all([
|
||||
ticker,
|
||||
SystemSetting(
|
||||
key="fundamental_data_sec_dolt_cutover_enabled",
|
||||
value="true",
|
||||
),
|
||||
SecFilingGap(
|
||||
cik=ticker.cik,
|
||||
accession="HIST-Q",
|
||||
form="10-Q",
|
||||
index_date=date.today().replace(day=1),
|
||||
reason="coregistrant_facts_rejected",
|
||||
first_seen_at=now,
|
||||
last_attempted_at=now,
|
||||
),
|
||||
])
|
||||
await db_session.flush()
|
||||
|
||||
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
|
||||
ticker.id
|
||||
}
|
||||
|
||||
db_session.add(
|
||||
FundamentalSnapshot(
|
||||
cik=ticker.cik,
|
||||
accession="LATER-Q",
|
||||
form="10-Q",
|
||||
filed_date=date.today(),
|
||||
accepted_at=datetime.now(timezone.utc),
|
||||
period_end=date.today(),
|
||||
fiscal_year=date.today().year,
|
||||
fiscal_period="Q2",
|
||||
)
|
||||
)
|
||||
await db_session.flush()
|
||||
|
||||
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == set()
|
||||
|
||||
|
||||
async def test_gap_without_index_date_uses_first_seen_date_for_supersession(
|
||||
db_session,
|
||||
):
|
||||
ticker = Ticker(symbol="DATELESS", cik="0000000045")
|
||||
first_seen = datetime(2026, 5, 1, 12, tzinfo=timezone.utc)
|
||||
db_session.add_all([
|
||||
ticker,
|
||||
SystemSetting(
|
||||
key="fundamental_data_sec_dolt_cutover_enabled",
|
||||
value="true",
|
||||
),
|
||||
SecFilingGap(
|
||||
cik=ticker.cik,
|
||||
accession="DATELESS-Q",
|
||||
form="10-Q",
|
||||
index_date=None,
|
||||
reason="not_in_companyfacts",
|
||||
first_seen_at=first_seen,
|
||||
last_attempted_at=first_seen,
|
||||
),
|
||||
])
|
||||
await db_session.flush()
|
||||
|
||||
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == {
|
||||
ticker.id
|
||||
}
|
||||
|
||||
db_session.add(
|
||||
FundamentalSnapshot(
|
||||
cik=ticker.cik,
|
||||
accession="LATER-DATELESS-Q",
|
||||
form="10-Q",
|
||||
filed_date=date(2026, 5, 2),
|
||||
accepted_at=datetime(2026, 5, 2, 12, tzinfo=timezone.utc),
|
||||
period_end=date(2026, 3, 31),
|
||||
fiscal_year=2026,
|
||||
fiscal_period="Q1",
|
||||
)
|
||||
)
|
||||
await db_session.flush()
|
||||
|
||||
assert await fundamentals_quality_service.blocked_ticker_ids(db_session) == set()
|
||||
|
||||
|
||||
async def test_ticker_quality_explains_no_xbrl_block(db_session):
|
||||
ticker = Ticker(symbol="NEWREG", cik="0000000044")
|
||||
db_session.add_all([
|
||||
ticker,
|
||||
SystemSetting(
|
||||
key="fundamental_data_sec_dolt_cutover_enabled",
|
||||
value="true",
|
||||
),
|
||||
DataImportRun(
|
||||
source="sec_facts",
|
||||
status="promoted",
|
||||
validation_json=json.dumps({
|
||||
"setup_blocked_ciks": [ticker.cik],
|
||||
"no_xbrl_ciks": [ticker.cik],
|
||||
"no_xbrl_filings": [],
|
||||
}),
|
||||
started_at=datetime.now(timezone.utc),
|
||||
),
|
||||
])
|
||||
await db_session.flush()
|
||||
|
||||
quality = await fundamentals_quality_service.ticker_quality(db_session, "NEWREG")
|
||||
assert quality.eligible is False
|
||||
assert quality.code == "no_xbrl_filings"
|
||||
assert "CIK override" in (quality.message or "")
|
||||
assert await fundamentals_quality_service.ticker_is_eligible(
|
||||
db_session, ticker.id
|
||||
) is False
|
||||
@@ -6,6 +6,8 @@ from datetime import date, timedelta
|
||||
|
||||
import pytest
|
||||
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.settings import IngestionProgress
|
||||
from app.models.ticker import Ticker
|
||||
from app.providers.protocol import OHLCVData
|
||||
from app.services import ingestion_service as svc
|
||||
@@ -18,9 +20,12 @@ async def session():
|
||||
yield s
|
||||
|
||||
|
||||
async def _add_ticker(session, symbol: str) -> None:
|
||||
session.add(Ticker(symbol=symbol))
|
||||
async def _add_ticker(session, symbol: str) -> Ticker:
|
||||
ticker = Ticker(symbol=symbol)
|
||||
session.add(ticker)
|
||||
await session.commit()
|
||||
await session.refresh(ticker)
|
||||
return ticker
|
||||
|
||||
|
||||
def _bars(symbol: str, n: int) -> list[OHLCVData]:
|
||||
@@ -51,6 +56,50 @@ async def test_happy_path_ingests_bars(session):
|
||||
assert result.records_ingested == 3
|
||||
|
||||
|
||||
async def test_incremental_fetch_overlaps_latest_session_and_updates_partial_bar(session):
|
||||
"""Once today exists, a live refresh must fetch and overwrite it again."""
|
||||
ticker = await _add_ticker(session, "LIVE")
|
||||
today = date.today()
|
||||
session.add_all([
|
||||
OHLCVRecord(
|
||||
ticker_id=ticker.id,
|
||||
date=today - timedelta(days=i),
|
||||
open=100.0,
|
||||
high=101.0,
|
||||
low=99.0,
|
||||
close=100.0,
|
||||
volume=1000,
|
||||
)
|
||||
for i in range(200)
|
||||
])
|
||||
session.add(IngestionProgress(ticker_id=ticker.id, last_ingested_date=today))
|
||||
await session.commit()
|
||||
|
||||
provider = MockMarketDataProvider(ohlcv_data=[
|
||||
OHLCVData(
|
||||
ticker="LIVE",
|
||||
date=today,
|
||||
open=100.0,
|
||||
high=124.0,
|
||||
low=99.0,
|
||||
close=123.0,
|
||||
volume=2000,
|
||||
)
|
||||
])
|
||||
result = await svc.fetch_and_ingest(session, provider, "LIVE")
|
||||
|
||||
assert provider.calls == [{
|
||||
"ticker": "LIVE",
|
||||
"start_date": today,
|
||||
"end_date": today,
|
||||
}]
|
||||
assert result.status == "complete"
|
||||
assert result.records_ingested == 1
|
||||
records = await svc.price_service.query_ohlcv(session, "LIVE", today, today)
|
||||
assert records[0].close == 123.0
|
||||
assert records[0].volume == 2000
|
||||
|
||||
|
||||
async def test_empty_fetch_with_existing_history_is_up_to_date(session):
|
||||
# Covered ticker, just no new bars in the window → complete, not no_data.
|
||||
await _add_ticker(session, "BBB")
|
||||
@@ -81,9 +130,34 @@ async def test_empty_fetch_with_stale_history_reports_stale(session):
|
||||
]
|
||||
await svc.fetch_and_ingest(session, MockMarketDataProvider(ohlcv_data=old), "SATS")
|
||||
|
||||
result = await svc.fetch_and_ingest(session, MockMarketDataProvider(ohlcv_data=[]), "SATS")
|
||||
# Incremental overlap means Alpaca can keep returning the final historical
|
||||
# bar. That is still stale: the latest session did not advance.
|
||||
result = await svc.fetch_and_ingest(
|
||||
session,
|
||||
MockMarketDataProvider(ohlcv_data=[old[-1]]),
|
||||
"SATS",
|
||||
)
|
||||
|
||||
assert result.status == "stale"
|
||||
assert result.records_ingested == 0
|
||||
assert result.records_ingested == 1
|
||||
assert result.last_date is not None
|
||||
assert "renamed" in (result.message or "").lower() or "halted" in (result.message or "").lower()
|
||||
|
||||
|
||||
async def test_ingest_can_skip_sr_refresh_when_scanner_follows(session, monkeypatch):
|
||||
await _add_ticker(session, "SCAN")
|
||||
calls: list[str] = []
|
||||
|
||||
async def fake_refresh(db, symbol):
|
||||
calls.append(symbol)
|
||||
|
||||
monkeypatch.setattr(svc, "_refresh_structural_sr", fake_refresh)
|
||||
result = await svc.fetch_and_ingest(
|
||||
session,
|
||||
MockMarketDataProvider(ohlcv_data=_bars("SCAN", 3)),
|
||||
"SCAN",
|
||||
refresh_sr=False,
|
||||
)
|
||||
|
||||
assert result.status == "complete"
|
||||
assert calls == []
|
||||
|
||||
@@ -56,6 +56,116 @@ async def test_create_and_list_open(session):
|
||||
assert row["symbol"] == "AAA"
|
||||
assert row["status"] == "open"
|
||||
assert row["current_price"] == 110.0 # marked to the latest close
|
||||
assert row["sessions_held"] == 0
|
||||
assert row["sessions_remaining"] == 30
|
||||
|
||||
|
||||
async def test_list_open_counts_post_entry_sessions_for_max_hold(session):
|
||||
await svc.set_exit_policy(session, mode="atr_trailing", hold_days=5)
|
||||
ticker_id = await _seed(session, "COUNT", close=110.0)
|
||||
trade = await svc.create_trade(
|
||||
session,
|
||||
1,
|
||||
symbol="COUNT",
|
||||
direction="long",
|
||||
entry_price=100.0,
|
||||
shares=10,
|
||||
stop_loss=95.0,
|
||||
target=120.0,
|
||||
)
|
||||
today = _today()
|
||||
trade.opened_at = datetime.combine(
|
||||
today - timedelta(days=5), datetime.min.time(), tzinfo=timezone.utc
|
||||
)
|
||||
session.add_all([
|
||||
OHLCVRecord(
|
||||
ticker_id=ticker_id,
|
||||
date=today - timedelta(days=4),
|
||||
open=101,
|
||||
high=102,
|
||||
low=100,
|
||||
close=101,
|
||||
volume=1,
|
||||
),
|
||||
OHLCVRecord(
|
||||
ticker_id=ticker_id,
|
||||
date=today - timedelta(days=2),
|
||||
open=102,
|
||||
high=103,
|
||||
low=101,
|
||||
close=102,
|
||||
volume=1,
|
||||
),
|
||||
])
|
||||
await session.commit()
|
||||
|
||||
row = (await svc.list_trades(session, 1, status="open"))[0]
|
||||
# Two added bars plus today's seeded bar; skipped calendar dates do not count.
|
||||
assert row["sessions_held"] == 3
|
||||
assert row["sessions_remaining"] == 2
|
||||
|
||||
|
||||
async def test_list_open_exposes_past_max_hold_after_policy_is_shortened(session):
|
||||
await svc.set_exit_policy(session, mode="time", hold_days=2)
|
||||
ticker_id = await _seed(session, "OVERDUE", close=110.0)
|
||||
trade = await svc.create_trade(
|
||||
session,
|
||||
1,
|
||||
symbol="OVERDUE",
|
||||
direction="long",
|
||||
entry_price=100.0,
|
||||
shares=10,
|
||||
stop_loss=95.0,
|
||||
target=120.0,
|
||||
)
|
||||
today = _today()
|
||||
trade.opened_at = datetime.combine(
|
||||
today - timedelta(days=5), datetime.min.time(), tzinfo=timezone.utc
|
||||
)
|
||||
session.add_all([
|
||||
OHLCVRecord(
|
||||
ticker_id=ticker_id,
|
||||
date=today - timedelta(days=4),
|
||||
open=101,
|
||||
high=102,
|
||||
low=100,
|
||||
close=101,
|
||||
volume=1,
|
||||
),
|
||||
OHLCVRecord(
|
||||
ticker_id=ticker_id,
|
||||
date=today - timedelta(days=2),
|
||||
open=102,
|
||||
high=103,
|
||||
low=101,
|
||||
close=102,
|
||||
volume=1,
|
||||
),
|
||||
])
|
||||
await session.commit()
|
||||
|
||||
row = (await svc.list_trades(session, 1, status="open"))[0]
|
||||
assert row["sessions_held"] == 3
|
||||
assert row["sessions_remaining"] == -1
|
||||
|
||||
|
||||
async def test_list_open_omits_countdown_without_max_hold_policy(session):
|
||||
await svc.set_exit_policy(session, mode="trailing")
|
||||
await _seed(session, "NOHOLD", close=110.0)
|
||||
await svc.create_trade(
|
||||
session,
|
||||
1,
|
||||
symbol="NOHOLD",
|
||||
direction="long",
|
||||
entry_price=100.0,
|
||||
shares=10,
|
||||
stop_loss=95.0,
|
||||
target=120.0,
|
||||
)
|
||||
|
||||
row = (await svc.list_trades(session, 1, status="open"))[0]
|
||||
assert row["sessions_held"] is None
|
||||
assert row["sessions_remaining"] is None
|
||||
|
||||
|
||||
async def test_create_trade_enforces_post_stop_gate_reset_at_service_boundary(session):
|
||||
|
||||
@@ -0,0 +1,905 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import pickle
|
||||
import sqlite3
|
||||
from datetime import date, timedelta
|
||||
|
||||
import pytest
|
||||
|
||||
from app.services import backtest_service as bt
|
||||
from scripts.portfolio_capacity_research import (
|
||||
ANCHOR_YEARS,
|
||||
RISK_FLOOR_ARMS,
|
||||
aggregate_results,
|
||||
bootstrap_median_interval,
|
||||
build_cells,
|
||||
build_cohort_manifest,
|
||||
iqr,
|
||||
summarize_simulation,
|
||||
validate_cohort_manifest,
|
||||
)
|
||||
from scripts.run_portfolio_construction_matrix import (
|
||||
CACHE_VERSION,
|
||||
STUDIES,
|
||||
_assert_clean_worktree,
|
||||
_build_candidate_cache,
|
||||
_checkpoint_state,
|
||||
_construction_candidate_view,
|
||||
_construction_universe_errors,
|
||||
_json_hash,
|
||||
_load_snapshot,
|
||||
_markdown,
|
||||
_operational_summary,
|
||||
_risk_floor_markdown,
|
||||
_worker_init,
|
||||
_worker_run_cell,
|
||||
_write_cell_checkpoint,
|
||||
)
|
||||
|
||||
|
||||
def _prices(ords: list[int], close: float = 100.0) -> tuple:
|
||||
closes = [close] * len(ords)
|
||||
return (
|
||||
ords,
|
||||
list(closes),
|
||||
[value + 1.0 for value in closes],
|
||||
[value - 1.0 for value in closes],
|
||||
list(closes),
|
||||
[1_000_000] * len(ords),
|
||||
)
|
||||
|
||||
|
||||
def _candidate(
|
||||
symbol: str,
|
||||
day: date,
|
||||
*,
|
||||
entry: float = 100.0,
|
||||
stop: float = 80.0,
|
||||
rank: float = 90.0,
|
||||
) -> dict:
|
||||
return {
|
||||
'qualified': True,
|
||||
'direction': 'long',
|
||||
'symbol': symbol,
|
||||
'date': day.isoformat(),
|
||||
'entry': entry,
|
||||
'stop': stop,
|
||||
'target': entry + 100.0,
|
||||
'momentum_percentile': rank,
|
||||
'activation_momentum_percentile': rank,
|
||||
'residual_high_vol_blend_80_20': rank,
|
||||
}
|
||||
|
||||
|
||||
def _business_days(start: date, end: date) -> list[date]:
|
||||
days: list[date] = []
|
||||
current = start
|
||||
while current <= end:
|
||||
if current.weekday() < 5:
|
||||
days.append(current)
|
||||
current += timedelta(days=1)
|
||||
return days
|
||||
|
||||
|
||||
def test_new_simulator_option_defaults_match_explicit_defaults():
|
||||
start = date(2025, 1, 6)
|
||||
ords = [start.toordinal() + offset for offset in range(8)]
|
||||
prices = {'AAA': _prices(ords)}
|
||||
candidates = [_candidate('AAA', start)]
|
||||
|
||||
legacy = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
3,
|
||||
include_trades=True,
|
||||
)
|
||||
explicit = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
3,
|
||||
max_positions=10,
|
||||
min_initial_risk_fraction=None,
|
||||
weekly_top_n_rebalance=False,
|
||||
measurement_start_date=None,
|
||||
hard_end_date=None,
|
||||
include_capacity_diagnostics=False,
|
||||
include_trades=True,
|
||||
)
|
||||
|
||||
assert legacy == explicit
|
||||
|
||||
|
||||
def test_load_snapshot_accepts_pre_sec_ticker_schema(tmp_path, monkeypatch):
|
||||
snapshot = tmp_path / 'legacy-research.sqlite'
|
||||
with sqlite3.connect(snapshot) as connection:
|
||||
connection.executescript(
|
||||
'''
|
||||
CREATE TABLE tickers (
|
||||
id INTEGER PRIMARY KEY,
|
||||
symbol VARCHAR(10) NOT NULL UNIQUE,
|
||||
name VARCHAR(120),
|
||||
created_at DATETIME NOT NULL
|
||||
);
|
||||
CREATE TABLE ohlcv_records (
|
||||
id INTEGER PRIMARY KEY,
|
||||
ticker_id INTEGER NOT NULL,
|
||||
date DATE NOT NULL,
|
||||
open FLOAT NOT NULL,
|
||||
high FLOAT NOT NULL,
|
||||
low FLOAT NOT NULL,
|
||||
close FLOAT NOT NULL,
|
||||
volume BIGINT NOT NULL,
|
||||
created_at DATETIME NOT NULL
|
||||
);
|
||||
CREATE TABLE research_rank_only (
|
||||
symbol VARCHAR(10) PRIMARY KEY
|
||||
);
|
||||
INSERT INTO tickers VALUES
|
||||
(1, 'LEGACY', 'Legacy Co', '2024-01-01 00:00:00'),
|
||||
(2, 'RANK', 'Rank Only Co', '2024-01-01 00:00:00');
|
||||
INSERT INTO ohlcv_records VALUES
|
||||
(1, 1, '2024-01-02', 100, 102, 99, 101, 1000000,
|
||||
'2024-01-02 00:00:00'),
|
||||
(2, 2, '2024-01-02', 50, 51, 49, 50, 500000,
|
||||
'2024-01-02 00:00:00');
|
||||
INSERT INTO research_rank_only VALUES ('RANK');
|
||||
'''
|
||||
)
|
||||
|
||||
async def recommendation_config(_db):
|
||||
return {}
|
||||
|
||||
async def activation_config(_db):
|
||||
return {'min_momentum_percentile': 80.0}
|
||||
|
||||
async def exit_policy(_db):
|
||||
return {'mode': 'atr_trailing', 'hold_days': 30, 'atr_multiplier': 3.0}
|
||||
|
||||
async def benchmark_closes(_db, *, days, refresh):
|
||||
assert days is None
|
||||
assert refresh is False
|
||||
return {date(2024, 1, 2): 100.0}
|
||||
|
||||
monkeypatch.setattr(
|
||||
'app.services.recommendation_service.get_recommendation_config',
|
||||
recommendation_config,
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
'app.services.admin_service.get_activation_config',
|
||||
activation_config,
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
'app.services.paper_trade_service.get_exit_policy',
|
||||
exit_policy,
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
'app.services.backtest_service._load_benchmark_closes_for_backtest',
|
||||
benchmark_closes,
|
||||
)
|
||||
|
||||
loaded = asyncio.run(_load_snapshot(snapshot, quiet=True))
|
||||
|
||||
assert loaded['symbols'] == ['LEGACY', 'RANK']
|
||||
assert loaded['construction_symbols'] == {'LEGACY'}
|
||||
assert loaded['prices']['LEGACY'] == (
|
||||
[date(2024, 1, 2).toordinal()],
|
||||
[100.0],
|
||||
[102.0],
|
||||
[99.0],
|
||||
[101.0],
|
||||
[1_000_000],
|
||||
)
|
||||
assert loaded['prices']['RANK'][4] == [50.0]
|
||||
assert loaded['construction_universe_manifest'][
|
||||
'construction_ticker_rows'
|
||||
] == 1
|
||||
assert loaded['construction_universe_manifest']['rank_only_ticker_rows'] == 1
|
||||
with sqlite3.connect(snapshot) as connection:
|
||||
columns = {
|
||||
row[1] for row in connection.execute('PRAGMA table_info(tickers)')
|
||||
}
|
||||
assert {'cik', 'sic', 'sic_description'}.isdisjoint(columns)
|
||||
|
||||
|
||||
def test_construction_view_filters_rank_only_rows_without_rebuilding_cache():
|
||||
manifest = {
|
||||
'ranking_ticker_rows': 506,
|
||||
'ranking_symbols_with_prices': 506,
|
||||
'construction_ticker_rows': 505,
|
||||
'construction_symbols_with_prices': 505,
|
||||
'rank_only_ticker_rows': 1,
|
||||
'rank_only_symbols_with_prices': 1,
|
||||
'rank_only_unknown_symbols': 0,
|
||||
}
|
||||
cached = {
|
||||
'key': {'version': 'existing-broad-cache'},
|
||||
'qualified_candidates': [
|
||||
{'symbol': 'PROD', 'date': '2025-01-02'},
|
||||
{'symbol': 'RANK', 'date': '2025-01-02'},
|
||||
],
|
||||
'qualified_long_count': 2,
|
||||
'daily_rank_map': {
|
||||
('RANK', '2025-01-02'): {'strategy_rank': 99.0},
|
||||
},
|
||||
}
|
||||
|
||||
view = _construction_candidate_view(
|
||||
cached,
|
||||
{
|
||||
'construction_symbols': {'PROD'},
|
||||
'construction_universe_manifest': manifest,
|
||||
},
|
||||
)
|
||||
|
||||
assert [row['symbol'] for row in view['qualified_candidates']] == ['PROD']
|
||||
assert view['raw_full_universe_qualified_long_count'] == 2
|
||||
assert view['filtered_rank_only_qualified_long_count'] == 1
|
||||
assert view['qualified_long_count'] == 1
|
||||
assert ('RANK', '2025-01-02') in view['daily_rank_map']
|
||||
assert len(cached['qualified_candidates']) == 2
|
||||
|
||||
|
||||
def test_existing_broad_candidate_cache_key_remains_reusable(tmp_path, monkeypatch):
|
||||
snapshot = tmp_path / 'research.sqlite'
|
||||
snapshot.write_bytes(b'snapshot-placeholder')
|
||||
cache_path = tmp_path / 'broad-cache.pkl'
|
||||
snapshot_data = {
|
||||
'recommendation_config': {'rr': 3.0},
|
||||
'activation': {'min_momentum_percentile': 80.0},
|
||||
'runtime_config': {'ranking_key': 'test'},
|
||||
'universe_manifest': {
|
||||
'ticker_rows': 4655,
|
||||
'symbols_with_prices': 4654,
|
||||
'symbols_sha256': 'symbols',
|
||||
},
|
||||
}
|
||||
key = {
|
||||
'version': CACHE_VERSION,
|
||||
'snapshot': str(snapshot.resolve()),
|
||||
'snapshot_sha256': 'snapshot-hash',
|
||||
'cadence': 'daily',
|
||||
'outcome_horizon_sessions': 0,
|
||||
'recommendation_config_hash': _json_hash(
|
||||
snapshot_data['recommendation_config']
|
||||
),
|
||||
'activation_hash': _json_hash(snapshot_data['activation']),
|
||||
'runtime_config': snapshot_data['runtime_config'],
|
||||
'universe_manifest': snapshot_data['universe_manifest'],
|
||||
}
|
||||
cached = {'key': key, 'qualified_candidates': [{'symbol': 'PROD'}]}
|
||||
cache_path.write_bytes(pickle.dumps(cached))
|
||||
monkeypatch.setattr(
|
||||
bt,
|
||||
'_replay_candidates_for_period',
|
||||
lambda *_args: pytest.fail('existing cache should avoid replay'),
|
||||
)
|
||||
|
||||
loaded = _build_candidate_cache(
|
||||
snapshot_data,
|
||||
snapshot=snapshot,
|
||||
snapshot_sha256='snapshot-hash',
|
||||
cache_path=cache_path,
|
||||
workers=1,
|
||||
quiet=True,
|
||||
)
|
||||
|
||||
assert loaded == cached
|
||||
|
||||
|
||||
def test_construction_universe_guard_rejects_leaked_broad_book():
|
||||
valid = {
|
||||
'ranking_ticker_rows': 4655,
|
||||
'construction_ticker_rows': 506,
|
||||
'construction_symbols_with_prices': 506,
|
||||
'rank_only_ticker_rows': 4149,
|
||||
'rank_only_unknown_symbols': 0,
|
||||
}
|
||||
assert _construction_universe_errors(valid) == []
|
||||
|
||||
leaked = {
|
||||
**valid,
|
||||
'construction_ticker_rows': 4655,
|
||||
'construction_symbols_with_prices': 4654,
|
||||
'rank_only_ticker_rows': 0,
|
||||
}
|
||||
errors = _construction_universe_errors(leaked)
|
||||
assert any('450-600' in error for error in errors)
|
||||
|
||||
|
||||
def test_unbounded_count_and_effective_risk_floor():
|
||||
start = date(2025, 1, 6)
|
||||
ords = [start.toordinal() + offset for offset in range(4)]
|
||||
symbols = [f'S{index}' for index in range(25)]
|
||||
prices = {symbol: _prices(ords) for symbol in symbols}
|
||||
candidates = [
|
||||
_candidate(symbol, start, stop=80.0, rank=100.0 - index)
|
||||
for index, symbol in enumerate(symbols)
|
||||
]
|
||||
|
||||
capped = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
30,
|
||||
max_positions=1,
|
||||
hard_end_date=start + timedelta(days=4),
|
||||
measurement_start_date=start,
|
||||
include_capacity_diagnostics=True,
|
||||
)
|
||||
unbounded = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
30,
|
||||
max_positions=None,
|
||||
min_initial_risk_fraction=0.005,
|
||||
hard_end_date=start + timedelta(days=4),
|
||||
measurement_start_date=start,
|
||||
include_capacity_diagnostics=True,
|
||||
)
|
||||
|
||||
assert capped is not None and unbounded is not None
|
||||
assert capped['peak_positions'] == 1
|
||||
assert capped['measurement_skipped_book_full'] == 24
|
||||
assert unbounded['peak_positions'] > 1
|
||||
assert unbounded['measurement_skipped_book_full'] == 0
|
||||
assert unbounded['skipped_min_initial_risk'] > 0
|
||||
assert unbounded['peak_positions'] == unbounded['trades']
|
||||
|
||||
|
||||
def test_measurement_window_carries_state_but_excludes_pre_anchor_trade_ev():
|
||||
start = date(2025, 1, 6)
|
||||
anchor = start + timedelta(days=2)
|
||||
hard_end = start + timedelta(days=7)
|
||||
ords = [
|
||||
start.toordinal() + offset
|
||||
for offset in range((hard_end - start).days)
|
||||
]
|
||||
prices = {
|
||||
'AAA': _prices(ords, 100.0),
|
||||
'BBB': _prices(ords, 100.0),
|
||||
}
|
||||
candidates = [
|
||||
_candidate('AAA', start),
|
||||
_candidate('BBB', anchor + timedelta(days=1)),
|
||||
]
|
||||
|
||||
sim = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
30,
|
||||
start_date=start,
|
||||
end_date=hard_end,
|
||||
measurement_start_date=anchor,
|
||||
hard_end_date=hard_end,
|
||||
include_curve=True,
|
||||
include_trades=True,
|
||||
)
|
||||
|
||||
assert sim is not None
|
||||
assert sim['simulation_start_date'] == start.isoformat()
|
||||
assert sim['start_date'] == anchor.isoformat()
|
||||
assert sim['measurement_start_positions'] == 1
|
||||
assert sim['trades'] == 1
|
||||
assert [trade['symbol'] for trade in sim['trade_details']] == ['BBB']
|
||||
assert sim['equity_curve'][0]['date'] == anchor.isoformat()
|
||||
|
||||
|
||||
def test_weekly_top10_uses_current_rank_for_both_sides_not_entry_rank():
|
||||
monday = date(2025, 1, 6)
|
||||
friday = date(2025, 1, 10)
|
||||
sessions = _business_days(monday, friday)
|
||||
ords = [session.toordinal() for session in sessions]
|
||||
prices = {
|
||||
'AAA': _prices(ords),
|
||||
'BBB': _prices(ords),
|
||||
}
|
||||
candidates = [
|
||||
_candidate('AAA', monday, rank=99.0),
|
||||
_candidate('BBB', friday, rank=10.0),
|
||||
]
|
||||
rank_map = {
|
||||
('AAA', friday.isoformat()): {'strategy_rank': 10.0},
|
||||
('BBB', friday.isoformat()): {'strategy_rank': 90.0},
|
||||
}
|
||||
|
||||
sim = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
30,
|
||||
max_positions=1,
|
||||
weekly_top_n_rebalance=True,
|
||||
daily_rank_map=rank_map,
|
||||
measurement_start_date=monday,
|
||||
hard_end_date=friday + timedelta(days=1),
|
||||
include_trades=True,
|
||||
include_capacity_diagnostics=True,
|
||||
)
|
||||
|
||||
assert sim is not None
|
||||
assert [trade['symbol'] for trade in sim['trade_details']] == ['AAA', 'BBB']
|
||||
assert sim['trade_details'][0]['reason'] == 'weekly_rebalance'
|
||||
event = sim['weekly_rebalance_events'][0]
|
||||
assert event['exited_symbols'] == ['AAA']
|
||||
assert event['selected_entrant_symbols'] == ['BBB']
|
||||
|
||||
|
||||
def test_weekly_top10_incumbent_wins_exact_current_rank_tie():
|
||||
monday = date(2025, 1, 6)
|
||||
friday = date(2025, 1, 10)
|
||||
sessions = _business_days(monday, friday)
|
||||
ords = [session.toordinal() for session in sessions]
|
||||
prices = {
|
||||
'AAA': _prices(ords),
|
||||
'BBB': _prices(ords),
|
||||
}
|
||||
candidates = [
|
||||
_candidate('AAA', monday, rank=10.0),
|
||||
_candidate('BBB', friday, rank=99.0),
|
||||
]
|
||||
rank_map = {
|
||||
('AAA', friday.isoformat()): {'strategy_rank': 80.0},
|
||||
('BBB', friday.isoformat()): {'strategy_rank': 80.0},
|
||||
}
|
||||
|
||||
sim = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
30,
|
||||
max_positions=1,
|
||||
weekly_top_n_rebalance=True,
|
||||
daily_rank_map=rank_map,
|
||||
measurement_start_date=monday,
|
||||
hard_end_date=friday + timedelta(days=1),
|
||||
include_trades=True,
|
||||
)
|
||||
|
||||
assert sim is not None
|
||||
assert [trade['symbol'] for trade in sim['trade_details']] == ['AAA']
|
||||
assert sim['trade_details'][0]['reason'] == 'open_at_end'
|
||||
assert sim['weekly_rebalance_events'][0]['replacements'] == 0
|
||||
|
||||
|
||||
def test_weekly_rebalance_exit_bypasses_cooldown_and_churn_is_counted():
|
||||
first_monday = date(2025, 1, 6)
|
||||
friday = date(2025, 1, 10)
|
||||
next_monday = date(2025, 1, 13)
|
||||
sessions = _business_days(first_monday, next_monday)
|
||||
ords = [session.toordinal() for session in sessions]
|
||||
prices = {
|
||||
'AAA': _prices(ords),
|
||||
'BBB': (
|
||||
ords,
|
||||
[100.0] * len(ords),
|
||||
[101.0] * len(ords),
|
||||
[99.0] * (len(ords) - 1) + [70.0],
|
||||
[100.0] * len(ords),
|
||||
[1_000_000] * len(ords),
|
||||
),
|
||||
}
|
||||
candidates = [
|
||||
_candidate('AAA', first_monday, rank=99.0),
|
||||
_candidate('BBB', friday, rank=10.0),
|
||||
_candidate('AAA', next_monday, rank=99.0),
|
||||
]
|
||||
rank_map = {
|
||||
('AAA', friday.isoformat()): {'strategy_rank': 10.0},
|
||||
('BBB', friday.isoformat()): {'strategy_rank': 90.0},
|
||||
}
|
||||
|
||||
sim = bt._simulate_portfolio(
|
||||
candidates,
|
||||
prices,
|
||||
None,
|
||||
'hold',
|
||||
30,
|
||||
max_positions=1,
|
||||
reentry_cooldown_sessions=5,
|
||||
weekly_top_n_rebalance=True,
|
||||
daily_rank_map=rank_map,
|
||||
measurement_start_date=first_monday,
|
||||
hard_end_date=next_monday + timedelta(days=1),
|
||||
include_trades=True,
|
||||
)
|
||||
|
||||
assert sim is not None
|
||||
assert [trade['symbol'] for trade in sim['trade_details']] == [
|
||||
'AAA',
|
||||
'BBB',
|
||||
'AAA',
|
||||
]
|
||||
assert sim['trade_details'][0]['reason'] == 'weekly_rebalance'
|
||||
assert sim['rebalance_reentries_within_5_sessions'] == 1
|
||||
assert sim['skipped_cooldown'] == 0
|
||||
|
||||
|
||||
def test_cohort_manifest_realizes_seven_frozen_clusters():
|
||||
sessions = _business_days(date(2016, 1, 4), date(2026, 7, 17))
|
||||
manifest = build_cohort_manifest(sessions)
|
||||
|
||||
assert validate_cohort_manifest(manifest) == []
|
||||
assert manifest['empty_cluster_count'] == 7
|
||||
assert manifest['warm_cluster_count'] == 7
|
||||
assert set(map(int, manifest['empty_cluster_counts'])) == set(ANCHOR_YEARS)
|
||||
assert all(
|
||||
int(count) >= 12 for count in manifest['warm_seed_counts'].values()
|
||||
)
|
||||
cells = build_cells(manifest)
|
||||
assert len(cells) == (
|
||||
len(manifest['empty_book']) + len(manifest['warm_book'])
|
||||
) * 4 * 2
|
||||
floor_cells = build_cells(manifest, arms=RISK_FLOOR_ARMS)
|
||||
assert len(floor_cells) == (
|
||||
len(manifest['empty_book']) + len(manifest['warm_book'])
|
||||
) * 2 * 2
|
||||
assert {row['arm_id'] for row in floor_cells} == {
|
||||
'cap10_incumbent',
|
||||
'cap10_min_risk_005',
|
||||
}
|
||||
|
||||
|
||||
def test_risk_floor_study_changes_only_the_effective_risk_floor():
|
||||
control, treatment = RISK_FLOOR_ARMS
|
||||
|
||||
assert control['max_positions'] == treatment['max_positions'] == 10
|
||||
assert (
|
||||
control['weekly_top_n_rebalance']
|
||||
== treatment['weekly_top_n_rebalance']
|
||||
is False
|
||||
)
|
||||
assert control['min_initial_risk_fraction'] is None
|
||||
assert treatment['min_initial_risk_fraction'] == 0.005
|
||||
assert STUDIES['risk-floor-ab']['arms'] == RISK_FLOOR_ARMS
|
||||
assert STUDIES['capacity-bracket']['arms'] != RISK_FLOOR_ARMS
|
||||
|
||||
|
||||
def test_zero_outcome_horizon_extends_rank_replay_to_last_session(monkeypatch):
|
||||
monkeypatch.setattr(bt, '_window_setups', lambda *_args, **_kwargs: [])
|
||||
count = bt.MIN_LOOKBACK + bt.HORIZON
|
||||
start = date(2025, 1, 1)
|
||||
ords = [start.toordinal() + offset for offset in range(count)]
|
||||
columns = _prices(ords)
|
||||
|
||||
legacy = bt._replay_candidates_for_period(
|
||||
'AAA',
|
||||
columns,
|
||||
{},
|
||||
{},
|
||||
None,
|
||||
date.min,
|
||||
'daily',
|
||||
True,
|
||||
True,
|
||||
)
|
||||
zero_horizon = bt._replay_candidates_for_period(
|
||||
'AAA',
|
||||
columns,
|
||||
{},
|
||||
{},
|
||||
None,
|
||||
date.min,
|
||||
'daily',
|
||||
True,
|
||||
True,
|
||||
0,
|
||||
)
|
||||
|
||||
assert len(zero_horizon) == len(legacy) + bt.HORIZON
|
||||
assert zero_horizon[-1]['date'] == date.fromordinal(ords[-1]).isoformat()
|
||||
|
||||
|
||||
def test_gain_to_pain_uses_all_monthly_returns_and_net_r():
|
||||
sim = {
|
||||
'measurement_start_equity': 100.0,
|
||||
'trade_details': [
|
||||
{
|
||||
'net_r': 1.0,
|
||||
'pnl': 10.0,
|
||||
'shares': 1.0,
|
||||
'entry': 100.0,
|
||||
'fill': 110.0,
|
||||
'transaction_cost': 0.0,
|
||||
},
|
||||
{
|
||||
'net_r': -0.5,
|
||||
'pnl': -5.0,
|
||||
'shares': 1.0,
|
||||
'entry': 100.0,
|
||||
'fill': 95.0,
|
||||
'transaction_cost': 0.0,
|
||||
},
|
||||
],
|
||||
'equity_curve': [
|
||||
{'date': '2025-01-31', 'equity': 110.0},
|
||||
{'date': '2025-02-28', 'equity': 99.0},
|
||||
],
|
||||
'trades': 2,
|
||||
'skipped_book_full': 0,
|
||||
}
|
||||
|
||||
summary = summarize_simulation(sim)
|
||||
|
||||
assert summary['ev_net_r'] == pytest.approx(0.25)
|
||||
assert summary['profit_factor'] == pytest.approx(2.0)
|
||||
# Monthly returns are +10% and -10%; all-return numerator is zero.
|
||||
assert summary['gain_to_pain'] == pytest.approx(0.0)
|
||||
|
||||
|
||||
def test_simple_cluster_bootstrap_is_deterministic_and_not_a_gate():
|
||||
first = bootstrap_median_interval(
|
||||
[1, 2, 3, 4, 5, 6, 7],
|
||||
seed_parts=('determinism',),
|
||||
replicates=500,
|
||||
)
|
||||
second = bootstrap_median_interval(
|
||||
[1, 2, 3, 4, 5, 6, 7],
|
||||
seed_parts=('determinism',),
|
||||
replicates=500,
|
||||
)
|
||||
|
||||
assert first == second
|
||||
assert first['point'] == 4
|
||||
assert first['p05'] <= first['point'] <= first['p95']
|
||||
|
||||
|
||||
def test_iqr_materializes_generator_before_both_quantiles():
|
||||
assert iqr(value for value in (0.0, 1.0, 2.0, 3.0)) == pytest.approx(1.5)
|
||||
|
||||
|
||||
def test_aggregate_reports_paired_years_and_separate_warm_iqrs():
|
||||
cells: list[dict] = []
|
||||
for cost in (0.1, 0.2):
|
||||
for cluster in ANCHOR_YEARS:
|
||||
for seed in range(3):
|
||||
path_id = f'warm-{cluster}-{seed}'
|
||||
for arm_id, shift in (
|
||||
('cap10_incumbent', 0.0),
|
||||
('cash_unbounded', 0.2),
|
||||
('cap10_weekly_top10', 0.1),
|
||||
('cap15_incumbent', 0.05),
|
||||
):
|
||||
cells.append({
|
||||
'arm_id': arm_id,
|
||||
'protocol': 'warm_book',
|
||||
'path_id': path_id,
|
||||
'cluster': cluster,
|
||||
'cost_per_side_pct': cost,
|
||||
'metrics': {
|
||||
'ev_net_r': seed + shift,
|
||||
'calmar': 1.0 + seed * 0.1 + shift,
|
||||
'profit_factor': 1.5 + shift,
|
||||
'gain_to_pain': 2.0 + shift,
|
||||
'sortino': 1.0 + shift,
|
||||
'cagr_pct': 10.0 + shift,
|
||||
'max_drawdown_pct': 5.0,
|
||||
'total_return_pct': 10.0 + shift,
|
||||
'sharpe': 1.0 + shift,
|
||||
},
|
||||
})
|
||||
for arm_id, shift in (
|
||||
('cap10_incumbent', 0.0),
|
||||
('cash_unbounded', 0.2),
|
||||
('cap10_weekly_top10', 0.1),
|
||||
('cap15_incumbent', 0.05),
|
||||
):
|
||||
cells.append({
|
||||
'arm_id': arm_id,
|
||||
'protocol': 'empty_book',
|
||||
'path_id': f'empty-{cluster}',
|
||||
'cluster': cluster,
|
||||
'cost_per_side_pct': cost,
|
||||
'metrics': {
|
||||
'ev_net_r': 1.0 + shift,
|
||||
'calmar': 2.0 + shift,
|
||||
'profit_factor': 1.5 + shift,
|
||||
'gain_to_pain': 2.0 + shift,
|
||||
'sortino': 1.0 + shift,
|
||||
'cagr_pct': 10.0 + shift,
|
||||
'max_drawdown_pct': 5.0,
|
||||
'total_return_pct': 10.0 + shift,
|
||||
'sharpe': 1.0 + shift,
|
||||
},
|
||||
})
|
||||
|
||||
report = aggregate_results(cells)
|
||||
|
||||
cash_empty = next(
|
||||
row
|
||||
for row in report['paired_per_year']
|
||||
if row['arm_id'] == 'cash_unbounded'
|
||||
and row['protocol'] == 'empty_book'
|
||||
and row['cost_per_side_pct'] == 0.1
|
||||
)
|
||||
assert cash_empty['headline']['ev_net_r']['paired_delta_median'] == pytest.approx(
|
||||
0.2
|
||||
)
|
||||
cash_paths = next(
|
||||
row
|
||||
for row in report['paired_path_distributions']
|
||||
if row['arm_id'] == 'cash_unbounded'
|
||||
and row['protocol'] == 'empty_book'
|
||||
and row['cost_per_side_pct'] == 0.1
|
||||
)
|
||||
assert cash_paths['metrics']['ev_net_r']['paired_delta_mean'] == pytest.approx(
|
||||
0.2
|
||||
)
|
||||
assert cash_paths['metrics']['ev_net_r']['positive_fraction'] == 1.0
|
||||
assert cash_paths['metrics']['ev_net_r']['identical_fraction'] == 0.0
|
||||
cash_warm = next(
|
||||
row
|
||||
for row in report['warm_seed_dispersion']
|
||||
if row['arm_id'] == 'cash_unbounded'
|
||||
and row['cost_per_side_pct'] == 0.1
|
||||
)
|
||||
assert set(cash_warm['headline']) == {'ev_net_r', 'calmar'}
|
||||
assert 'D' not in cash_warm
|
||||
assert cash_warm['headline']['ev_net_r']['median_iqr_ratio'] == pytest.approx(
|
||||
1.0
|
||||
)
|
||||
assert cash_warm['headline']['calmar']['median_iqr_ratio'] == pytest.approx(
|
||||
1.0
|
||||
)
|
||||
assert cash_warm['headline']['ev_net_r']['bootstrap_90']['n'] == 7
|
||||
markdown = _markdown({
|
||||
'generated_at': '2026-08-05T00:00:00Z',
|
||||
'analysis': report,
|
||||
'operational_summary': _operational_summary(cells),
|
||||
'validation': {
|
||||
'construction_universe_manifest': {
|
||||
'construction_symbols_with_prices': 506,
|
||||
'rank_only_symbols_with_prices': 4148,
|
||||
'ranking_symbols_with_prices': 4654,
|
||||
},
|
||||
'candidate_rank_coverage': {
|
||||
'construction_qualified_longs': 5000,
|
||||
'filtered_rank_only_qualified_longs': 137000,
|
||||
},
|
||||
},
|
||||
})
|
||||
assert 'ΔGain-to-Pain' in markdown
|
||||
assert '0.10% per fill' in markdown
|
||||
assert '0.20% per fill' in markdown
|
||||
assert 'Tradable setup symbols with prices: 506.' in markdown
|
||||
assert 'Rank-only qualified rows removed: 137000.' in markdown
|
||||
assert 'formal promotion gate' in markdown
|
||||
|
||||
focused_cells = [
|
||||
row
|
||||
for row in cells
|
||||
if row['arm_id'] == 'cap10_incumbent'
|
||||
] + [
|
||||
{
|
||||
**row,
|
||||
'arm_id': 'cap10_min_risk_005',
|
||||
}
|
||||
for row in cells
|
||||
if row['arm_id'] == 'cash_unbounded'
|
||||
]
|
||||
focused_analysis = aggregate_results(
|
||||
focused_cells,
|
||||
arms=RISK_FLOOR_ARMS,
|
||||
include_warm_dispersion=False,
|
||||
)
|
||||
assert focused_analysis['warm_seed_dispersion'] == []
|
||||
focused_markdown = _risk_floor_markdown({
|
||||
'generated_at': '2026-08-05T00:00:00Z',
|
||||
'arms': list(RISK_FLOOR_ARMS),
|
||||
'protocols': ['empty_book', 'warm_book'],
|
||||
'costs_per_side_pct': [0.1, 0.2],
|
||||
'analysis': focused_analysis,
|
||||
'operational_summary': _operational_summary(
|
||||
focused_cells,
|
||||
arms=RISK_FLOOR_ARMS,
|
||||
),
|
||||
})
|
||||
assert '# Effective initial-risk floor A/B' in focused_markdown
|
||||
assert 'Mean dEV' in focused_markdown
|
||||
assert 'Identical' in focused_markdown
|
||||
assert 'Mean dGtP' in focused_markdown
|
||||
assert 'Mean dCalmar/MAR' in focused_markdown
|
||||
assert 'Floor rejects' in focused_markdown
|
||||
assert 'not independent evidence' in focused_markdown
|
||||
|
||||
|
||||
def test_synthetic_worker_matrix_covers_four_arms_protocols_and_costs(monkeypatch):
|
||||
monkeypatch.setenv('BACKTEST_SNAPSHOT_OFFLINE', '0')
|
||||
monkeypatch.setenv('BACKTEST_ALLOW_SPAWN', '0')
|
||||
start = date(2025, 1, 6)
|
||||
sessions = _business_days(start, date(2025, 1, 17))
|
||||
ords = [session.toordinal() for session in sessions]
|
||||
symbols = [f'S{index}' for index in range(12)]
|
||||
prices = {symbol: _prices(ords) for symbol in symbols}
|
||||
candidates = [
|
||||
_candidate(symbol, start, rank=99.0 - index)
|
||||
for index, symbol in enumerate(symbols[:11])
|
||||
]
|
||||
friday = date(2025, 1, 10)
|
||||
candidates.append(_candidate('S11', friday, rank=99.0))
|
||||
rank_map = {
|
||||
(symbol, friday.isoformat()): {
|
||||
'strategy_rank': 100.0 if symbol == 'S11' else float(index)
|
||||
}
|
||||
for index, symbol in enumerate(symbols)
|
||||
}
|
||||
_worker_init({
|
||||
'qualified_candidates': candidates,
|
||||
'daily_rank_map': rank_map,
|
||||
'prices': prices,
|
||||
'benchmark_closes': None,
|
||||
'ranking_key': 'residual_high_vol_blend_80_20',
|
||||
'exit_policy': 'hold',
|
||||
'hold_days': 30,
|
||||
'risk_per_trade': 0.01,
|
||||
'atr_trail_multiplier': 3.0,
|
||||
})
|
||||
|
||||
rows = []
|
||||
for protocol, measurement_start in (
|
||||
('empty_book', start),
|
||||
('warm_book', date(2025, 1, 8)),
|
||||
):
|
||||
for cost in (0.1, 0.2):
|
||||
for arm_id in (
|
||||
'cap10_incumbent',
|
||||
'cash_unbounded',
|
||||
'cap10_weekly_top10',
|
||||
'cap15_incumbent',
|
||||
):
|
||||
rows.append(_worker_run_cell({
|
||||
'cell_id': f'{arm_id}|{protocol}|{cost}',
|
||||
'arm_id': arm_id,
|
||||
'protocol': protocol,
|
||||
'path_id': f'{protocol}-synthetic',
|
||||
'cluster': 2025,
|
||||
'simulation_start': start.isoformat(),
|
||||
'measurement_start': measurement_start.isoformat(),
|
||||
'hard_end_exclusive': date(2025, 1, 14).isoformat(),
|
||||
'cost_per_side_pct': cost,
|
||||
}))
|
||||
|
||||
assert len(rows) == 16
|
||||
assert {row['arm_id'] for row in rows} == {
|
||||
'cap10_incumbent',
|
||||
'cash_unbounded',
|
||||
'cap10_weekly_top10',
|
||||
'cap15_incumbent',
|
||||
}
|
||||
assert {row['protocol'] for row in rows} == {'empty_book', 'warm_book'}
|
||||
assert {row['cost_per_side_pct'] for row in rows} == {0.1, 0.2}
|
||||
assert all('ev_net_r' in row['metrics'] for row in rows)
|
||||
|
||||
|
||||
def test_checkpoint_resume_rejects_fingerprint_mismatch(tmp_path):
|
||||
checkpoint = tmp_path / 'checkpoint'
|
||||
completed = _checkpoint_state(checkpoint, 'fingerprint-a', resume=False)
|
||||
assert completed == {}
|
||||
_write_cell_checkpoint(
|
||||
checkpoint,
|
||||
{'cell_id': 'one', 'metrics': {'ev_net_r': 1.0}},
|
||||
)
|
||||
resumed = _checkpoint_state(checkpoint, 'fingerprint-a', resume=True)
|
||||
assert set(resumed) == {'one'}
|
||||
with pytest.raises(SystemExit, match='fingerprint mismatch'):
|
||||
_checkpoint_state(checkpoint, 'fingerprint-b', resume=True)
|
||||
|
||||
|
||||
def test_dirty_worktree_guard(monkeypatch):
|
||||
monkeypatch.setattr(
|
||||
'scripts.run_portfolio_construction_matrix._git_output',
|
||||
lambda *_args: ' M changed.py',
|
||||
)
|
||||
with pytest.raises(SystemExit, match='dirty worktree'):
|
||||
_assert_clean_worktree()
|
||||
@@ -1,4 +1,4 @@
|
||||
"""Pure-function tests for the v2 Regime Monitor contract."""
|
||||
"""Pure-function tests for the v3 Regime Monitor contract."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -12,23 +12,30 @@ from sqlalchemy import select
|
||||
|
||||
from app.models.regime_snapshot import RegimeSnapshot
|
||||
from app.routers import market as market_router
|
||||
from app.services import regime_monitor_service as rms
|
||||
from app.services import breadth_service, regime_monitor_service as rms
|
||||
from app.services.regime_monitor_service import (
|
||||
DEFAULT_CONFIG,
|
||||
HY_OAS_ELEVATED,
|
||||
HY_OAS_MILD,
|
||||
HY_OAS_STRESSED,
|
||||
STATE_BANDS,
|
||||
WARNING_BANDS,
|
||||
WARNING_WEIGHTS,
|
||||
_compute_index,
|
||||
_fundamental_scores_asof,
|
||||
_score_pillars,
|
||||
band_for,
|
||||
breadth_level_score,
|
||||
drawdown_pct,
|
||||
f2_credit_spreads,
|
||||
fundamental_overlay,
|
||||
p1_trend_break,
|
||||
p2_death_cross,
|
||||
p3_drawdown,
|
||||
p4_relative_strength,
|
||||
p5_volatility,
|
||||
score_warning_sensors,
|
||||
w3_credit_impulse,
|
||||
warning_sensor_scores,
|
||||
)
|
||||
|
||||
|
||||
@@ -39,11 +46,15 @@ def _dated(values: list[float], end: date = date(2026, 6, 26)) -> list[tuple[dat
|
||||
]
|
||||
|
||||
|
||||
def test_band_for_keeps_documented_boundaries():
|
||||
assert band_for(10) == "stable"
|
||||
assert band_for(30) == "watch"
|
||||
assert band_for(60) == "elevated"
|
||||
assert band_for(80) == "breaking"
|
||||
def test_band_for_is_per_axis():
|
||||
assert band_for(10, STATE_BANDS) == "stable"
|
||||
assert band_for(20, STATE_BANDS) == "watch"
|
||||
assert band_for(50, STATE_BANDS) == "elevated"
|
||||
assert band_for(80, STATE_BANDS) == "breaking"
|
||||
# Warning's realized range is far narrower, so it gets its own thresholds.
|
||||
assert band_for(45, STATE_BANDS) == "watch"
|
||||
assert band_for(45, WARNING_BANDS) == "elevated"
|
||||
assert band_for(60, WARNING_BANDS) == "breaking"
|
||||
|
||||
|
||||
def test_price_sensors_are_stress_only():
|
||||
@@ -56,8 +67,97 @@ def test_price_sensors_are_stress_only():
|
||||
assert (p2_death_cross(bearish, bearish) or 0) > 0
|
||||
assert p2_death_cross(healthy, healthy) == 0
|
||||
|
||||
closes = [100.0] * 252 + [80.0]
|
||||
assert p3_drawdown(closes, [100.0] * 253) == 100.0
|
||||
|
||||
def test_drawdown_sensor_keeps_headroom_past_a_twenty_percent_fall():
|
||||
"""v2 pegged at 100 on a 20% drawdown, losing all resolution deeper in."""
|
||||
flat = [100.0] * 253
|
||||
down_20 = [100.0] * 252 + [80.0]
|
||||
down_30 = [100.0] * 252 + [70.0]
|
||||
down_45 = [100.0] * 252 + [55.0]
|
||||
|
||||
assert drawdown_pct(down_20) == pytest.approx(20.0)
|
||||
leader_only_20 = p3_drawdown(down_20, flat)
|
||||
leader_only_30 = p3_drawdown(down_30, flat)
|
||||
assert leader_only_20 < leader_only_30 < 100.0
|
||||
|
||||
# Full scale needs both legs at the deepest anchor, not one at 20%.
|
||||
assert p3_drawdown(down_45, down_45) == 100.0
|
||||
assert p3_drawdown(flat, flat) == 0.0
|
||||
|
||||
|
||||
def test_drawdown_blends_leader_and_confirm_instead_of_taking_the_max():
|
||||
"""max() let the more volatile leader own the whole price pillar."""
|
||||
flat = [100.0] * 253
|
||||
down = [100.0] * 252 + [72.0]
|
||||
both = p3_drawdown(down, down)
|
||||
leader_only = p3_drawdown(down, flat)
|
||||
assert leader_only == pytest.approx(both * 2.0 / 3.0)
|
||||
|
||||
|
||||
def test_credit_impulse_scores_widening_only():
|
||||
assert w3_credit_impulse([3.0] * 40) == 0.0
|
||||
# Tightening is not stress.
|
||||
assert w3_credit_impulse([4.0] * 21 + [3.0]) == 0.0
|
||||
# +35% over the lookback is full scale; half of it is half the score.
|
||||
assert w3_credit_impulse([3.0] * 21 + [3.0 * 1.35]) == pytest.approx(100.0)
|
||||
assert w3_credit_impulse([3.0] * 21 + [3.0 * 1.175]) == pytest.approx(50.0)
|
||||
# Fires while the OAS *level* is still far below the 3.5 mild anchor. This
|
||||
# is the pairing that lets the level stay purely anchored: dynamics live on
|
||||
# the Warning axis rather than being smuggled into State as a percentile.
|
||||
assert f2_credit_spreads([2.0] * 21 + [2.7]) == 0.0
|
||||
assert (w3_credit_impulse([2.0] * 21 + [2.7]) or 0) > 0
|
||||
assert w3_credit_impulse([3.0] * 5) is None
|
||||
|
||||
|
||||
def test_snapshot_records_upstream_history_spans():
|
||||
"""Guards the silent-truncation failure mode that caused this change."""
|
||||
end = date(2026, 6, 26)
|
||||
rising = [100.0 + index * 0.2 for index in range(700)]
|
||||
prices = {"SMH": _dated(rising, end), "QQQ": _dated(rising, end), "SPY": _dated(rising, end)}
|
||||
oas = [(end - timedelta(days=index), 4.0) for index in reversed(range(100))]
|
||||
|
||||
result = _compute_index(
|
||||
prices, [(end, 20.0)], oas, {"f1_score": None, "f3_score": None},
|
||||
copy.deepcopy(DEFAULT_CONFIG), end, [(end, 55.0)], [(end, 20.0)], {end: 25},
|
||||
)
|
||||
|
||||
assert result["data_quality"]["credit_history_days"] == 99
|
||||
assert result["data_quality"]["vix_history_days"] == 0
|
||||
|
||||
|
||||
def test_divergence_still_registers_when_price_confirms_the_breadth_loss():
|
||||
"""v2's hard price gate zeroed this sensor during every decline.
|
||||
|
||||
On 2026-07-24 the basket shed 10 points of participation in 20 sessions
|
||||
while SMH fell 11.9%, and Warning printed exactly 0 as a result.
|
||||
"""
|
||||
days = [date(2026, 1, 1) + timedelta(days=index) for index in range(21)]
|
||||
breadth = {day: 70.0 for day in days[:1]} | {day: 70.0 - index for index, day in enumerate(days)}
|
||||
|
||||
holding = [(day, 100.0) for day in days]
|
||||
falling = [(day, 100.0 - index * 0.9) for index, day in enumerate(days)]
|
||||
|
||||
masked = breadth_service.compute_divergence_series(breadth, holding)[days[-1]]
|
||||
confirmed = breadth_service.compute_divergence_series(breadth, falling)[days[-1]]
|
||||
|
||||
assert masked > confirmed > 0
|
||||
assert confirmed == pytest.approx(masked * breadth_service.DIVERGENCE_CONFIRMED_FLOOR)
|
||||
|
||||
|
||||
def test_warning_score_renormalises_over_available_sensors():
|
||||
full = {"breadth_divergence": 40.0, "relative_strength": 0.0, "credit_impulse": 20.0}
|
||||
assert score_warning_sensors(full) == pytest.approx(
|
||||
(40 * 45 + 0 * 30 + 20 * 25) / 100
|
||||
)
|
||||
partial = {"breadth_divergence": 40.0, "relative_strength": None, "credit_impulse": None}
|
||||
assert score_warning_sensors(partial) == 40.0
|
||||
assert score_warning_sensors(dict.fromkeys(full, None)) is None
|
||||
|
||||
|
||||
def test_warning_sensor_scores_covers_every_weighted_pillar():
|
||||
"""Guards the study/monitor shared definition against silent drift."""
|
||||
sensors = warning_sensor_scores(10.0, [100.0] * 70, [100.0] * 70, [3.0] * 40)
|
||||
assert set(sensors) == set(WARNING_WEIGHTS)
|
||||
|
||||
|
||||
def test_relative_strength_flat_or_better_is_zero():
|
||||
@@ -77,12 +177,24 @@ def test_volatility_and_breadth_zero_points():
|
||||
assert breadth_level_score(None) is None
|
||||
|
||||
|
||||
def test_credit_uses_named_anchors_and_constant_series_is_not_extreme():
|
||||
assert f2_credit_spreads([HY_OAS_MILD] * 100) == 0
|
||||
assert f2_credit_spreads([HY_OAS_ELEVATED] * 100) == 35.0
|
||||
assert f2_credit_spreads([HY_OAS_STRESSED] * 100) == 70.0
|
||||
rising = [3.0 + index * 0.01 for index in range(100)]
|
||||
assert (f2_credit_spreads(rising) or 0) > f2_credit_spreads([3.0] * 100)
|
||||
def test_credit_level_is_anchored_and_ignores_the_reference_window():
|
||||
"""The percentile leg is gone: the anchors already encode the long run.
|
||||
|
||||
It ranked the level against whatever history the upstream series happened to
|
||||
serve, and that silently shrank from 10 years to 3 in April 2026 -- three
|
||||
uniformly tight years, against which an unremarkable spread scored as an
|
||||
extreme. Identical inputs must now score identically regardless of window.
|
||||
"""
|
||||
assert f2_credit_spreads([HY_OAS_MILD] * 100) == 0.0
|
||||
assert f2_credit_spreads([HY_OAS_ELEVATED] * 100) == 50.0
|
||||
assert f2_credit_spreads([HY_OAS_STRESSED] * 100) == 100.0
|
||||
assert f2_credit_spreads([]) is None
|
||||
|
||||
# A level at the "mild" anchor is zero stress even when it tops its window.
|
||||
tight_window = [2.6] * 400 + [HY_OAS_MILD]
|
||||
assert f2_credit_spreads(tight_window) == 0.0
|
||||
# Only the latest observation matters; history cannot move the reading.
|
||||
assert f2_credit_spreads([9.0] * 400 + [3.0]) == f2_credit_spreads([2.6] * 400 + [3.0])
|
||||
|
||||
|
||||
def test_score_pillars_gates_band_below_75_percent_coverage():
|
||||
@@ -98,31 +210,75 @@ def test_score_pillars_gates_band_below_75_percent_coverage():
|
||||
assert result["band"] is None
|
||||
|
||||
|
||||
def test_fundamentals_never_replay_before_effective_date_and_expire():
|
||||
def test_fundamental_overlay_never_replays_before_effective_date_and_expires():
|
||||
overrides = {
|
||||
"f1_score": 0.0,
|
||||
"f3_score": 100.0,
|
||||
"capex": {"GOOGL": "raising"},
|
||||
"good_news_stock_down": "yes",
|
||||
"fetched_at": "2026-06-01T10:00:00+00:00",
|
||||
"effective_date": "2026-06-02",
|
||||
}
|
||||
config = {**DEFAULT_CONFIG, "fundamental_staleness_days": 80}
|
||||
assert _fundamental_scores_asof(overrides, config, date(2026, 6, 1))[:2] == (None, None)
|
||||
assert _fundamental_scores_asof(overrides, config, date(2026, 6, 2))[:2] == (0.0, 100.0)
|
||||
assert _fundamental_scores_asof(overrides, config, date(2026, 8, 22))[:2] == (None, None)
|
||||
|
||||
pending = fundamental_overlay(overrides, config, date(2026, 6, 1))
|
||||
assert pending["pending"] is True
|
||||
assert pending["available"] is False
|
||||
assert pending["capex"] is None
|
||||
# The effective date is still reported so a pending refresh is visible.
|
||||
assert pending["effective_date"] == "2026-06-02"
|
||||
|
||||
live = fundamental_overlay(overrides, config, date(2026, 6, 2))
|
||||
assert live["available"] is True
|
||||
assert live["good_news_stock_down"] == "yes"
|
||||
assert live["earnings_stress"] == 100.0
|
||||
|
||||
expired = fundamental_overlay(overrides, config, date(2026, 8, 22))
|
||||
assert expired["stale"] is True
|
||||
assert expired["available"] is False
|
||||
|
||||
|
||||
def test_capex_score_is_derived_from_company_categories():
|
||||
def test_fundamentals_do_not_move_the_warning_score():
|
||||
"""The v3 complaint: a maxed-out LLM read must not silently do nothing.
|
||||
|
||||
It no longer feeds Warning at all, so Warning is identical either way and
|
||||
the observation is reported beside the score instead of buried in it.
|
||||
"""
|
||||
end = date(2026, 6, 26)
|
||||
rising = [100.0 + index * 0.2 for index in range(700)]
|
||||
prices = {"SMH": _dated(rising, end), "QQQ": _dated(rising, end), "SPY": _dated(rising, end)}
|
||||
args = (prices, [(end, 20.0)], [(end - timedelta(days=i), 4.0) for i in reversed(range(100))])
|
||||
tail = (copy.deepcopy(DEFAULT_CONFIG), end, [(end, 55.0)], [(end, 20.0)], {end: 25})
|
||||
|
||||
quiet = _compute_index(*args, {"f1_score": None, "f3_score": None}, *tail)
|
||||
screaming = _compute_index(
|
||||
*args,
|
||||
{
|
||||
"f1_score": 100.0,
|
||||
"f3_score": 100.0,
|
||||
"capex": dict.fromkeys(DEFAULT_CONFIG["tickers"]["hyperscalers"], "cutting"),
|
||||
"good_news_stock_down": "yes",
|
||||
"effective_date": "2026-06-01",
|
||||
},
|
||||
*tail,
|
||||
)
|
||||
|
||||
assert quiet["warning"]["score"] == screaming["warning"]["score"]
|
||||
assert {p["id"] for p in quiet["warning"]["pillars"]} == set(WARNING_WEIGHTS)
|
||||
assert screaming["fundamental_overlay"]["available"] is True
|
||||
assert screaming["fundamental_overlay"]["capex_stress"] == 100.0
|
||||
|
||||
|
||||
def test_capex_score_separates_holding_from_raising():
|
||||
"""v2 mapped raising and holding both to 0, so a boom read identical to a
|
||||
deceleration and the sensor carried no information."""
|
||||
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
|
||||
assert rms._score_capex_states(dict.fromkeys(names, "raising"), names) == 0.0
|
||||
assert rms._score_capex_states(dict.fromkeys(names, "holding"), names) == 50.0
|
||||
assert rms._score_capex_states(dict.fromkeys(names, "cutting"), names) == 100.0
|
||||
assert rms._score_capex_states(
|
||||
dict.fromkeys(names, "holding"), names
|
||||
) == 0.0
|
||||
assert rms._score_capex_states(
|
||||
{names[0]: "cutting", **dict.fromkeys(names[1:], "holding")}, names
|
||||
) == 25.0
|
||||
assert rms._score_capex_states(
|
||||
{names[0]: "cutting", names[1]: "holding", names[2]: "holding", names[3]: "unknown"},
|
||||
names,
|
||||
) == 33.3
|
||||
{names[0]: "raising", **dict.fromkeys(names[1:], "holding")}, names
|
||||
) == 37.5
|
||||
assert rms._score_capex_states(
|
||||
{names[0]: "cutting", names[1]: "holding", names[2]: "unknown", names[3]: "unknown"},
|
||||
names,
|
||||
@@ -135,7 +291,7 @@ def test_fundamental_api_rejects_numeric_ordinal_overrides():
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_legacy_numeric_fundamentals_do_not_leak_into_v2(monkeypatch):
|
||||
async def test_legacy_numeric_fundamentals_do_not_leak_into_v3(monkeypatch):
|
||||
async def fake_value(_db, _key):
|
||||
return json.dumps({"f1_score": 75.0, "f3_score": 75.0, "source": "manual"})
|
||||
|
||||
@@ -143,16 +299,47 @@ async def test_legacy_numeric_fundamentals_do_not_leak_into_v2(monkeypatch):
|
||||
|
||||
result = await rms.get_fundamental_overrides(object())
|
||||
|
||||
assert result["methodology"] == "v2"
|
||||
assert result["methodology"] == "v3"
|
||||
assert result["f1_score"] is None
|
||||
assert result["f3_score"] is None
|
||||
assert result["good_news_stock_down"] == "mixed"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_v2_observation_survives_the_methodology_bump(monkeypatch):
|
||||
"""A snapshot reseed must not throw away a hand/LLM-collected observation.
|
||||
|
||||
The categorical format is unchanged, so the stored capex map is still valid;
|
||||
only the capex scale moved, and f1 is recomputed from the categories.
|
||||
"""
|
||||
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
|
||||
|
||||
async def fake_value(_db, _key):
|
||||
return json.dumps({
|
||||
"methodology": "v2",
|
||||
"f1_score": 0.0, # stale v2 scale, must be recomputed
|
||||
"f3_score": 100.0,
|
||||
"capex": {names[0]: "raising", **dict.fromkeys(names[1:], "holding")},
|
||||
"good_news_stock_down": "yes",
|
||||
"source": "gemini",
|
||||
"fetched_at": "2026-07-24T14:25:47+00:00",
|
||||
"effective_date": "2026-07-27",
|
||||
})
|
||||
|
||||
monkeypatch.setattr(rms.settings_store, "get_value", fake_value)
|
||||
|
||||
result = await rms.get_fundamental_overrides(object())
|
||||
|
||||
assert result["source"] == "gemini"
|
||||
assert result["good_news_stock_down"] == "yes"
|
||||
assert result["effective_date"] == "2026-07-27"
|
||||
assert result["f1_score"] == 37.5 # recomputed on the v3 scale, not the stored 0.0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_unlock_does_not_redate_a_fundamental_observation(monkeypatch):
|
||||
stored = {
|
||||
"methodology": "v2",
|
||||
"methodology": "v3",
|
||||
"f1_score": 100.0,
|
||||
"f3_score": 0.0,
|
||||
"capex": dict.fromkeys(DEFAULT_CONFIG["tickers"]["hyperscalers"], "cutting"),
|
||||
@@ -185,7 +372,7 @@ async def test_unlock_does_not_redate_a_fundamental_observation(monkeypatch):
|
||||
async def test_manual_fundamentals_are_categorical_and_derived(monkeypatch):
|
||||
names = DEFAULT_CONFIG["tickers"]["hyperscalers"]
|
||||
current = {
|
||||
"methodology": "v2",
|
||||
"methodology": "v3",
|
||||
"f1_score": None,
|
||||
"f3_score": None,
|
||||
"capex": dict.fromkeys(names, "unknown"),
|
||||
@@ -212,7 +399,7 @@ async def test_manual_fundamentals_are_categorical_and_derived(monkeypatch):
|
||||
object(), capex=capex, good_news_stock_down="mixed"
|
||||
)
|
||||
|
||||
assert result["f1_score"] == 25.0
|
||||
assert result["f1_score"] == 62.5 # one cutting (100) + three holding (50)
|
||||
assert result["f3_score"] is None
|
||||
assert result["good_news_stock_down"] == "mixed"
|
||||
assert result["source"] == "manual"
|
||||
@@ -222,10 +409,10 @@ async def test_manual_fundamentals_are_categorical_and_derived(monkeypatch):
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_prior_v2_snapshot_is_immutable_without_explicit_rebuild(db_session):
|
||||
async def test_prior_snapshot_is_immutable_without_explicit_rebuild(db_session):
|
||||
snapshot_date = date(2026, 6, 26)
|
||||
first = {
|
||||
"methodology": "v2",
|
||||
"methodology": "v3",
|
||||
"date": snapshot_date.isoformat(),
|
||||
"state": {"score": 10.0, "band": "stable"},
|
||||
"warning": {"score": 20.0, "band": "stable"},
|
||||
@@ -281,7 +468,7 @@ async def test_routine_can_refresh_latest_trading_session_after_civil_day_rolls(
|
||||
return {}, {}
|
||||
|
||||
async def fake_latest(_db):
|
||||
return object(), {"methodology": "v2"}
|
||||
return object(), {"methodology": "v3"}
|
||||
|
||||
async def fake_upsert(_db, result, *, rewrite_existing_v2):
|
||||
rewrites.append(rewrite_existing_v2)
|
||||
@@ -296,7 +483,7 @@ async def test_routine_can_refresh_latest_trading_session_after_civil_day_rolls(
|
||||
monkeypatch.setattr(rms, "_fetch_prices", fake_prices)
|
||||
monkeypatch.setattr(rms, "_fetch_fred_series", fake_fred)
|
||||
monkeypatch.setattr(rms.breadth_service, "compute_breadth_details", fake_breadth)
|
||||
monkeypatch.setattr(rms, "_latest_v2_row", fake_latest)
|
||||
monkeypatch.setattr(rms, "_latest_snapshot_row", fake_latest)
|
||||
monkeypatch.setattr(rms, "_upsert_snapshot", fake_upsert)
|
||||
|
||||
result = await rms.update_regime_monitor(FakeDB())
|
||||
@@ -364,6 +551,6 @@ def test_compute_index_uses_one_max_price_vote_and_has_no_combined_score():
|
||||
price = next(p for p in result["state"]["pillars"] if p["id"] == "price")
|
||||
sensor_scores = [sensor["score"] for sensor in price["sensors"] if sensor["score"] is not None]
|
||||
assert price["score"] == max(sensor_scores)
|
||||
assert result["methodology"] == "v2"
|
||||
assert result["methodology"] == "v3"
|
||||
assert "combined" not in result
|
||||
assert result["basket"]["members_available"] == 25
|
||||
|
||||
@@ -1,36 +1,46 @@
|
||||
"""Tests for v2 State/Warning quadrant hysteresis and basket reseeding keys."""
|
||||
"""Tests for v3 State/Warning quadrant hysteresis and basket reseeding keys.
|
||||
|
||||
v3 dividers are per axis (State 50, Warning 40) because the two scores have
|
||||
different realized ranges -- Warning never exceeded 64.9 in the 408 calibration
|
||||
sessions, so a shared 60 left the whole upper half of that axis unreachable.
|
||||
"""
|
||||
|
||||
from app.services.alert_service import (
|
||||
QUAD_X_DIV,
|
||||
QUAD_Y_DIV,
|
||||
_classify_quadrant,
|
||||
_parse_quadrant_log_key,
|
||||
_quadrant_log_key,
|
||||
)
|
||||
|
||||
|
||||
def test_fresh_classification_uses_60_60_boundaries():
|
||||
def test_fresh_classification_uses_per_axis_boundaries():
|
||||
assert (QUAD_X_DIV, QUAD_Y_DIV) == (50.0, 40.0)
|
||||
assert _classify_quadrant(20, 90, None) == "1"
|
||||
assert _classify_quadrant(70, 90, None) == "2"
|
||||
assert _classify_quadrant(20, 30, None) == "3"
|
||||
assert _classify_quadrant(70, 30, None) == "4"
|
||||
# A Warning of 45 is above its own divider but below State's.
|
||||
assert _classify_quadrant(45, 45, None) == "1"
|
||||
|
||||
|
||||
def test_warning_axis_hysteresis():
|
||||
assert _classify_quadrant(20, 62, prev="3") == "3"
|
||||
assert _classify_quadrant(20, 66, prev="3") == "1"
|
||||
assert _classify_quadrant(20, 58, prev="1") == "1"
|
||||
assert _classify_quadrant(20, 54, prev="1") == "3"
|
||||
assert _classify_quadrant(20, 42, prev="3") == "3"
|
||||
assert _classify_quadrant(20, 46, prev="3") == "1"
|
||||
assert _classify_quadrant(20, 38, prev="1") == "1"
|
||||
assert _classify_quadrant(20, 34, prev="1") == "3"
|
||||
|
||||
|
||||
def test_state_axis_hysteresis():
|
||||
assert _classify_quadrant(63, 30, prev="3") == "3"
|
||||
assert _classify_quadrant(66, 30, prev="3") == "4"
|
||||
assert _classify_quadrant(57, 30, prev="4") == "4"
|
||||
assert _classify_quadrant(54, 30, prev="4") == "3"
|
||||
assert _classify_quadrant(53, 30, prev="3") == "3"
|
||||
assert _classify_quadrant(56, 30, prev="3") == "4"
|
||||
assert _classify_quadrant(47, 30, prev="4") == "4"
|
||||
assert _classify_quadrant(44, 30, prev="4") == "3"
|
||||
|
||||
|
||||
def test_boundary_sitting_does_not_flip():
|
||||
for quadrant in ("1", "2", "3", "4"):
|
||||
assert _classify_quadrant(60, 60, prev=quadrant) == quadrant
|
||||
assert _classify_quadrant(QUAD_X_DIV, QUAD_Y_DIV, prev=quadrant) == quadrant
|
||||
|
||||
|
||||
def test_quadrant_key_carries_basket_hash_and_parses_legacy_keys():
|
||||
|
||||
@@ -1,285 +0,0 @@
|
||||
"""Regression: scanner must not headline the most distant (max raw R:R) level.
|
||||
|
||||
Historical bug: provisional candidate pick used max R:R / quality only. Production
|
||||
headline is probability-based primary after enhance_trade_setup — near levels
|
||||
with real reach-probability beat far lotteries.
|
||||
|
||||
**Validates: Requirements 1.1, 1.3, 1.4, 2.1, 2.3, 2.4**
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import date, timedelta
|
||||
|
||||
import pytest
|
||||
from hypothesis import given, settings, HealthCheck, strategies as st
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.sr_level import SRLevel
|
||||
from app.models.ticker import Ticker
|
||||
from app.services.rr_scanner_service import scan_ticker
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Session fixture that allows scan_ticker to commit
|
||||
# ---------------------------------------------------------------------------
|
||||
# The default db_session fixture wraps in session.begin() which conflicts
|
||||
# with scan_ticker's internal commit(). We use a plain session instead.
|
||||
|
||||
@pytest.fixture
|
||||
async def scan_session() -> AsyncSession:
|
||||
"""Provide a DB session compatible with scan_ticker (which commits)."""
|
||||
from tests.conftest import _test_session_factory
|
||||
|
||||
async with _test_session_factory() as session:
|
||||
yield session
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _make_ohlcv_bars(
|
||||
ticker_id: int,
|
||||
num_bars: int = 20,
|
||||
base_close: float = 100.0,
|
||||
) -> list[OHLCVRecord]:
|
||||
"""Generate realistic OHLCV bars with small daily variation.
|
||||
|
||||
Produces bars where close ≈ base_close, with enough range for ATR
|
||||
computation (needs >= 15 bars). The ATR will be roughly 2.0.
|
||||
"""
|
||||
bars: list[OHLCVRecord] = []
|
||||
start = date(2024, 1, 1)
|
||||
for i in range(num_bars):
|
||||
close = base_close + (i % 3 - 1) * 0.5 # oscillate ±0.5
|
||||
bars.append(OHLCVRecord(
|
||||
ticker_id=ticker_id,
|
||||
date=start + timedelta(days=i),
|
||||
open=close - 0.3,
|
||||
high=close + 1.0,
|
||||
low=close - 1.0,
|
||||
close=close,
|
||||
volume=100_000,
|
||||
))
|
||||
return bars
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Deterministic test: strong-near vs weak-far (long setup)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_long_prefers_strong_near_over_weak_far(scan_session: AsyncSession):
|
||||
"""With a strong nearby resistance and a weak distant resistance, the
|
||||
probability primary should be the nearby level — NOT the far lottery.
|
||||
"""
|
||||
ticker = Ticker(symbol="EXPLR")
|
||||
scan_session.add(ticker)
|
||||
await scan_session.flush()
|
||||
|
||||
# 20 bars closing around 100
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
scan_session.add_all(bars)
|
||||
|
||||
# With ATR=2.0 and multiplier=1.5, risk=3.0.
|
||||
# R:R threshold=1.5 → min reward=4.5 → min target=104.5
|
||||
# Strong nearby resistance: price=105, strength=90 (R:R≈1.67, quality≈0.66)
|
||||
near_level = SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=105.0,
|
||||
type="resistance",
|
||||
strength=90,
|
||||
detection_method="volume_profile",
|
||||
)
|
||||
# Weak distant resistance: price=130, strength=5 (R:R=10, quality≈0.58)
|
||||
far_level = SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=130.0,
|
||||
type="resistance",
|
||||
strength=5,
|
||||
detection_method="volume_profile",
|
||||
)
|
||||
scan_session.add_all([near_level, far_level])
|
||||
await scan_session.flush()
|
||||
|
||||
setups = await scan_ticker(
|
||||
scan_session,
|
||||
"EXPLR",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=[near_level, far_level],
|
||||
)
|
||||
|
||||
long_setups = [s for s in setups if s.direction == "long"]
|
||||
assert len(long_setups) == 1, "Expected exactly one long setup"
|
||||
|
||||
selected_target = long_setups[0].target
|
||||
# The scanner must NOT pick the most distant level (130)
|
||||
assert selected_target != pytest.approx(130.0, abs=0.01), (
|
||||
"Bug: scanner picked the weak distant level (130) instead of the "
|
||||
"strong nearby level (105)"
|
||||
)
|
||||
# Probability primary should pick the strong nearby level
|
||||
assert selected_target == pytest.approx(105.0, abs=0.01)
|
||||
primaries = [t for t in long_setups[0].targets if t.get("is_primary")]
|
||||
assert len(primaries) == 1
|
||||
assert primaries[0]["price"] == pytest.approx(105.0, abs=0.01)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Deterministic test: strong-near vs weak-far (short setup)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_short_prefers_strong_near_over_weak_far(scan_session: AsyncSession):
|
||||
"""Short-side mirror: strong nearby support should be preferred over
|
||||
weak distant support.
|
||||
"""
|
||||
ticker = Ticker(symbol="EXPLS")
|
||||
scan_session.add(ticker)
|
||||
await scan_session.flush()
|
||||
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
scan_session.add_all(bars)
|
||||
|
||||
# With ATR=2.0 and multiplier=1.5, risk=3.0.
|
||||
# R:R threshold=1.5 → min reward=4.5 → min target below 95.5
|
||||
# Strong nearby support: price=95, strength=85 (R:R≈1.67, quality≈0.64)
|
||||
near_level = SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=95.0,
|
||||
type="support",
|
||||
strength=85,
|
||||
detection_method="pivot_point",
|
||||
)
|
||||
# Weak distant support: price=70, strength=5 (R:R=10, quality≈0.58)
|
||||
far_level = SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=70.0,
|
||||
type="support",
|
||||
strength=5,
|
||||
detection_method="pivot_point",
|
||||
)
|
||||
scan_session.add_all([near_level, far_level])
|
||||
await scan_session.flush()
|
||||
|
||||
setups = await scan_ticker(
|
||||
scan_session,
|
||||
"EXPLS",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=[near_level, far_level],
|
||||
)
|
||||
|
||||
short_setups = [s for s in setups if s.direction == "short"]
|
||||
assert len(short_setups) == 1, "Expected exactly one short setup"
|
||||
|
||||
selected_target = short_setups[0].target
|
||||
assert selected_target != pytest.approx(70.0, abs=0.01), (
|
||||
"Bug: scanner picked the weak distant level (70) instead of the "
|
||||
"strong nearby level (95)"
|
||||
)
|
||||
assert selected_target == pytest.approx(95.0, abs=0.01)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Hypothesis property test: selection is NOT always the most distant level
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@st.composite
|
||||
def strong_near_weak_far_pair(draw: st.DrawFn) -> dict:
|
||||
"""Generate a (strong-near, weak-far) resistance pair above entry=100.
|
||||
|
||||
Guarantees:
|
||||
- near_price < far_price (both above entry)
|
||||
- near_strength >> far_strength
|
||||
- Both meet the R:R threshold of 1.5 given typical ATR ≈ 2 → risk ≈ 3
|
||||
"""
|
||||
# Near level: 5–15 above entry (R:R ≈ 1.7–5.0 with risk≈3)
|
||||
near_dist = draw(st.floats(min_value=5.0, max_value=15.0))
|
||||
near_strength = draw(st.integers(min_value=70, max_value=100))
|
||||
|
||||
# Far level: 25–60 above entry (R:R ≈ 8.3–20 with risk≈3)
|
||||
far_dist = draw(st.floats(min_value=25.0, max_value=60.0))
|
||||
far_strength = draw(st.integers(min_value=1, max_value=15))
|
||||
|
||||
return {
|
||||
"near_price": 100.0 + near_dist,
|
||||
"near_strength": near_strength,
|
||||
"far_price": 100.0 + far_dist,
|
||||
"far_strength": far_strength,
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@given(pair=strong_near_weak_far_pair())
|
||||
@settings(
|
||||
max_examples=15,
|
||||
deadline=None,
|
||||
suppress_health_check=[HealthCheck.function_scoped_fixture],
|
||||
)
|
||||
async def test_property_scanner_does_not_always_pick_most_distant(
|
||||
pair: dict,
|
||||
scan_session: AsyncSession,
|
||||
):
|
||||
"""**Validates: Requirements 1.1, 1.3, 1.4, 2.1, 2.3, 2.4**
|
||||
|
||||
Property: when a strong nearby resistance exists alongside a weak distant
|
||||
resistance, the scanner does NOT always select the most distant level.
|
||||
|
||||
On unfixed code this would fail for every example because max-R:R always
|
||||
picks the farthest level.
|
||||
"""
|
||||
from tests.conftest import _test_engine, _test_session_factory
|
||||
|
||||
# Each hypothesis example needs a fresh DB state
|
||||
async with _test_engine.begin() as conn:
|
||||
from app.database import Base
|
||||
await conn.run_sync(Base.metadata.drop_all)
|
||||
await conn.run_sync(Base.metadata.create_all)
|
||||
|
||||
async with _test_session_factory() as session:
|
||||
ticker = Ticker(symbol="PROP")
|
||||
session.add(ticker)
|
||||
await session.flush()
|
||||
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
session.add_all(bars)
|
||||
|
||||
near_level = SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=pair["near_price"],
|
||||
type="resistance",
|
||||
strength=pair["near_strength"],
|
||||
detection_method="volume_profile",
|
||||
)
|
||||
far_level = SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=pair["far_price"],
|
||||
type="resistance",
|
||||
strength=pair["far_strength"],
|
||||
detection_method="volume_profile",
|
||||
)
|
||||
session.add_all([near_level, far_level])
|
||||
await session.commit()
|
||||
|
||||
setups = await scan_ticker(
|
||||
session,
|
||||
"PROP",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=[near_level, far_level],
|
||||
)
|
||||
|
||||
long_setups = [s for s in setups if s.direction == "long"]
|
||||
assert len(long_setups) == 1, "Expected exactly one long setup"
|
||||
|
||||
selected_target = long_setups[0].target
|
||||
most_distant = round(pair["far_price"], 4)
|
||||
|
||||
# The fixed scanner should prefer the strong nearby level, not the
|
||||
# most distant weak one.
|
||||
assert selected_target != pytest.approx(most_distant, abs=0.01), (
|
||||
f"Bug: scanner picked the most distant level ({most_distant}) "
|
||||
f"with strength={pair['far_strength']} over the nearby level "
|
||||
f"({round(pair['near_price'], 4)}) with strength={pair['near_strength']}"
|
||||
)
|
||||
@@ -1,375 +0,0 @@
|
||||
"""Fix-checking tests for R:R scanner probability-based primary selection.
|
||||
|
||||
Verify that after enhance_trade_setup the headline target is the most likely
|
||||
worthwhile primary (R:R + probability floors), for both long and short setups.
|
||||
The pre-enhance quality loop only seeds a provisional target.
|
||||
|
||||
**Validates: Requirements 2.1, 2.2, 2.3, 2.4**
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import date, timedelta
|
||||
|
||||
import pytest
|
||||
from hypothesis import given, settings, HealthCheck, strategies as st
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.sr_level import SRLevel
|
||||
from app.models.ticker import Ticker
|
||||
from app.services.rr_scanner_service import scan_ticker
|
||||
|
||||
|
||||
def _assert_primary_is_most_likely_worthwhile(setup) -> None:
|
||||
"""Headline = starred primary = max(probability, rr) among floor-clearing targets."""
|
||||
targets = setup.targets
|
||||
assert targets, "expected generated targets"
|
||||
primaries = [t for t in targets if t.get("is_primary")]
|
||||
assert len(primaries) == 1, "exactly one primary target expected"
|
||||
primary = primaries[0]
|
||||
assert setup.target == pytest.approx(primary["price"], abs=0.01)
|
||||
|
||||
# Mirrors recommendation_service._select_primary_target floors.
|
||||
worthwhile = [
|
||||
t for t in targets
|
||||
if float(t["rr_ratio"]) >= 1.5 and float(t["probability"]) >= 20.0
|
||||
]
|
||||
pool = worthwhile or targets
|
||||
best = max(pool, key=lambda t: (t["probability"], t["rr_ratio"]))
|
||||
assert primary["price"] == pytest.approx(best["price"], abs=0.01)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Session fixture (plain session, not wrapped in begin())
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.fixture
|
||||
async def scan_session() -> AsyncSession:
|
||||
"""Provide a DB session compatible with scan_ticker (which commits)."""
|
||||
from tests.conftest import _test_session_factory
|
||||
|
||||
async with _test_session_factory() as session:
|
||||
yield session
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _make_ohlcv_bars(
|
||||
ticker_id: int,
|
||||
num_bars: int = 20,
|
||||
base_close: float = 100.0,
|
||||
) -> list[OHLCVRecord]:
|
||||
"""Generate OHLCV bars closing around base_close with ATR ≈ 2.0."""
|
||||
bars: list[OHLCVRecord] = []
|
||||
start = date(2024, 1, 1)
|
||||
for i in range(num_bars):
|
||||
close = base_close + (i % 3 - 1) * 0.5 # oscillate ±0.5
|
||||
bars.append(OHLCVRecord(
|
||||
ticker_id=ticker_id,
|
||||
date=start + timedelta(days=i),
|
||||
open=close - 0.3,
|
||||
high=close + 1.0,
|
||||
low=close - 1.0,
|
||||
close=close,
|
||||
volume=100_000,
|
||||
))
|
||||
return bars
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Hypothesis strategy: multiple resistance levels above entry for longs
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@st.composite
|
||||
def long_candidate_levels(draw: st.DrawFn) -> list[dict]:
|
||||
"""Generate 2-5 resistance levels above entry_price=100.
|
||||
|
||||
All levels meet the R:R threshold of 1.5 given ATR≈2, risk≈3,
|
||||
so min reward=4.5, min target=104.5.
|
||||
"""
|
||||
num_levels = draw(st.integers(min_value=2, max_value=5))
|
||||
levels = []
|
||||
for _ in range(num_levels):
|
||||
# Distance from entry: 5 to 50 (all above 4.5 threshold)
|
||||
distance = draw(st.floats(min_value=5.0, max_value=50.0))
|
||||
strength = draw(st.integers(min_value=0, max_value=100))
|
||||
levels.append({
|
||||
"price": 100.0 + distance,
|
||||
"strength": strength,
|
||||
})
|
||||
return levels
|
||||
|
||||
|
||||
@st.composite
|
||||
def short_candidate_levels(draw: st.DrawFn) -> list[dict]:
|
||||
"""Generate 2-5 support levels below entry_price=100.
|
||||
|
||||
All levels meet the R:R threshold of 1.5 given ATR≈2, risk≈3,
|
||||
so min reward=4.5, max target=95.5.
|
||||
"""
|
||||
num_levels = draw(st.integers(min_value=2, max_value=5))
|
||||
levels = []
|
||||
for _ in range(num_levels):
|
||||
# Distance below entry: 5 to 50 (all above 4.5 threshold)
|
||||
distance = draw(st.floats(min_value=5.0, max_value=50.0))
|
||||
strength = draw(st.integers(min_value=0, max_value=100))
|
||||
levels.append({
|
||||
"price": 100.0 - distance,
|
||||
"strength": strength,
|
||||
})
|
||||
return levels
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Property test: long setup selects probability-based primary
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@given(levels=long_candidate_levels())
|
||||
@settings(
|
||||
max_examples=20,
|
||||
deadline=None,
|
||||
suppress_health_check=[HealthCheck.function_scoped_fixture],
|
||||
)
|
||||
async def test_property_long_selects_probability_primary(
|
||||
levels: list[dict],
|
||||
scan_session: AsyncSession,
|
||||
):
|
||||
"""**Validates: Requirements 2.1, 2.3, 2.4**
|
||||
|
||||
Property: when multiple resistance levels meet the R:R threshold,
|
||||
the headline after enhance is the probability-based primary.
|
||||
"""
|
||||
from tests.conftest import _test_engine, _test_session_factory
|
||||
from app.database import Base
|
||||
|
||||
# Fresh DB state per hypothesis example
|
||||
async with _test_engine.begin() as conn:
|
||||
await conn.run_sync(Base.metadata.drop_all)
|
||||
await conn.run_sync(Base.metadata.create_all)
|
||||
|
||||
async with _test_session_factory() as session:
|
||||
ticker = Ticker(symbol="FIXL")
|
||||
session.add(ticker)
|
||||
await session.flush()
|
||||
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
session.add_all(bars)
|
||||
|
||||
sr_levels = []
|
||||
for lv in levels:
|
||||
sr_levels.append(SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=lv["price"],
|
||||
type="resistance",
|
||||
strength=lv["strength"],
|
||||
detection_method="volume_profile",
|
||||
))
|
||||
session.add_all(sr_levels)
|
||||
await session.commit()
|
||||
|
||||
setups = await scan_ticker(
|
||||
session,
|
||||
"FIXL",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=sr_levels,
|
||||
)
|
||||
|
||||
long_setups = [s for s in setups if s.direction == "long"]
|
||||
assert len(long_setups) == 1, "Expected exactly one long setup"
|
||||
|
||||
_assert_primary_is_most_likely_worthwhile(long_setups[0])
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Property test: short setup selects probability-based primary
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@given(levels=short_candidate_levels())
|
||||
@settings(
|
||||
max_examples=20,
|
||||
deadline=None,
|
||||
suppress_health_check=[HealthCheck.function_scoped_fixture],
|
||||
)
|
||||
async def test_property_short_selects_probability_primary(
|
||||
levels: list[dict],
|
||||
scan_session: AsyncSession,
|
||||
):
|
||||
"""**Validates: Requirements 2.2, 2.3, 2.4**
|
||||
|
||||
Property: when multiple support levels meet the R:R threshold,
|
||||
the headline after enhance is the probability-based primary.
|
||||
"""
|
||||
from tests.conftest import _test_engine, _test_session_factory
|
||||
from app.database import Base
|
||||
|
||||
# Fresh DB state per hypothesis example
|
||||
async with _test_engine.begin() as conn:
|
||||
await conn.run_sync(Base.metadata.drop_all)
|
||||
await conn.run_sync(Base.metadata.create_all)
|
||||
|
||||
async with _test_session_factory() as session:
|
||||
ticker = Ticker(symbol="FIXS")
|
||||
session.add(ticker)
|
||||
await session.flush()
|
||||
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
session.add_all(bars)
|
||||
|
||||
sr_levels = []
|
||||
for lv in levels:
|
||||
sr_levels.append(SRLevel(
|
||||
ticker_id=ticker.id,
|
||||
price_level=lv["price"],
|
||||
type="support",
|
||||
strength=lv["strength"],
|
||||
detection_method="pivot_point",
|
||||
))
|
||||
session.add_all(sr_levels)
|
||||
await session.commit()
|
||||
|
||||
setups = await scan_ticker(
|
||||
session,
|
||||
"FIXS",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=sr_levels,
|
||||
)
|
||||
|
||||
short_setups = [s for s in setups if s.direction == "short"]
|
||||
assert len(short_setups) == 1, "Expected exactly one short setup"
|
||||
|
||||
_assert_primary_is_most_likely_worthwhile(short_setups[0])
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Deterministic test: 3 levels with known quality scores (long)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_deterministic_long_three_levels(scan_session: AsyncSession):
|
||||
"""**Validates: Requirements 2.1, 2.3, 2.4**
|
||||
|
||||
Concrete example with 3 resistance levels of known quality scores.
|
||||
Entry=100, ATR≈2, risk≈3.
|
||||
|
||||
Level A: price=105, strength=90 → rr=5/3≈1.67, dist=5
|
||||
quality = 0.35*(1.67/10) + 0.35*(90/100) + 0.30*(1-5/100)
|
||||
= 0.35*0.167 + 0.35*0.9 + 0.30*0.95
|
||||
= 0.0585 + 0.315 + 0.285 = 0.6585
|
||||
|
||||
Level B: price=112, strength=50 → rr=12/3=4.0, dist=12
|
||||
quality = 0.35*(4/10) + 0.35*(50/100) + 0.30*(1-12/100)
|
||||
= 0.35*0.4 + 0.35*0.5 + 0.30*0.88
|
||||
= 0.14 + 0.175 + 0.264 = 0.579
|
||||
|
||||
Level C: price=130, strength=10 → rr=30/3=10.0, dist=30
|
||||
quality = 0.35*(10/10) + 0.35*(10/100) + 0.30*(1-30/100)
|
||||
= 0.35*1.0 + 0.35*0.1 + 0.30*0.7
|
||||
= 0.35 + 0.035 + 0.21 = 0.595
|
||||
|
||||
Expected winner: Level A (quality=0.6585)
|
||||
"""
|
||||
ticker = Ticker(symbol="DET3L")
|
||||
scan_session.add(ticker)
|
||||
await scan_session.flush()
|
||||
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
scan_session.add_all(bars)
|
||||
|
||||
level_a = SRLevel(
|
||||
ticker_id=ticker.id, price_level=105.0, type="resistance",
|
||||
strength=90, detection_method="volume_profile",
|
||||
)
|
||||
level_b = SRLevel(
|
||||
ticker_id=ticker.id, price_level=112.0, type="resistance",
|
||||
strength=50, detection_method="volume_profile",
|
||||
)
|
||||
level_c = SRLevel(
|
||||
ticker_id=ticker.id, price_level=130.0, type="resistance",
|
||||
strength=10, detection_method="volume_profile",
|
||||
)
|
||||
scan_session.add_all([level_a, level_b, level_c])
|
||||
await scan_session.flush()
|
||||
|
||||
setups = await scan_ticker(
|
||||
scan_session,
|
||||
"DET3L",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=[level_a, level_b, level_c],
|
||||
)
|
||||
|
||||
long_setups = [s for s in setups if s.direction == "long"]
|
||||
assert len(long_setups) == 1, "Expected exactly one long setup"
|
||||
|
||||
_assert_primary_is_most_likely_worthwhile(long_setups[0])
|
||||
# Near/strong level A wins on reach-probability over far lottery C.
|
||||
assert long_setups[0].target == pytest.approx(105.0, abs=0.01), (
|
||||
f"Expected primary=105.0 (near, high reach-prob), got {long_setups[0].target}"
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Deterministic test: 3 levels with known quality scores (short)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_deterministic_short_three_levels(scan_session: AsyncSession):
|
||||
"""**Validates: Requirements 2.2, 2.3, 2.4**
|
||||
|
||||
Concrete example with 3 support levels of known quality scores.
|
||||
Entry=100, ATR≈2, risk≈3.
|
||||
|
||||
Level A: price=95, strength=85 → rr=5/3≈1.67, dist=5
|
||||
quality = 0.35*(1.67/10) + 0.35*(85/100) + 0.30*(1-5/100)
|
||||
= 0.0585 + 0.2975 + 0.285 = 0.641
|
||||
|
||||
Level B: price=88, strength=45 → rr=12/3=4.0, dist=12
|
||||
quality = 0.35*(4/10) + 0.35*(45/100) + 0.30*(1-12/100)
|
||||
= 0.14 + 0.1575 + 0.264 = 0.5615
|
||||
|
||||
Level C: price=70, strength=8 → rr=30/3=10.0, dist=30
|
||||
quality = 0.35*(10/10) + 0.35*(8/100) + 0.30*(1-30/100)
|
||||
= 0.35 + 0.028 + 0.21 = 0.588
|
||||
|
||||
Expected winner: Level A (quality=0.641)
|
||||
"""
|
||||
ticker = Ticker(symbol="DET3S")
|
||||
scan_session.add(ticker)
|
||||
await scan_session.flush()
|
||||
|
||||
bars = _make_ohlcv_bars(ticker.id, num_bars=20, base_close=100.0)
|
||||
scan_session.add_all(bars)
|
||||
|
||||
level_a = SRLevel(
|
||||
ticker_id=ticker.id, price_level=95.0, type="support",
|
||||
strength=85, detection_method="pivot_point",
|
||||
)
|
||||
level_b = SRLevel(
|
||||
ticker_id=ticker.id, price_level=88.0, type="support",
|
||||
strength=45, detection_method="pivot_point",
|
||||
)
|
||||
level_c = SRLevel(
|
||||
ticker_id=ticker.id, price_level=70.0, type="support",
|
||||
strength=8, detection_method="pivot_point",
|
||||
)
|
||||
scan_session.add_all([level_a, level_b, level_c])
|
||||
await scan_session.flush()
|
||||
|
||||
setups = await scan_ticker(
|
||||
scan_session,
|
||||
"DET3S",
|
||||
rr_threshold=1.5,
|
||||
gate_levels_override=[level_a, level_b, level_c],
|
||||
)
|
||||
|
||||
short_setups = [s for s in setups if s.direction == "short"]
|
||||
assert len(short_setups) == 1, "Expected exactly one short setup"
|
||||
|
||||
_assert_primary_is_most_likely_worthwhile(short_setups[0])
|
||||
assert short_setups[0].target == pytest.approx(95.0, abs=0.01), (
|
||||
f"Expected primary=95.0 (near, high reach-prob), got {short_setups[0].target}"
|
||||
)
|
||||
@@ -23,6 +23,8 @@ from sqlalchemy.ext.asyncio import AsyncSession
|
||||
from app.models.ohlcv import OHLCVRecord
|
||||
from app.models.paper_trade import PaperTrade
|
||||
from app.models.signal_context_snapshot import SignalContextSnapshot
|
||||
from app.models.sec_filing_gap import SecFilingGap
|
||||
from app.models.settings import SystemSetting
|
||||
from app.models.sr_level import SRLevel
|
||||
from app.models.ticker import Ticker
|
||||
from app.models.trade_setup import TradeSetup
|
||||
@@ -513,6 +515,45 @@ async def test_get_trade_setups_excludes_stale_rows(db_session: AsyncSession):
|
||||
assert stale_rows == []
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_get_trade_setups_hides_active_sec_filing_gap(
|
||||
db_session: AsyncSession,
|
||||
):
|
||||
now = datetime.now(timezone.utc)
|
||||
ticker = Ticker(symbol="SECWAIT", cik="0000000042")
|
||||
db_session.add(ticker)
|
||||
await db_session.flush()
|
||||
db_session.add_all([
|
||||
SystemSetting(
|
||||
key="fundamental_data_sec_dolt_cutover_enabled",
|
||||
value="true",
|
||||
),
|
||||
SecFilingGap(
|
||||
cik=ticker.cik,
|
||||
accession="0000000042-26-000001",
|
||||
form="10-Q",
|
||||
index_date=date.today(),
|
||||
reason="not_in_companyfacts",
|
||||
first_seen_at=now,
|
||||
last_attempted_at=now,
|
||||
),
|
||||
TradeSetup(
|
||||
ticker_id=ticker.id,
|
||||
direction="long",
|
||||
entry_price=100.0,
|
||||
stop_loss=97.0,
|
||||
target=109.0,
|
||||
rr_ratio=3.0,
|
||||
composite_score=70.0,
|
||||
confidence_score=80.0,
|
||||
detected_at=now,
|
||||
),
|
||||
])
|
||||
await db_session.flush()
|
||||
|
||||
assert await get_trade_setups(db_session, symbol="SECWAIT") == []
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_get_trade_setups_can_exclude_tickers_with_open_paper_trades(
|
||||
db_session: AsyncSession,
|
||||
|
||||
@@ -108,3 +108,61 @@ async def test_scan_error_does_not_stop_later_tickers(session, monkeypatch):
|
||||
await rr_scanner_service.scan_all_tickers(session)
|
||||
|
||||
assert scanned == ["AAA", "BBB"]
|
||||
|
||||
|
||||
async def test_scan_skips_ticker_with_incomplete_sec_fundamentals(
|
||||
session, monkeypatch
|
||||
):
|
||||
ticker = Ticker(symbol="BLOCKED", cik="0000000001")
|
||||
session.add(ticker)
|
||||
await session.commit()
|
||||
|
||||
async def _blocked(db):
|
||||
return {ticker.id}
|
||||
|
||||
async def _unexpected_scan(*args, **kwargs):
|
||||
raise AssertionError("fundamentals-incomplete ticker was scanned")
|
||||
|
||||
monkeypatch.setattr(
|
||||
rr_scanner_service.fundamentals_quality_service,
|
||||
"blocked_ticker_ids",
|
||||
_blocked,
|
||||
)
|
||||
monkeypatch.setattr(rr_scanner_service, "scan_ticker", _unexpected_scan)
|
||||
|
||||
assert await rr_scanner_service.scan_all_tickers(session) == []
|
||||
|
||||
|
||||
async def test_scan_quality_failure_blocks_closed_and_emits_event(
|
||||
session, monkeypatch
|
||||
):
|
||||
session.add(Ticker(symbol="BLOCKED"))
|
||||
await session.commit()
|
||||
|
||||
async def _boom(db):
|
||||
raise ValueError("bad quality metadata")
|
||||
|
||||
async def _unexpected_scan(*args, **kwargs):
|
||||
raise AssertionError("ticker was scanned without a quality decision")
|
||||
|
||||
events: list[dict] = []
|
||||
|
||||
async def _capture_event(**kwargs):
|
||||
events.append(kwargs)
|
||||
|
||||
monkeypatch.setattr(
|
||||
rr_scanner_service.fundamentals_quality_service,
|
||||
"blocked_ticker_ids",
|
||||
_boom,
|
||||
)
|
||||
monkeypatch.setattr(rr_scanner_service, "scan_ticker", _unexpected_scan)
|
||||
monkeypatch.setattr(
|
||||
rr_scanner_service.system_event_service,
|
||||
"log_event_standalone",
|
||||
_capture_event,
|
||||
)
|
||||
|
||||
assert await rr_scanner_service.scan_all_tickers(session) == []
|
||||
assert [event["code"] for event in events] == [
|
||||
"fundamentals_quality_unavailable"
|
||||
]
|
||||
|
||||
@@ -85,6 +85,7 @@ class TestTradingDayCrons:
|
||||
(
|
||||
("schedule_dolt_earnings_cron", 2, 30),
|
||||
("schedule_sec_fundamentals_cron", 4, 0),
|
||||
("schedule_fundamentals_parity_cron", 5, 30),
|
||||
),
|
||||
)
|
||||
def test_shadow_imports_run_daily_at_expected_et_time(
|
||||
|
||||
@@ -5,18 +5,24 @@ from types import SimpleNamespace
|
||||
import pytest
|
||||
|
||||
from app.scheduler import (
|
||||
_DAILY_PIPELINE_STEPS,
|
||||
_NEAR_CLOSE_PIPELINE_STEPS,
|
||||
_consume_backtest_options,
|
||||
_consume_backtest_target_model,
|
||||
_parse_frequency,
|
||||
_resume_tickers,
|
||||
_last_successful,
|
||||
_run_shadow_import,
|
||||
collect_fundamentals,
|
||||
run_fundamentals_parity_report,
|
||||
run_sec_fundamentals_import,
|
||||
configure_scheduler,
|
||||
get_job_runtime_snapshot,
|
||||
queue_backtest_options,
|
||||
queue_backtest_target_model,
|
||||
scheduler,
|
||||
)
|
||||
from app.services.data_import import STATUS_DEFERRED
|
||||
|
||||
|
||||
def test_manual_backtest_target_model_is_one_shot():
|
||||
@@ -39,6 +45,14 @@ def test_manual_backtest_options_are_one_shot_and_default_back_to_weekly():
|
||||
assert _consume_backtest_options() == ("production_gtl", "weekly")
|
||||
|
||||
|
||||
def test_only_near_close_fetch_skips_redundant_sr_refresh():
|
||||
assert dict(_DAILY_PIPELINE_STEPS)["data_collector"] == "collect_ohlcv"
|
||||
assert (
|
||||
dict(_NEAR_CLOSE_PIPELINE_STEPS)["data_collector"]
|
||||
== "collect_ohlcv_for_scan"
|
||||
)
|
||||
|
||||
|
||||
class TestParseFrequency:
|
||||
def test_hourly(self):
|
||||
assert _parse_frequency("hourly") == {"hours": 1}
|
||||
@@ -112,6 +126,7 @@ class TestConfigureScheduler:
|
||||
"fundamental_collector",
|
||||
"dolt_earnings_import",
|
||||
"sec_fundamentals_import",
|
||||
"fundamentals_parity_report",
|
||||
"rr_scanner",
|
||||
"shadow_book",
|
||||
"ticker_universe_sync",
|
||||
@@ -145,6 +160,7 @@ class TestConfigureScheduler:
|
||||
"fundamental_collector",
|
||||
"dolt_earnings_import",
|
||||
"sec_fundamentals_import",
|
||||
"fundamentals_parity_report",
|
||||
"market_regime",
|
||||
"near_close_pipeline",
|
||||
"regime_monitor",
|
||||
@@ -165,6 +181,41 @@ class _SessionContext:
|
||||
return None
|
||||
|
||||
|
||||
class TestFundamentalCollector:
|
||||
@staticmethod
|
||||
def _session_factory():
|
||||
return _SessionContext()
|
||||
|
||||
async def test_skips_legacy_provider_when_cutover_is_active(self, monkeypatch):
|
||||
async def enabled(db, job_name):
|
||||
return True
|
||||
|
||||
async def cutover_enabled(db):
|
||||
return True
|
||||
|
||||
async def unexpected_ticker_lookup(db):
|
||||
raise AssertionError("legacy ticker lookup must not run after cutover")
|
||||
|
||||
monkeypatch.setattr("app.scheduler.async_session_factory", self._session_factory)
|
||||
monkeypatch.setattr("app.scheduler._is_job_enabled", enabled)
|
||||
monkeypatch.setattr(
|
||||
"app.scheduler.fundamental_data_refresh_service.is_enabled",
|
||||
cutover_enabled,
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
"app.scheduler._get_fundamental_priority_tickers",
|
||||
unexpected_ticker_lookup,
|
||||
)
|
||||
|
||||
await collect_fundamentals()
|
||||
|
||||
runtime = get_job_runtime_snapshot("fundamental_collector")
|
||||
assert runtime["status"] == "skipped"
|
||||
assert runtime["processed"] == 0
|
||||
assert runtime["total"] == 0
|
||||
assert runtime["message"] == "SEC + Dolt fundamentals cutover is active"
|
||||
|
||||
|
||||
class TestShadowImportJobs:
|
||||
@staticmethod
|
||||
def _session_factory():
|
||||
@@ -210,6 +261,28 @@ class TestShadowImportJobs:
|
||||
assert runtime["processed"] == 0
|
||||
assert runtime["message"] == "validation failed"
|
||||
|
||||
async def test_deferred_run_is_visible_without_error_status(self, monkeypatch):
|
||||
async def enabled(db, job_name):
|
||||
return True
|
||||
|
||||
async def imported(importer):
|
||||
return SimpleNamespace(
|
||||
status=STATUS_DEFERRED,
|
||||
revision="abcdef1234567890",
|
||||
error_details="Company Facts publication lag; retrying",
|
||||
)
|
||||
|
||||
monkeypatch.setattr("app.scheduler.async_session_factory", self._session_factory)
|
||||
monkeypatch.setattr("app.scheduler._is_job_enabled", enabled)
|
||||
monkeypatch.setattr("app.scheduler.run_import", imported)
|
||||
|
||||
await _run_shadow_import("sec_fundamentals_import", object())
|
||||
|
||||
runtime = get_job_runtime_snapshot("sec_fundamentals_import")
|
||||
assert runtime["status"] == STATUS_DEFERRED
|
||||
assert runtime["processed"] == 0
|
||||
assert runtime["message"] == "Company Facts publication lag; retrying"
|
||||
|
||||
async def test_source_lock_surfaces_skipped(self, monkeypatch):
|
||||
async def enabled(db, job_name):
|
||||
return True
|
||||
@@ -243,3 +316,121 @@ class TestShadowImportJobs:
|
||||
runtime = get_job_runtime_snapshot("sec_fundamentals_import")
|
||||
assert runtime["status"] == "skipped"
|
||||
assert runtime["message"] == "Disabled"
|
||||
|
||||
async def test_sec_failure_still_runs_activated_local_refresh(self, monkeypatch):
|
||||
calls = []
|
||||
|
||||
async def enabled(db, job_name):
|
||||
return True
|
||||
|
||||
async def unavailable(importer):
|
||||
raise RuntimeError("SEC unavailable")
|
||||
|
||||
async def refreshed(db):
|
||||
calls.append(db)
|
||||
return {
|
||||
"enabled": True,
|
||||
"refreshed": 511,
|
||||
"score_inputs_changed": 2,
|
||||
"dimension_scores_staled": 2,
|
||||
"composite_scores_staled": 2,
|
||||
}
|
||||
|
||||
monkeypatch.setattr("app.scheduler.async_session_factory", self._session_factory)
|
||||
monkeypatch.setattr("app.scheduler._is_job_enabled", enabled)
|
||||
monkeypatch.setattr("app.scheduler.run_import", unavailable)
|
||||
monkeypatch.setattr(
|
||||
"app.scheduler.fundamental_data_refresh_service.refresh_if_enabled",
|
||||
refreshed,
|
||||
)
|
||||
|
||||
await run_sec_fundamentals_import()
|
||||
|
||||
assert len(calls) == 1
|
||||
runtime = get_job_runtime_snapshot("sec_fundamentals_import")
|
||||
assert runtime["status"] == "error"
|
||||
assert runtime["message"] == "SEC unavailable"
|
||||
|
||||
async def test_sec_success_surfaces_activated_refresh_summary(self, monkeypatch):
|
||||
async def enabled(db, job_name):
|
||||
return True
|
||||
|
||||
async def imported(importer):
|
||||
return SimpleNamespace(
|
||||
status="no_op", revision="abcdef1234567890", error_details=None
|
||||
)
|
||||
|
||||
async def refreshed(db):
|
||||
return {
|
||||
"enabled": True,
|
||||
"refreshed": 511,
|
||||
"score_inputs_changed": 2,
|
||||
"dimension_scores_staled": 2,
|
||||
"composite_scores_staled": 2,
|
||||
}
|
||||
|
||||
monkeypatch.setattr("app.scheduler.async_session_factory", self._session_factory)
|
||||
monkeypatch.setattr("app.scheduler._is_job_enabled", enabled)
|
||||
monkeypatch.setattr("app.scheduler.run_import", imported)
|
||||
monkeypatch.setattr(
|
||||
"app.scheduler.fundamental_data_refresh_service.refresh_if_enabled",
|
||||
refreshed,
|
||||
)
|
||||
|
||||
await run_sec_fundamentals_import()
|
||||
|
||||
runtime = get_job_runtime_snapshot("sec_fundamentals_import")
|
||||
assert runtime["status"] == "completed"
|
||||
assert runtime["message"] == (
|
||||
"no_op · abcdef123456 · cache 511 · 2 score inputs changed"
|
||||
)
|
||||
|
||||
async def test_disabled_sec_job_does_not_run_local_refresh(self, monkeypatch):
|
||||
async def disabled(db, job_name):
|
||||
return False
|
||||
|
||||
async def should_not_run(*args, **kwargs):
|
||||
raise AssertionError("disabled SEC job ran work")
|
||||
|
||||
monkeypatch.setattr("app.scheduler.async_session_factory", self._session_factory)
|
||||
monkeypatch.setattr("app.scheduler._is_job_enabled", disabled)
|
||||
monkeypatch.setattr("app.scheduler.run_import", should_not_run)
|
||||
monkeypatch.setattr(
|
||||
"app.scheduler.fundamental_data_refresh_service.refresh_if_enabled",
|
||||
should_not_run,
|
||||
)
|
||||
|
||||
await run_sec_fundamentals_import()
|
||||
|
||||
runtime = get_job_runtime_snapshot("sec_fundamentals_import")
|
||||
assert runtime["status"] == "skipped"
|
||||
assert runtime["message"] == "Disabled"
|
||||
|
||||
|
||||
async def test_fundamentals_parity_job_surfaces_report_summary(monkeypatch):
|
||||
async def enabled(db, job_name):
|
||||
return True
|
||||
|
||||
async def generated(db, report_dir):
|
||||
return (
|
||||
{
|
||||
"generated_at": "2026-07-23T10:30:00+00:00",
|
||||
"summary": {
|
||||
"universe_count": 511,
|
||||
"fundamental_score_material_changes": 12,
|
||||
},
|
||||
},
|
||||
{"json": "report.json", "csv": "report.csv"},
|
||||
)
|
||||
|
||||
monkeypatch.setattr("app.scheduler.async_session_factory", TestShadowImportJobs._session_factory)
|
||||
monkeypatch.setattr("app.scheduler._is_job_enabled", enabled)
|
||||
monkeypatch.setattr(
|
||||
"app.scheduler.fundamentals_parity_service.generate_and_store", generated
|
||||
)
|
||||
|
||||
await run_fundamentals_parity_report()
|
||||
|
||||
runtime = get_job_runtime_snapshot("fundamentals_parity_report")
|
||||
assert runtime["status"] == "completed"
|
||||
assert runtime["message"] == "511 tickers · 12 material score changes"
|
||||
|
||||
@@ -207,6 +207,69 @@ async def test_only_404_is_treated_as_missing():
|
||||
assert await c.latest_index_date(today=date(2026, 7, 22)) is None
|
||||
|
||||
|
||||
# The two shapes a real SEC 403 takes (captured live 2026-07-30). /Archives is
|
||||
# S3-backed with no ListBucket grant, so an ABSENT file comes back as S3's
|
||||
# AccessDenied XML; a genuine fair-access rejection is the WAF interstitial.
|
||||
S3_ACCESS_DENIED = (
|
||||
'<?xml version="1.0" encoding="UTF-8"?>'
|
||||
"<Error><Code>AccessDenied</Code><Message>Access Denied</Message>"
|
||||
"<RequestId>5AWQBRAEX3NPAPHB</RequestId><HostId>MHFbU0a3k0ER</HostId></Error>"
|
||||
)
|
||||
WAF_HTML = (
|
||||
"<!DOCTYPE html><html><head><title>SEC.gov | Your Request Originates from "
|
||||
"an Undeclared Automated Tool</title></head><body>...</body></html>"
|
||||
)
|
||||
|
||||
|
||||
def _forbidden(body: str, content_type: str) -> httpx.Response:
|
||||
return httpx.Response(
|
||||
403, content=body.encode(), headers={"Content-Type": content_type}
|
||||
)
|
||||
|
||||
|
||||
async def test_archives_access_denied_is_absent_not_forbidden():
|
||||
# SEC publishes no daily index on weekends, and the bucket reports the absent
|
||||
# key as 403/AccessDenied. Treating that as fatal wedged the importer on the
|
||||
# first Saturday of an incremental walk (2026-07-25); it must read as "missing".
|
||||
def handler(request: httpx.Request) -> httpx.Response:
|
||||
url = str(request.url)
|
||||
if url.endswith("QTR2/index.json"):
|
||||
return httpx.Response(200, json={"directory": {"item": [{"name": "form.20260630.idx"}]}})
|
||||
return _forbidden(S3_ACCESS_DENIED, "application/xml")
|
||||
|
||||
def client() -> SecClient:
|
||||
return SecClient(
|
||||
transport=httpx.MockTransport(handler), spacing_seconds=0, max_retries=0
|
||||
)
|
||||
|
||||
async with client() as c:
|
||||
assert await c.daily_index(date(2026, 7, 25)) == []
|
||||
async with client() as c:
|
||||
# QTR3 absent → the previous-quarter fallback now actually fires.
|
||||
assert await c.latest_index_date(today=date(2026, 7, 22)) == date(2026, 6, 30)
|
||||
|
||||
|
||||
async def test_archives_waf_rejection_stays_forbidden():
|
||||
# A real UA/pattern rejection is served for files that DO exist — never
|
||||
# downgrade it, or a blocked run would look like an empty index.
|
||||
def handler(request: httpx.Request) -> httpx.Response:
|
||||
return _forbidden(WAF_HTML, "text/html")
|
||||
|
||||
async with SecClient(transport=httpx.MockTransport(handler), spacing_seconds=0) as c:
|
||||
with pytest.raises(SecForbiddenError):
|
||||
await c.daily_index(date(2026, 7, 21))
|
||||
|
||||
|
||||
async def test_access_denied_outside_archives_stays_forbidden():
|
||||
# The downgrade is gated on the Archives prefix; data.sec.gov is not S3-backed.
|
||||
def handler(request: httpx.Request) -> httpx.Response:
|
||||
return _forbidden(S3_ACCESS_DENIED, "application/xml")
|
||||
|
||||
async with SecClient(transport=httpx.MockTransport(handler), spacing_seconds=0) as c:
|
||||
with pytest.raises(SecForbiddenError):
|
||||
await c.companyfacts(320193)
|
||||
|
||||
|
||||
async def test_fair_access_validation_on_real_client():
|
||||
# Placeholder email rejected.
|
||||
with pytest.raises(SecError):
|
||||
|
||||
@@ -270,3 +270,271 @@ async def test_live_apple_parse_invariants():
|
||||
# shares cover-date differs from period_end
|
||||
latest = max(rows, key=lambda r: r.period_end)
|
||||
assert latest.shares_outstanding_date != latest.period_end
|
||||
|
||||
|
||||
# -- revenue concept coverage (A5 parity findings) ---------------------------
|
||||
|
||||
def _one_filing(concepts: dict, *, start: str, end: str, fp: str):
|
||||
"""A single 10-Q whose facts are the given {concept: value} at one YTD span."""
|
||||
facts = {
|
||||
name: {"units": {"USD": [_dur(start, end, val, "X", fp=fp)]}}
|
||||
for name, val in concepts.items()
|
||||
}
|
||||
companyfacts = {"cik": 19617, "facts": {"us-gaap": facts}}
|
||||
filings = {
|
||||
"X": FilingMeta(
|
||||
date.fromisoformat(end), date(2026, 5, 1), datetime(2026, 5, 1, 10, tzinfo=UTC), "10-Q"
|
||||
)
|
||||
}
|
||||
return parse_snapshots(companyfacts, filings, {"X"})
|
||||
|
||||
|
||||
def test_revenue_reads_banks_total_revenue_tag():
|
||||
# JPM/GS/WFC tag RevenuesNetOfInterestExpense in every 10-Q and never (or
|
||||
# only annually) `Revenues` -- previously null, so revenue growth was too.
|
||||
res = _one_filing(
|
||||
{"RevenuesNetOfInterestExpense": 49836}, start="2026-01-01", end="2026-03-31", fp="Q1"
|
||||
)
|
||||
assert res.rows[0].revenue == 49836
|
||||
|
||||
|
||||
def test_revenue_reads_including_assessed_tax_variant():
|
||||
# ARE/KHC tag only the Including variant.
|
||||
res = _one_filing(
|
||||
{"RevenueFromContractWithCustomerIncludingAssessedTax": 671},
|
||||
start="2026-01-01", end="2026-03-31", fp="Q1",
|
||||
)
|
||||
assert res.rows[0].revenue == 671
|
||||
|
||||
|
||||
def test_revenue_concept_priority_is_unchanged_by_the_added_tags():
|
||||
# The new entries are appended, so any issuer that already resolved keeps
|
||||
# the same concept -- only issuers that resolved to nothing gain a value.
|
||||
res = _one_filing(
|
||||
{
|
||||
"RevenueFromContractWithCustomerExcludingAssessedTax": 100,
|
||||
"RevenueFromContractWithCustomerIncludingAssessedTax": 110,
|
||||
"RevenuesNetOfInterestExpense": 120,
|
||||
"Revenues": 130,
|
||||
},
|
||||
start="2026-01-01", end="2026-03-31", fp="Q1",
|
||||
)
|
||||
assert res.rows[0].revenue == 100
|
||||
|
||||
|
||||
def test_four_four_five_q3_ytd_span_is_accepted():
|
||||
# A 12/12/12/16-week filer's YTD-Q3 is 36 weeks = 251 days (COST 2026 Q3),
|
||||
# which missed the old 20-day tolerance around 273 by ~2 and dropped Q3
|
||||
# every year -- breaking the quarter chain and nulling TTM and YoY.
|
||||
res = _one_filing(
|
||||
{"RevenueFromContractWithCustomerExcludingAssessedTax": 207431},
|
||||
start="2025-09-01", end="2026-05-10", fp="Q3",
|
||||
)
|
||||
assert (date(2026, 5, 10) - date(2025, 9, 1)).days == 251
|
||||
assert res.rows[0].revenue == 207431
|
||||
|
||||
|
||||
def test_eps_falls_back_to_continuing_operations_variant():
|
||||
# REG tags only this variant on every filing; FCX tags it in its 10-K while
|
||||
# using EarningsPerShareDiluted in its 10-Qs.
|
||||
companyfacts = {
|
||||
"cik": 910606,
|
||||
"facts": {"us-gaap": {"IncomeLossFromContinuingOperationsPerDilutedShare": {
|
||||
"units": {"USD/shares": [_dur("2026-01-01", "2026-03-31", 1.81, "X", fp="Q1")]}
|
||||
}}},
|
||||
}
|
||||
filings = {"X": FilingMeta(date(2026, 3, 31), date(2026, 5, 1),
|
||||
datetime(2026, 5, 1, 10, tzinfo=UTC), "10-Q")}
|
||||
res = parse_snapshots(companyfacts, filings, {"X"})
|
||||
assert res.rows[0].diluted_eps == 1.81
|
||||
|
||||
|
||||
def test_eps_concept_priority_is_unchanged_by_the_added_tag():
|
||||
companyfacts = {
|
||||
"cik": 831259,
|
||||
"facts": {"us-gaap": {
|
||||
"EarningsPerShareDiluted": {
|
||||
"units": {"USD/shares": [_dur("2026-01-01", "2026-03-31", 0.61, "X", fp="Q1")]}},
|
||||
"IncomeLossFromContinuingOperationsPerDilutedShare": {
|
||||
"units": {"USD/shares": [_dur("2026-01-01", "2026-03-31", 0.75, "X", fp="Q1")]}},
|
||||
}},
|
||||
}
|
||||
filings = {"X": FilingMeta(date(2026, 3, 31), date(2026, 5, 1),
|
||||
datetime(2026, 5, 1, 10, tzinfo=UTC), "10-Q")}
|
||||
res = parse_snapshots(companyfacts, filings, {"X"})
|
||||
assert res.rows[0].diluted_eps == 0.61
|
||||
|
||||
|
||||
# -- period identity from the fiscal calendar, not SEC's fy/fp ---------------
|
||||
|
||||
from app.services.sec_facts_parser import _period_identity # noqa: E402
|
||||
|
||||
|
||||
def _meta(end: str, form: str = "10-Q") -> FilingMeta:
|
||||
d = date.fromisoformat(end)
|
||||
return FilingMeta(d, d, datetime(d.year, d.month, d.day, tzinfo=UTC), form)
|
||||
|
||||
|
||||
def test_a_10q_is_never_labelled_fy():
|
||||
# BXP: a 10-Q for period end 2026-03-31 carried fy/fp saying "2026 FY", which
|
||||
# collided with the real annual row and measured a 90-day fact against the
|
||||
# 365-day FY expectation.
|
||||
fy, fp = _period_identity(_meta("2026-03-31"), "1231")
|
||||
assert (fy, fp) == (2026, "Q1")
|
||||
|
||||
|
||||
def test_december_filer_years_do_not_collide():
|
||||
# FRT: two 10-Ks, ending 2024-12-31 and 2025-12-31, both labelled "2024 FY".
|
||||
assert _period_identity(_meta("2024-12-31", "10-K"), "1231") == (2024, "FY")
|
||||
assert _period_identity(_meta("2025-12-31", "10-K"), "1231") == (2025, "FY")
|
||||
|
||||
|
||||
def test_january_year_end_groups_its_quarters():
|
||||
# CRM/CRWD/WDAY: the year ending 2026-01-31 and its own quarters must share a
|
||||
# fiscal year, and must not collide with the year ending 2025-01-31.
|
||||
assert _period_identity(_meta("2026-01-31", "10-K"), "0131") == (2026, "FY")
|
||||
assert _period_identity(_meta("2025-01-31", "10-K"), "0131") == (2025, "FY")
|
||||
assert _period_identity(_meta("2025-04-30"), "0131") == (2026, "Q1")
|
||||
assert _period_identity(_meta("2025-07-31"), "0131") == (2026, "Q2")
|
||||
assert _period_identity(_meta("2025-10-31"), "0131") == (2026, "Q3")
|
||||
|
||||
|
||||
def test_mid_year_end_orders_correctly():
|
||||
# STX: the year ending 2025-06-27 was labelled "2027 FY" and sorted after
|
||||
# quarters that precede it.
|
||||
assert _period_identity(_meta("2025-06-27", "10-K"), "0627") == (2025, "FY")
|
||||
assert _period_identity(_meta("2025-10-03"), "0627") == (2026, "Q1")
|
||||
assert _period_identity(_meta("2026-01-02"), "0627") == (2026, "Q2")
|
||||
assert _period_identity(_meta("2026-04-03"), "0627") == (2026, "Q3")
|
||||
|
||||
|
||||
def test_four_four_five_quarters_place_correctly():
|
||||
# COST: a 12/12/12/16-week year leaves Q3 112 days from the year end, not 91.
|
||||
assert _period_identity(_meta("2025-11-23"), "0830") == (2026, "Q1")
|
||||
assert _period_identity(_meta("2026-02-15"), "0830") == (2026, "Q2")
|
||||
assert _period_identity(_meta("2026-05-10"), "0830") == (2026, "Q3")
|
||||
assert _period_identity(_meta("2026-08-30", "10-K"), "0830") == (2026, "FY")
|
||||
|
||||
|
||||
def test_year_end_crossing_january_still_groups_one_year():
|
||||
# DPZ (fiscalYearEnd 0102): the label shifts by one against Domino's own
|
||||
# naming, which is fine -- a year and its quarters must simply agree.
|
||||
year, _ = _period_identity(_meta("2025-12-28", "10-K"), "0102")
|
||||
assert (year, "FY") == _period_identity(_meta("2025-12-28", "10-K"), "0102")
|
||||
assert _period_identity(_meta("2025-03-23"), "0102") == (year, "Q1")
|
||||
assert _period_identity(_meta("2025-06-15"), "0102") == (year, "Q2")
|
||||
assert _period_identity(_meta("2025-09-07"), "0102") == (year, "Q3")
|
||||
|
||||
|
||||
def test_missing_fiscal_calendar_falls_back_to_filing_context():
|
||||
assert _period_identity(_meta("2026-03-31"), None) == (None, None)
|
||||
# ...and parse_snapshots then uses the fy/fp path, preserving old behaviour.
|
||||
res = parse_snapshots(COMPANYFACTS, FILINGS, {"B"})
|
||||
assert (res.rows[0].fiscal_year, res.rows[0].fiscal_period) == (2026, "Q2")
|
||||
|
||||
|
||||
|
||||
def test_eps_falls_back_to_basic_only_when_no_diluted_variant_exists():
|
||||
# PPL's 2026 Q1 tags no diluted EPS at all, only basic -- one missing period
|
||||
# broke the quarter chain and nulled TTM.
|
||||
companyfacts = {
|
||||
"cik": 922224,
|
||||
"facts": {"us-gaap": {"EarningsPerShareBasic": {
|
||||
"units": {"USD/shares": [_dur("2026-01-01", "2026-03-31", 0.60, "X", fp="Q1")]}
|
||||
}}},
|
||||
}
|
||||
filings = {"X": FilingMeta(date(2026, 3, 31), date(2026, 5, 1),
|
||||
datetime(2026, 5, 1, 10, tzinfo=UTC), "10-Q")}
|
||||
res = parse_snapshots(companyfacts, filings, {"X"}, fiscal_year_end="1231")
|
||||
assert res.rows[0].diluted_eps == 0.60
|
||||
|
||||
|
||||
def test_diluted_still_wins_over_basic_when_both_present():
|
||||
companyfacts = {
|
||||
"cik": 320193,
|
||||
"facts": {"us-gaap": {
|
||||
"EarningsPerShareDiluted": {
|
||||
"units": {"USD/shares": [_dur("2026-01-01", "2026-03-31", 1.36, "X", fp="Q1")]}},
|
||||
"EarningsPerShareBasic": {
|
||||
"units": {"USD/shares": [_dur("2026-01-01", "2026-03-31", 1.40, "X", fp="Q1")]}},
|
||||
}},
|
||||
}
|
||||
filings = {"X": FilingMeta(date(2026, 3, 31), date(2026, 5, 1),
|
||||
datetime(2026, 5, 1, 10, tzinfo=UTC), "10-Q")}
|
||||
res = parse_snapshots(companyfacts, filings, {"X"}, fiscal_year_end="1231")
|
||||
assert res.rows[0].diluted_eps == 1.36
|
||||
|
||||
|
||||
def test_weighted_average_shares_prefers_the_shortest_span():
|
||||
# A 10-Q carries both the quarter's average and the YTD one. The shorter
|
||||
# window sits closer to the current count, which is what market cap wants.
|
||||
companyfacts = {
|
||||
"cik": 1326801,
|
||||
"facts": {"us-gaap": {"WeightedAverageNumberOfDilutedSharesOutstanding": {
|
||||
"units": {"shares": [
|
||||
_dur("2026-01-01", "2026-09-30", 2_600_000_000, "X", fp="Q3"), # YTD
|
||||
_dur("2026-07-01", "2026-09-30", 2_564_000_000, "X", fp="Q3"), # quarter
|
||||
]}
|
||||
}}},
|
||||
}
|
||||
filings = {"X": FilingMeta(date(2026, 9, 30), date(2026, 11, 1),
|
||||
datetime(2026, 11, 1, 10, tzinfo=UTC), "10-Q")}
|
||||
res = parse_snapshots(companyfacts, filings, {"X"}, fiscal_year_end="1231")
|
||||
assert res.rows[0].weighted_avg_diluted_shares == 2_564_000_000
|
||||
|
||||
|
||||
def test_weighted_average_shares_falls_back_to_the_basic_and_diluted_concept():
|
||||
companyfacts = {
|
||||
"cik": 1326801,
|
||||
"facts": {"us-gaap": {"WeightedAverageNumberOfSharesOutstandingBasicAndDiluted": {
|
||||
"units": {"shares": [_dur("2026-07-01", "2026-09-30", 500_000, "X", fp="Q3")]}
|
||||
}}},
|
||||
}
|
||||
filings = {"X": FilingMeta(date(2026, 9, 30), date(2026, 11, 1),
|
||||
datetime(2026, 11, 1, 10, tzinfo=UTC), "10-Q")}
|
||||
res = parse_snapshots(companyfacts, filings, {"X"}, fiscal_year_end="1231")
|
||||
assert res.rows[0].weighted_avg_diluted_shares == 500_000
|
||||
|
||||
|
||||
# -- fiscal year end resolution (submissions.fiscalYearEnd is unreliable) -----
|
||||
|
||||
from app.services.sec_facts_parser import resolve_fiscal_year_end # noqa: E402
|
||||
|
||||
|
||||
def test_the_issuers_own_10k_overrides_a_wrong_declared_year_end():
|
||||
# Franklin Resources declares 1231 while every 10-K ends 09-30. Trusting the
|
||||
# declaration labelled its fiscal Q2 (Mar) as Q1, colliding with the real
|
||||
# fiscal Q1 (Dec) and destroying the quarter chain.
|
||||
filings = {
|
||||
"K": _meta("2025-09-30", "10-K"),
|
||||
"Q": _meta("2025-12-31"),
|
||||
}
|
||||
assert resolve_fiscal_year_end(filings, "1231") == "0930"
|
||||
|
||||
|
||||
def test_declared_year_end_is_used_when_no_annual_filing_is_present():
|
||||
assert resolve_fiscal_year_end({"Q": _meta("2026-03-31")}, "1231") == "1231"
|
||||
assert resolve_fiscal_year_end({}, None) is None
|
||||
|
||||
|
||||
def test_a_wrong_declared_year_end_no_longer_collides_two_periods():
|
||||
"""End to end: BEN's Dec and Mar quarters must land on distinct keys."""
|
||||
def _q(accn, start, end, val):
|
||||
return _dur(start, end, val, accn, fp="Q1")
|
||||
|
||||
companyfacts = {
|
||||
"cik": 38777,
|
||||
"facts": {"us-gaap": {"Revenues": {"units": {"USD": [
|
||||
_q("Q1", "2025-10-01", "2025-12-31", 2327), # fiscal Q1
|
||||
_q("Q2", "2025-10-01", "2026-03-31", 4622), # fiscal Q2 YTD
|
||||
]}}}},
|
||||
}
|
||||
filings = {
|
||||
"K": _meta("2025-09-30", "10-K"),
|
||||
"Q1": _meta("2025-12-31"),
|
||||
"Q2": _meta("2026-03-31"),
|
||||
}
|
||||
res = parse_snapshots(companyfacts, filings, {"Q1", "Q2"}, fiscal_year_end="1231")
|
||||
keys = {(r.fiscal_year, r.fiscal_period) for r in res.rows}
|
||||
assert len(keys) == 2, f"periods collided on one key: {keys}"
|
||||
assert keys == {(2026, "Q1"), (2026, "Q2")}
|
||||
|
||||
@@ -3,6 +3,7 @@ import framework with a fake SEC client (no network)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import tempfile
|
||||
from datetime import date, datetime, timezone
|
||||
@@ -14,9 +15,20 @@ from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker, create_asyn
|
||||
from app.database import Base
|
||||
import app.models # noqa: F401
|
||||
from app.models.fundamental_snapshot import FundamentalSnapshot
|
||||
from app.models.sec_filing_gap import SecFilingGap
|
||||
from app.models.system_event import SystemEvent
|
||||
from app.models.ticker import Ticker
|
||||
from app.services.data_import import STATUS_FAILED, STATUS_PROMOTED, run_import
|
||||
from app.services.sec_fundamentals_importer import SecFundamentalsImporter
|
||||
from app.services.data_import import (
|
||||
STATUS_DEFERRED,
|
||||
STATUS_FAILED,
|
||||
STATUS_PROMOTED,
|
||||
run_import,
|
||||
)
|
||||
from app.services.sec_fundamentals_importer import (
|
||||
SecFundamentalsImporter,
|
||||
StagedFundamentals,
|
||||
)
|
||||
from app.services.sec_universe import ResolvedUniverse
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
@@ -189,7 +201,7 @@ async def test_incremental_adds_only_new_filing(engine):
|
||||
assert q2.fiscal_period == "Q2" and q2.revenue == 254940
|
||||
|
||||
|
||||
async def test_consistency_gate_fails_when_facts_lag_index(engine):
|
||||
async def test_consistency_gate_defers_without_alert_when_facts_lag_index(engine):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
@@ -211,9 +223,536 @@ async def test_consistency_gate_fails_when_facts_lag_index(engine):
|
||||
)
|
||||
run = await run_import(_importer(incr, today=date(2026, 5, 3)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_FAILED
|
||||
assert "Company Facts" in (run.error_details or "")
|
||||
assert run.status == STATUS_DEFERRED
|
||||
# The gate blocks every later run until it clears, so run history still has
|
||||
# to name the filing and say why it could not be resolved.
|
||||
details = run.error_details or ""
|
||||
assert "GHOST" in details and "not_in_companyfacts" in details
|
||||
assert "2026-05-01" in details # index date the filing was seen on
|
||||
summary = json.loads(run.validation_json or "{}")
|
||||
assert summary["missing_xbrl_count"] == 1
|
||||
assert summary["missing_xbrl"][0]["accession"] == "GHOST"
|
||||
assert summary["missing_xbrl"][0]["form"] == "10-Q"
|
||||
assert await _count(factory, FundamentalSnapshot) == 2 # nothing new written
|
||||
assert await _count(factory, SystemEvent) == 0 # expected SEC lag does not alert
|
||||
|
||||
|
||||
async def test_companyfacts_lag_does_not_mask_second_validation_failure():
|
||||
importer = SecFundamentalsImporter(today=date(2026, 5, 3))
|
||||
importer._latest_index_date = date(2026, 5, 2)
|
||||
staged = StagedFundamentals(
|
||||
resolved=ResolvedUniverse(),
|
||||
missing_xbrl=[{
|
||||
"cik": "0000320193",
|
||||
"accession": "GHOST",
|
||||
"form": "10-Q",
|
||||
"index_date": date(2026, 5, 1),
|
||||
"age_days": 2,
|
||||
"reason": "not_in_companyfacts",
|
||||
}],
|
||||
invalid_payloads=[{
|
||||
"cik": "0000789019",
|
||||
"reason": "missing facts structure",
|
||||
}],
|
||||
)
|
||||
|
||||
result = await importer.validate(None, staged)
|
||||
|
||||
assert not result.ok
|
||||
assert not result.retryable
|
||||
assert len(result.messages) == 2
|
||||
|
||||
|
||||
async def test_deferred_alert_names_aged_out_accessions_separately():
|
||||
importer = SecFundamentalsImporter(today=date(2026, 5, 6))
|
||||
importer._latest_index_date = date(2026, 5, 5)
|
||||
staged = StagedFundamentals(
|
||||
resolved=ResolvedUniverse(),
|
||||
missing_xbrl=[
|
||||
{
|
||||
"cik": "0000320193",
|
||||
"accession": "YOUNG",
|
||||
"form": "10-Q",
|
||||
"index_date": date(2026, 5, 5),
|
||||
"age_days": 1,
|
||||
"reason": "not_in_companyfacts",
|
||||
},
|
||||
{
|
||||
"cik": "0000789019",
|
||||
"accession": "AGED-OUT",
|
||||
"form": "10-Q",
|
||||
"index_date": date(2026, 5, 1),
|
||||
"age_days": 5,
|
||||
"reason": "not_in_companyfacts",
|
||||
},
|
||||
],
|
||||
)
|
||||
|
||||
result = await importer.validate(None, staged)
|
||||
|
||||
assert result.retryable
|
||||
assert len(result.messages) == 1 and "YOUNG" in result.messages[0]
|
||||
assert len(result.deferred_alert_messages) == 1
|
||||
assert "AGED-OUT" in result.deferred_alert_messages[0]
|
||||
|
||||
|
||||
async def test_gate_separates_missing_submissions_from_missing_facts(engine):
|
||||
"""An index row the issuer's own filing list does not carry is a different
|
||||
failure from a Company-Facts lag, and must not be reported as one."""
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
|
||||
incr = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)}, # submissions never lists ORPHAN
|
||||
latest_index=date(2026, 5, 2),
|
||||
daily={date(2026, 5, 1): [{"form": "10-Q", "cik": 320193, "accession": "ORPHAN"}]},
|
||||
)
|
||||
run = await run_import(_importer(incr, today=date(2026, 5, 3)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_FAILED
|
||||
assert "ORPHAN" in (run.error_details or "")
|
||||
assert "not_in_submissions" in (run.error_details or "")
|
||||
summary = json.loads(run.validation_json or "{}")
|
||||
assert summary["missing_xbrl"][0]["reason"] == "not_in_submissions"
|
||||
|
||||
|
||||
def _coregistrant_client(share_fact):
|
||||
"""Incremental client where Q2A's facts landed in co-registrant 99999's file
|
||||
instead of the filer's own — the NEE-via-FPL / DOW-via-Dow-Chemical shape."""
|
||||
cf_q2 = _rev("2025-09-28", "2026-03-28", 254940, 2026, "Q2", "Q2A")
|
||||
return FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={
|
||||
320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1]), # filer's own: no Q2A
|
||||
99999: _companyfacts([cf_q2], [share_fact], cik=99999),
|
||||
},
|
||||
submissions={320193: _submissions(SUB_FILINGS + [
|
||||
_filing("Q2A", "10-Q", "2026-03-28", "2026-05-01", "2026-05-01T10:01:00.000Z")])},
|
||||
latest_index=date(2026, 5, 2),
|
||||
daily={date(2026, 5, 1): [
|
||||
# one combined filing, listed by the index under both co-registrants
|
||||
{"form": "10-Q", "cik": 320193, "accession": "Q2A"},
|
||||
{"form": "10-Q", "cik": 99999, "accession": "Q2A"},
|
||||
]},
|
||||
)
|
||||
|
||||
|
||||
async def test_recovers_facts_misfiled_under_coregistrant(engine):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
|
||||
# 14687 shares is continuous with the issuer's own history (14681 last quarter).
|
||||
incr = _coregistrant_client(_shares("2026-04-17", 14687, "Q2A", 2026, "Q2"))
|
||||
run = await run_import(_importer(incr, today=date(2026, 5, 3)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_PROMOTED
|
||||
summary = json.loads(run.validation_json or "{}")
|
||||
assert summary["recovered_count"] == 1
|
||||
assert summary["recovered_from_coregistrant"][0]["source_cik"] == "0000099999"
|
||||
assert summary["missing_xbrl_count"] == 0
|
||||
|
||||
async with factory() as s:
|
||||
q2 = (await s.execute(
|
||||
select(FundamentalSnapshot).where(FundamentalSnapshot.accession == "Q2A")
|
||||
)).scalar_one()
|
||||
codes = (await s.execute(select(SystemEvent.code))).scalars().all()
|
||||
# Stamped to the issuer that filed, NOT the co-registrant whose file it came from.
|
||||
assert q2.cik == "0000320193"
|
||||
assert q2.revenue == 254940 and q2.shares_outstanding == 14687
|
||||
assert "coregistrant_recovery" not in codes
|
||||
|
||||
|
||||
async def test_coregistrant_recovery_rejects_discontinuous_share_count(engine):
|
||||
"""A co-registrant shell's standalone facts must never be stored as the
|
||||
parent's — a token float is the signature and it has to be refused."""
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
|
||||
incr = _coregistrant_client(_shares("2026-04-17", 100, "Q2A", 2026, "Q2"))
|
||||
run = await run_import(_importer(incr, today=date(2026, 5, 3)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_FAILED
|
||||
assert "coregistrant_facts_rejected" in (run.error_details or "")
|
||||
assert await _count(factory, FundamentalSnapshot) == 2 # nothing recovered
|
||||
|
||||
|
||||
async def test_unresolved_filing_stops_blocking_after_retry_window(engine):
|
||||
"""A filing SEC has misfiled must not wedge every later import forever."""
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
|
||||
incr = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])}, # no GHOST
|
||||
submissions={320193: _submissions(SUB_FILINGS + [
|
||||
_filing("GHOST", "10-Q", "2026-03-28", "2026-05-01", "2026-05-01T10:01:00.000Z")])},
|
||||
latest_index=date(2026, 5, 2),
|
||||
daily={date(2026, 5, 1): [{"form": "10-Q", "cik": 320193, "accession": "GHOST"}]},
|
||||
)
|
||||
# 9 days after the index date — well past the retry window.
|
||||
run = await run_import(_importer(incr, today=date(2026, 5, 10)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_PROMOTED # promoted around it, not blocked by it
|
||||
assert run.source_max_date == date(2026, 5, 2) # and the index advances
|
||||
summary = json.loads(run.validation_json or "{}")
|
||||
assert summary["missing_xbrl_count"] == 1 and summary["missing_xbrl_blocking"] == 0
|
||||
assert await _count(factory, SecFilingGap) == 1
|
||||
|
||||
async with factory() as s:
|
||||
events = (await s.execute(select(SystemEvent))).scalars().all()
|
||||
unresolved = [e for e in events if e.code == "unresolved_filing"]
|
||||
assert len(unresolved) == 1 and "GHOST" in unresolved[0].message
|
||||
assert "automatic SEC retry" in unresolved[0].message
|
||||
|
||||
# The next scheduled run retries even though the SEC daily-index revision
|
||||
# has not changed. Company Facts is a separate SEC product and may catch up
|
||||
# independently, so the generic revision no-op must not suppress this work.
|
||||
still_missing_run = await run_import(
|
||||
_importer(incr, today=date(2026, 5, 11)),
|
||||
engine=engine,
|
||||
)
|
||||
assert still_missing_run.status == STATUS_PROMOTED
|
||||
assert still_missing_run.revision is None
|
||||
assert await _count(factory, SecFilingGap) == 1
|
||||
|
||||
# A later normal scheduled import retries only the queued issuer. Once SEC
|
||||
# publishes the accession in Company Facts, it is inserted and unblocked
|
||||
# without a full-universe reparse or operator action.
|
||||
cf_ghost = _rev("2025-09-28", "2026-03-28", 254940, 2026, "Q2", "GHOST")
|
||||
sh_ghost = _shares("2026-04-17", 14687, "GHOST", 2026, "Q2")
|
||||
healed = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={
|
||||
320193: _companyfacts(
|
||||
[CF_K, CF_Q1, cf_ghost],
|
||||
[SH_K, SH_Q1, sh_ghost],
|
||||
)
|
||||
},
|
||||
submissions={
|
||||
320193: _submissions(SUB_FILINGS + [
|
||||
_filing(
|
||||
"GHOST",
|
||||
"10-Q",
|
||||
"2026-03-28",
|
||||
"2026-05-01",
|
||||
"2026-05-01T10:01:00.000Z",
|
||||
)
|
||||
])
|
||||
},
|
||||
latest_index=date(2026, 5, 2),
|
||||
)
|
||||
healed_run = await run_import(
|
||||
_importer(healed, today=date(2026, 5, 12)),
|
||||
engine=engine,
|
||||
)
|
||||
|
||||
assert healed_run.status == STATUS_PROMOTED
|
||||
assert await _count(factory, FundamentalSnapshot) == 3
|
||||
assert await _count(factory, SecFilingGap) == 0
|
||||
|
||||
|
||||
async def test_queued_gap_without_index_date_retries_without_wedging_and_escalates_once(
|
||||
engine,
|
||||
):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
|
||||
old = datetime(2026, 4, 1, tzinfo=timezone.utc)
|
||||
async with factory() as db:
|
||||
db.add(SecFilingGap(
|
||||
cik="0000320193",
|
||||
accession="DATELESS",
|
||||
form="10-Q",
|
||||
index_date=None,
|
||||
reason="not_in_companyfacts",
|
||||
first_seen_at=old,
|
||||
last_attempted_at=old,
|
||||
))
|
||||
await db.commit()
|
||||
|
||||
missing = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS + [
|
||||
_filing(
|
||||
"DATELESS",
|
||||
"10-Q",
|
||||
"2026-03-28",
|
||||
"2026-05-01",
|
||||
"2026-05-01T10:01:00.000Z",
|
||||
)
|
||||
])},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
first = await run_import(
|
||||
_importer(missing, today=date(2026, 5, 20)), engine=engine
|
||||
)
|
||||
second = await run_import(
|
||||
_importer(missing, today=date(2026, 5, 21)), engine=engine
|
||||
)
|
||||
|
||||
assert first.status == STATUS_PROMOTED
|
||||
assert second.status == STATUS_PROMOTED
|
||||
async with factory() as db:
|
||||
gap = (await db.execute(select(SecFilingGap))).scalar_one()
|
||||
events = (
|
||||
await db.execute(
|
||||
select(SystemEvent).where(SystemEvent.code == "filing_gap_aged")
|
||||
)
|
||||
).scalars().all()
|
||||
assert gap.escalated_at is not None
|
||||
assert len(events) == 1
|
||||
|
||||
|
||||
async def test_queued_filing_reclassified_non_xbrl_is_removed(engine):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
now = datetime.now(timezone.utc)
|
||||
async with factory() as db:
|
||||
db.add(SecFilingGap(
|
||||
cik="0000320193",
|
||||
accession="NONX",
|
||||
form="10-Q/A",
|
||||
index_date=date(2026, 5, 1),
|
||||
reason="not_in_companyfacts",
|
||||
first_seen_at=now,
|
||||
last_attempted_at=now,
|
||||
))
|
||||
await db.commit()
|
||||
|
||||
client = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS + [
|
||||
_filing(
|
||||
"NONX",
|
||||
"10-Q/A",
|
||||
"2026-03-28",
|
||||
"2026-05-01",
|
||||
"2026-05-01T10:01:00.000Z",
|
||||
is_xbrl=False,
|
||||
)
|
||||
])},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
run = await run_import(_importer(client, today=date(2026, 5, 20)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_PROMOTED
|
||||
assert await _count(factory, SecFilingGap) == 0
|
||||
|
||||
|
||||
async def test_queued_parser_skip_stays_blocked_with_actionable_reason(
|
||||
engine, monkeypatch
|
||||
):
|
||||
from app.services.sec_facts_parser import ParseResult
|
||||
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
now = datetime.now(timezone.utc)
|
||||
async with factory() as db:
|
||||
db.add(SecFilingGap(
|
||||
cik="0000320193",
|
||||
accession="BADPARSE",
|
||||
form="10-Q",
|
||||
index_date=date(2026, 5, 1),
|
||||
reason="not_in_companyfacts",
|
||||
first_seen_at=now,
|
||||
last_attempted_at=now,
|
||||
))
|
||||
await db.commit()
|
||||
|
||||
bad_fact = _rev(
|
||||
"2025-09-28", "2026-03-28", 254940, 2026, "Q2", "BADPARSE"
|
||||
)
|
||||
bad_share = _shares("2026-04-17", 14687, "BADPARSE", 2026, "Q2")
|
||||
client = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={
|
||||
320193: _companyfacts(
|
||||
[CF_K, CF_Q1, bad_fact], [SH_K, SH_Q1, bad_share]
|
||||
)
|
||||
},
|
||||
submissions={320193: _submissions(SUB_FILINGS + [
|
||||
_filing(
|
||||
"BADPARSE",
|
||||
"10-Q",
|
||||
"2026-03-28",
|
||||
"2026-05-01",
|
||||
"2026-05-01T10:01:00.000Z",
|
||||
)
|
||||
])},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
|
||||
def skip_parse(*args, **kwargs):
|
||||
return ParseResult(skipped_filings=[{
|
||||
"accession": "BADPARSE",
|
||||
"reason": "unparseable",
|
||||
}])
|
||||
|
||||
monkeypatch.setattr("app.services.sec_facts_parser.parse_snapshots", skip_parse)
|
||||
run = await run_import(_importer(client, today=date(2026, 5, 20)), engine=engine)
|
||||
|
||||
assert run.status == STATUS_PROMOTED
|
||||
async with factory() as db:
|
||||
gap = (await db.execute(select(SecFilingGap))).scalar_one()
|
||||
assert gap.reason == "parser_unusable"
|
||||
|
||||
|
||||
async def test_new_parser_skip_gets_grace_then_enters_retry_queue(
|
||||
engine, monkeypatch
|
||||
):
|
||||
from app.services.sec_facts_parser import ParseResult
|
||||
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
backfill = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K, CF_Q1], [SH_K, SH_Q1])},
|
||||
submissions={320193: _submissions(SUB_FILINGS)},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
await run_import(_importer(backfill), engine=engine)
|
||||
|
||||
bad_fact = _rev(
|
||||
"2025-09-28", "2026-03-28", 254940, 2026, "Q2", "NEWBAD"
|
||||
)
|
||||
bad_share = _shares("2026-04-17", 14687, "NEWBAD", 2026, "Q2")
|
||||
client = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={
|
||||
320193: _companyfacts(
|
||||
[CF_K, CF_Q1, bad_fact], [SH_K, SH_Q1, bad_share]
|
||||
)
|
||||
},
|
||||
submissions={320193: _submissions(SUB_FILINGS + [
|
||||
_filing(
|
||||
"NEWBAD",
|
||||
"10-Q",
|
||||
"2026-03-28",
|
||||
"2026-05-01",
|
||||
"2026-05-01T10:01:00.000Z",
|
||||
)
|
||||
])},
|
||||
latest_index=date(2026, 5, 2),
|
||||
daily={
|
||||
date(2026, 5, 1): [{
|
||||
"form": "10-Q",
|
||||
"cik": 320193,
|
||||
"accession": "NEWBAD",
|
||||
}]
|
||||
},
|
||||
)
|
||||
|
||||
def skip_parse(*args, **kwargs):
|
||||
return ParseResult(skipped_filings=[{
|
||||
"accession": "NEWBAD",
|
||||
"reason": "unparseable",
|
||||
}])
|
||||
|
||||
monkeypatch.setattr("app.services.sec_facts_parser.parse_snapshots", skip_parse)
|
||||
|
||||
young = await run_import(
|
||||
_importer(client, today=date(2026, 5, 3)), engine=engine
|
||||
)
|
||||
assert young.status == STATUS_DEFERRED
|
||||
assert "parser_unusable" in (young.error_details or "")
|
||||
assert await _count(factory, SecFilingGap) == 0
|
||||
|
||||
aged = await run_import(
|
||||
_importer(client, today=date(2026, 5, 5)), engine=engine
|
||||
)
|
||||
assert aged.status == STATUS_PROMOTED
|
||||
async with factory() as db:
|
||||
gap = (await db.execute(select(SecFilingGap))).scalar_one()
|
||||
assert gap.accession == "NEWBAD"
|
||||
assert gap.reason == "parser_unusable"
|
||||
|
||||
|
||||
async def test_validation_caps_details_but_keeps_complete_blocked_cik_set():
|
||||
importer = SecFundamentalsImporter(today=date(2026, 5, 20))
|
||||
importer._latest_index_date = date(2026, 5, 19)
|
||||
staged = StagedFundamentals(
|
||||
resolved=ResolvedUniverse(),
|
||||
missing_xbrl=[
|
||||
{
|
||||
"cik": f"{i:010d}",
|
||||
"accession": f"MISS-{i}",
|
||||
"form": "10-Q",
|
||||
"index_date": date(2026, 5, 1),
|
||||
"age_days": 19,
|
||||
"reason": "not_in_companyfacts",
|
||||
}
|
||||
for i in range(60)
|
||||
],
|
||||
no_xbrl_filings=[
|
||||
{"cik": f"{i + 100:010d}", "name": f"New {i}"}
|
||||
for i in range(60)
|
||||
],
|
||||
recovered=[
|
||||
{"cik": f"{i:010d}", "accession": f"REC-{i}", "source_cik": "1"}
|
||||
for i in range(60)
|
||||
],
|
||||
)
|
||||
|
||||
result = await importer.validate(None, staged)
|
||||
|
||||
assert len(result.summary["missing_xbrl"]) == 50
|
||||
assert len(result.summary["no_xbrl_filings"]) == 50
|
||||
assert len(result.summary["no_xbrl_ciks"]) == 60
|
||||
assert len(result.summary["recovered_from_coregistrant"]) == 50
|
||||
assert len(result.summary["setup_blocked_ciks"]) == 120
|
||||
|
||||
|
||||
async def test_non_xbrl_amendment_skipped_not_failed(engine):
|
||||
@@ -411,3 +950,181 @@ async def test_discrepancy_in_shares_is_detected_and_reported(engine):
|
||||
assert k.shares_outstanding == 999.0 and k.import_run_id == 1 # immutable — not overwritten
|
||||
events = (await s.execute(select(SystemEvent).where(SystemEvent.code == "snapshot_discrepancy"))).scalars().all()
|
||||
assert len(events) == 1 and events[0].severity == "warning"
|
||||
|
||||
|
||||
# --- reparse: rewriting rows a fixed parser reconstructs differently --------
|
||||
|
||||
# A 4-4-5 filer's YTD-Q3 span (36 weeks = 251 days). The old 20-day tolerance
|
||||
# around 273 rejected it and stored revenue=None; 25 accepts it. Reparsing with
|
||||
# the fixed parser is exactly the situation this mode exists for.
|
||||
CF_Q3_445 = _rev("2025-09-01", "2026-05-10", 207431, 2026, "Q3", "Q3F")
|
||||
SUB_445 = [_filing("Q3F", "10-Q", "2026-05-10", "2026-06-01", "2026-06-01T10:01:00.000Z")]
|
||||
|
||||
|
||||
def _445_client():
|
||||
return FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_Q3_445], [_shares("2026-05-15", 100, "Q3F", 2026, "Q3")])},
|
||||
submissions={320193: _submissions(SUB_445)},
|
||||
latest_index=date(2026, 6, 1),
|
||||
)
|
||||
|
||||
|
||||
async def _import_with_old_tolerance(engine, monkeypatch):
|
||||
"""Seed the DB the way the pre-fix parser did: Q3 revenue rejected -> null."""
|
||||
from app.services import sec_facts_parser
|
||||
|
||||
monkeypatch.setattr(sec_facts_parser, "_YTD_TOLERANCE_DAYS", 20)
|
||||
run = await run_import(_importer(_445_client(), today=date(2026, 6, 2)), engine=engine)
|
||||
monkeypatch.undo()
|
||||
return run
|
||||
|
||||
|
||||
async def test_reparse_rewrites_rows_the_fixed_parser_reads_differently(engine, monkeypatch):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
first = await _import_with_old_tolerance(engine, monkeypatch)
|
||||
|
||||
async with factory() as s:
|
||||
stale = (await s.execute(select(FundamentalSnapshot))).scalar_one()
|
||||
assert stale.revenue is None, "precondition: the old parser stored a null"
|
||||
|
||||
# Reparse with the current (fixed) parser. force=True because SEC has not
|
||||
# changed -- the staleness is on our side, so the revision gate would no-op.
|
||||
run = await run_import(
|
||||
SecFundamentalsImporter(
|
||||
client_factory=lambda: _445_client(), today=date(2026, 6, 2), reparse=True
|
||||
),
|
||||
engine=engine,
|
||||
force=True,
|
||||
)
|
||||
|
||||
assert run.status == STATUS_PROMOTED
|
||||
assert '"updated": 1' in run.row_counts_json
|
||||
async with factory() as s:
|
||||
fixed = (await s.execute(select(FundamentalSnapshot))).scalar_one()
|
||||
assert fixed.revenue == 207431 # rewritten in place
|
||||
assert fixed.accession == stale.accession
|
||||
assert fixed.import_run_id == run.id # rewrite is attributable
|
||||
assert fixed.import_run_id != first.id
|
||||
|
||||
|
||||
async def test_reparse_leaves_unchanged_rows_untouched(engine):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
first = await run_import(_importer(_445_client(), today=date(2026, 6, 2)), engine=engine)
|
||||
|
||||
run = await run_import(
|
||||
SecFundamentalsImporter(
|
||||
client_factory=lambda: _445_client(), today=date(2026, 6, 2), reparse=True
|
||||
),
|
||||
engine=engine,
|
||||
force=True,
|
||||
)
|
||||
|
||||
assert '"updated": 0' in run.row_counts_json
|
||||
async with factory() as s:
|
||||
row = (await s.execute(select(FundamentalSnapshot))).scalar_one()
|
||||
assert row.import_run_id == first.id # provenance preserved, no needless rewrite
|
||||
|
||||
|
||||
async def test_without_reparse_a_differing_row_stays_immutable(engine, monkeypatch):
|
||||
"""The default contract is unchanged: report the discrepancy, never mutate."""
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
await _import_with_old_tolerance(engine, monkeypatch)
|
||||
|
||||
importer = SecFundamentalsImporter(
|
||||
client_factory=lambda: _445_client(), today=date(2026, 6, 2), reparse=True
|
||||
)
|
||||
async with _factory(engine)() as db:
|
||||
await importer.detect_revision(db)
|
||||
staged = await importer.stage(db)
|
||||
importer.reparse = False # same staged diff, default disposition
|
||||
counts = await importer.promote(db, staged, run_id=999)
|
||||
await db.commit()
|
||||
|
||||
assert staged.discrepancies, "the diff should still be detected and reported"
|
||||
assert counts["updated"] == 0
|
||||
async with factory() as s:
|
||||
row = (await s.execute(select(FundamentalSnapshot))).scalar_one()
|
||||
assert row.revenue is None # untouched
|
||||
|
||||
|
||||
async def test_force_bypasses_the_unchanged_revision_no_op(engine):
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
await run_import(_importer(_445_client(), today=date(2026, 6, 2)), engine=engine)
|
||||
|
||||
same = _importer(_445_client(), today=date(2026, 6, 2))
|
||||
assert (await run_import(same, engine=engine)).status == "no_op"
|
||||
|
||||
forced = SecFundamentalsImporter(
|
||||
client_factory=lambda: _445_client(), today=date(2026, 6, 2), reparse=True
|
||||
)
|
||||
assert (await run_import(forced, engine=engine, force=True)).status == STATUS_PROMOTED
|
||||
|
||||
|
||||
# --- CIK resolution: successor registrants with no filings -----------------
|
||||
|
||||
async def test_issuer_with_no_xbrl_filings_is_reported_not_silent(engine):
|
||||
"""XOM resolved to CIK 2115436 'ExxonMobil Holdings Corp', which has zero
|
||||
filings, so it produced no snapshots and nothing said why."""
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
client = FakeSecClient(
|
||||
tickers={"AAPL": 320193},
|
||||
companyfacts={320193: _companyfacts([CF_K], [SH_K])},
|
||||
submissions={320193: {**_submissions([]), "name": "Shell Holdings Corp"}},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
importer = _importer(client)
|
||||
async with factory() as db:
|
||||
await importer.detect_revision(db)
|
||||
staged = await importer.stage(db)
|
||||
result = await importer.validate(db, staged)
|
||||
|
||||
assert result.summary["no_xbrl_filings_count"] == 1
|
||||
assert staged.no_xbrl_filings[0]["cik"] == "0000320193"
|
||||
assert staged.no_xbrl_filings[0]["name"] == "Shell Holdings Corp"
|
||||
|
||||
|
||||
async def test_cik_override_pins_a_ticker_to_the_real_filer(engine):
|
||||
from app.models.settings import SystemSetting
|
||||
from app.services.sec_universe import CIK_OVERRIDES_KEY, resolve_ciks
|
||||
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
async with factory() as s:
|
||||
s.add(SystemSetting(key=CIK_OVERRIDES_KEY, value='{"AAPL": 34088}'))
|
||||
await s.commit()
|
||||
|
||||
client = FakeSecClient(
|
||||
tickers={"AAPL": 320193}, # SEC points at the wrong registrant
|
||||
companyfacts={}, submissions={}, latest_index=date(2026, 1, 31),
|
||||
)
|
||||
async with factory() as db:
|
||||
resolved = await resolve_ciks(db, client)
|
||||
|
||||
assert resolved.symbol_to_cik["AAPL"] == 34088
|
||||
assert resolved.cik_updates == [(1, "0000034088")]
|
||||
|
||||
|
||||
async def test_malformed_cik_override_is_ignored_not_fatal(engine):
|
||||
from app.models.settings import SystemSetting
|
||||
from app.services.sec_universe import CIK_OVERRIDES_KEY, resolve_ciks
|
||||
|
||||
factory = _factory(engine)
|
||||
await _seed(factory, ["AAPL"])
|
||||
async with factory() as s:
|
||||
s.add(SystemSetting(key=CIK_OVERRIDES_KEY, value="not json at all"))
|
||||
await s.commit()
|
||||
|
||||
client = FakeSecClient(
|
||||
tickers={"AAPL": 320193}, companyfacts={}, submissions={},
|
||||
latest_index=date(2026, 1, 31),
|
||||
)
|
||||
async with factory() as db:
|
||||
resolved = await resolve_ciks(db, client)
|
||||
|
||||
assert resolved.symbol_to_cik["AAPL"] == 320193 # fell back to company_tickers
|
||||
|
||||
Reference in New Issue
Block a user