An AI equity-research analyst that turns SEC filings, fundamentals and earnings calls into cited theses with falsifiable kill criteria — then monitors them, and keeps score.
Built solo with Claude Code (Opus + Haiku in the product; agents under my direction for most of the code). What I'd point to is not the volume but the boundary: deterministic finance math and routing in code, judgment in the model, and an eval harness and CI gates that check the model's work.
Personal research tool, not investment advice. It has no track record worth claiming yet — see Keeping score.
One of nine deep-dive categories for Vertiv (VRT). Every finding carries a [Source: …] tag —
an FMP endpoint, an SEC XBRL fact, or the deterministic quant layer.
| Research run | ~$2 and ~7 minutes (VRT, two risk loops: $2.07, 6.6 min, 25 model calls) |
| Stack | FastAPI · async SQLAlchemy · Postgres · Next.js 16 · Claude Opus 5.5 + Haiku 4.5 |
| Tests | 1,100 backend tests, an end-to-end research run on a real Postgres, Playwright + axe on every page, a live smoke of all 19 model call paths |
| Evals | thesis grounding, citation and consistency checks; rerun dispersion; a cross-model rubric judge |
| History | 657 commits since April 2026, 82% co-authored with Claude; 5 ADRs |
1. Cited research → thesis → kill criteria → monitoring. Push a ticker through a four-phase pipeline. A fast screen, then nine deep-dive categories in parallel, each fed FMP statements, transcripts, SEC filing excerpts, FRED macro data and a deterministic quant layer (Piotroski, Altman, Beneish, accruals). A thesis step turns them into a stance (long / avoid / short), 12-month price targets, catalysts and kill criteria. A risk step stress-tests it. The status board then tracks every live thesis — health, the next catalyst, which kill criteria have tripped — and the performance page scores each call against SPY, a beta-adjusted SPY, the sector ETF and the theme.
2. SEC filings → supply-chain graph. Filings are pulled from EDGAR and split into sections; a Haiku pass extracts customers, suppliers, partners and competitors with verbatim quotes; names are resolved to tickers by fuzzy match against EDGAR's company list, with a curation queue for the rest. The graph feeds the deep-dive prompts ("use these as anchors, don't re-quote") and a two-hop explorer, so a question like "what does my universe say about CoreWeave's suppliers?" has an answer.
3. Measuring the model. Every model call is recorded with its tokens, latency and cost; theses are checked against the data they were written from; repeated runs measure how stable the output is; a second model grades quality. See Evidence.
Also: an editable three-statement model per ticker, AI-seeded, with a reverse DCF that solves what the market price implies; a thesis-refresh loop after earnings; an 8-K and insider/congress-trade scanner; a trade journal. docs/ARCHITECTURE.md is the full reference.
flowchart LR
subgraph sources[Data]
FMP[FMP: statements, estimates,<br/>transcripts, 13F, insiders]
SEC[SEC EDGAR: filings, XBRL]
FRED[FRED: macro series]
end
subgraph api[FastAPI]
P[Research pipeline<br/>explicit state machine]
F[Filings: extract → resolve → graph]
M[3-statement model + reverse DCF]
W[Monitoring: status board, catalysts,<br/>8-K scan, outcome tracking]
L[llm.py: structured outputs,<br/>caching, telemetry]
end
DB[(Postgres:<br/>runs, relationships,<br/>outcomes, llm_calls)]
UI[Next.js 16 app] <-->|REST + SSE| api
sources --> api
api <--> DB
L --> C[Claude API]
flowchart LR
QS[quick_screen<br/>Haiku] --> DD[deep_dive<br/>9 categories · Opus]
DD --> TH[thesis<br/>stance · targets · kill criteria]
TH --> RK[risk stress test]
RK -->|model asks, R/R < 2,<br/>loops < 2 — decided in code| DD
RK --> DONE[completed] --> OUT[outcome tracking<br/>next trading day's close]
- The model judges; code decides. Reward/risk is computed from the price and the model's targets,
not taken from the model; whether the risk step loops back is a rule in
graph/routing.py(before, the prompt asked the model to apply the rules itself, and 18 of 22 runs looped). The quick screen's verdict is the sum of its dimension scores, not the model's own total. - Deterministic finance stays deterministic. The three-statement model, DCF, reverse DCF, quant fingerprint and outcome maths are plain Python with tests. The model proposes forecast drivers; the engine balances the statements (no plug) and values them.
- Structured outputs everywhere, and no invented defaults. Every JSON-producing call goes through one function using the API's native structured outputs. Truncation and refusals raise instead of being parsed around; an over-long list is clamped to its bound because the API treats length limits as hints; nothing is filled in when output is missing.
- An explicit state machine, not a framework. LangGraph was compiled but never invoked; I removed it (ADR-0004), and a follow-up phase that ran every time but never acted — 0 of 273 priority-1 questions ever qualified (ADR-0005).
- Model tiers by job. Opus 5.5 at medium effort for synthesis; Haiku 4.5 for screening, classification and extraction.
- Caching designed for fan-out. The nine deep-dive calls share one cached prefix (system prompt + data). Parallel requests can't read an entry still being written, so the first category starts alone and the other eight launch when its response begins streaming.
- Citations as a data type. Every data-client method returns
(data, Citation); the report renders them next to the claims they support.
Evals (backend/evals)
The thesis step is evaluated on frozen inputs so a prompt change can be compared like for like.
| Check | Result |
|---|---|
| Calls across 10 tickers × 3 reruns | first run: 29 avoid, 1 short, 0 long, conviction 55 in 21 of 30 — the step didn't discriminate. After the fix: 5 long, 15 avoid, 10 short; 9 distinct convictions |
| Numbers in a thesis found in the raw FMP/FRED data | 68% median. Measured against the prompt it's 100% — inflated, because the prompt is mostly the model's own deep-dive text |
| Stance vs targets vs price; kill criteria and pre-mortem present | all 30 pass (9 of 21 older stored theses lacked kill criteria) |
| Thesis evidence carrying a source tag | 0% |
| Rubric judge (Claude Sonnet 5, a different model from the writer) | ~4.1–4.9 on both runs — it can't tell the collapsed run from the fixed one, so it needs calibrating |
The first full run ($5.08) found a real product weakness, not a pass. Three causes: the schema had the model write its conviction before its stance and targets, so it committed to a middle score first; the prompt told it to anchor targets on the current price; and it defined avoid as "no edge either way". The fix reorders the schema (valuation → targets → stance → conviction), makes the model value the business before looking at the price, and ties the stance to its own targets. The rerun ($5.64) moved the median gap between base target and price from 4% to 10% and the calls from 0/29/1 to 5/15/10 long/avoid/short. Conviction still never exceeds 55 — the next target (details).
| Phase | Calls | Cost |
|---|---|---|
| Quick screen (Haiku) | 1 | $0.01 |
| Deep dive incl. transcript analysis (two loop-backs) | 18 | $1.41 |
| Thesis (×3) | 3 | $0.47 |
| Risk stress test (×3) | 3 | $0.17 |
| One VRT run | 25 | $2.07 · 6.6 min |
Restructuring the prompts for caching and reusing transcript analysis across loops took the same run from $2.92 and 38 calls to $2.07 and 25 (cache hits 6.6% → 28.5% of prompt tokens). About 70% of what remains is output (thinking included), so effort level is the next lever, not caching.
- Backend: ruff and 1,100 unit tests; migrations must apply to an empty Postgres and match the models.
- An end-to-end research run on that database, with recorded FMP responses and a fake model — every phase, a loop-back, the cache fan-out and the telemetry — with live HTTP blocked (1 second).
- Every model call site must have a probe in the live smoke test, which runs nightly (19 paths, $0.08).
- Frontend: types, lint, 72 unit tests, and Playwright on every page plus a finished report: no console errors, no serious or critical axe findings, no exceptions list.
Each finished call is recorded at the next trading day's close and snapshotted at 1 day to 6 months, scored in the direction of the call against four benchmarks. With ten live calls so far, the honest reading is that there's no detectable edge — the early outperformance is mostly theme beta, and the three-month numbers are worse than the one-month ones. The point is the harness, not the returns: from here on every call is recorded automatically before its outcome is known.
I directed the work and Claude Code wrote most of it: written specs and implementation plans before building (34 and 39 so far), ADRs for the decisions that matter, a review pass before merging, and periodic audits of my own AI-written code. The biggest lesson came from one of those audits: a fix that lived only in a commit message (assistant prefill breaks on newer Claude models) came back at five call sites. It's now a guard test and a nightly live smoke. PROCESS.md is the longer account.
Needs Python 3.12, Node 24, Postgres, and API keys for Anthropic and FMP (FRED optional).
# .env at the repo root: ANTHROPIC_API_KEY, FMP_API_KEY, X_BEARER_TOKEN, DATABASE_URL, DATABASE_URL_SYNC
python -m venv backend/venv && source backend/venv/bin/activate
pip install -r backend/requirements.txt
(cd backend && alembic upgrade head)
uvicorn backend.app.main:app --reload # from the repo root
cd frontend && npm install && npm run dev # http://localhost:3000Desktop-first, single user, no auth — a local tool. Tests: see CLAUDE.md.
| Where | What |
|---|---|
backend/app/graph/ |
pipeline phases, prompts, routing, llm.py |
backend/app/services/ |
filings, model engine, outcomes, telemetry, schedulers |
backend/evals/ |
eval harness |
frontend/app/, frontend/components/ |
Next.js pages and components |
docs/adr/ |
architecture decisions |
| docs/ARCHITECTURE.md | full reference |
MIT licensed — see LICENSE.



