AI-native quantitative research platform focused on reproducibility, statistical validation, and production-grade financial data engineering. This is research infrastructure — not a trading app, and it makes no claim to generate alpha. Its value is methodological rigor: purged cross-validation, walk-forward analysis, the Probability of Backtest Overfitting (PBO), and a multiple-testing-adjusted Sharpe margin adapted from Bailey and López de Prado (2014). FINDING-007 records why that margin is not the paper's statistic; ADR-054 implements the paper's probability form beside it and corrects every surface that called the margin by the paper's name.
Full design rationale and every decision:
docs/ARCHITECTURE.md.
Most backtesting projects report the strategies they found. This one reports what its own gate does, because a graduation criterion whose error rates are unknown cannot support any claim made with it. Every number below is produced by a committed workflow and can be re-run.
| question | answer | how |
|---|---|---|
| Type-I error — how often does the whole pipeline graduate a symbol with no edge by construction? | 0/200 under both iid-normal and bootstrap nulls, with ADR-050's dispersion and judged at the hunt's own 7400-bar history (ADR-063). Max DSR -0.368 (iid) / -0.253 (bootstrap) | null-calibration.yml (ADR-036/037/050/051/063) |
| Do those false graduates survive the universe-deflation bar? | 0 of 200, in both nulls | same run |
| Power — does the gate detect a planted edge at production parity? | Yes, and it is measured: 66% at AR(1) oracle Sharpe 3.9 (33/50 also clear the deflation bar), 40% at oracle 4.0 in the reverting direction, 0% at oracle 1.3. Lengthening the searched history to 1990 raised every cell that had an edge to find — 34/22/14/64% → 40/36/24/66% — and moved none down (ADR-063) | power-calibration.yml (ADR-041/049/050/051/063) |
| Is the effect size those rates are quoted against achievable? | Not always. Charged the same 10bp turnover cost the catalog pays, the oracle at |φ| = 0.10 is +0.02 / −0.09 — the two zero-power cells contained no tradeable edge at all | power-calibration.yml (ADR-055) |
| Resolution — what must an edge actually be to be found here? | A true annualized Sharpe of 2.13 at the 607-symbol universe and the 4.3-year holdout the pool reports today, falling to 1.82 as the re-searched cohort reaches the 5.9-year holdout ADR-063 bought. The bar falls as 1/sqrt(T) in holdout length but only sqrt(ln N) in universe size, so history is the lever |
scripts/pool_report.py (ADR-043/063) |
| How many discovered strategies clear that bar today? | 0 of 41. They are forward-tested on paper, never recommended | GET /api/v1/pool-report |
| Does what the search proposes beat a no-edge surrogate? | No. In the 5,400-bar cohort (1,155 matched), walk-forward +0.532 and purged-CV +0.575 do not separate from bootstrap medians +0.652 / +0.661. In the 7,400-bar cohort (254 matched), +0.460 / +0.601 likewise do not separate from +0.617 / +0.639 | scripts/pool_report.py vs null-calibration.yml (ADR-051/064) |
| How much of that out-of-sample number is drift? | All of it, on a null. Measured per symbol on 200 nulls each: walk-forward excess is −0.011 (bootstrap) and +0.000 (iid-normal); purged-CV excess is −0.000 under both modes | ADR-068/078, null-calibration.yml |
| What does the search ADD, with drift removed from both sides? | Walk-forward adds −0.127 on 254 matched experiments (89 symbols, 7,346 bars), against −0.011 / +0.000 on the nulls. Purged CV adds −0.000 on its first 88-symbol measured cohort, against −0.000 / −0.000 | scripts/pool_report.py (ADR-068/072/078; FINDING-017) |
| Is that difference distinguishable from the null's? | Walk-forward: yes, and in the wrong direction — −0.116 [−0.174, −0.064] against bootstrap and −0.127 [−0.179, −0.072] against iid-normal. Purged CV: no — +0.000 [−0.048, +0.002] / +0.000 [−0.048, +0.000]. These symbol-clustered intervals remain lower bounds on width; ADR-081 pre-registers 400 whole-panel null replicates but has not spent that measurement | scripts/pool_report.py (ADR-075/078/081) |
| Did lengthening the searched history make selection worse? | Not decidably. Paired within symbol across ADR-063's window change, the finalist's out-of-sample Sharpe moved −0.038 [−0.060, −0.009] on 368 symbols while in-sample moved +0.012 [−0.005, +0.034], and the search picks a different strategy on 257 of 368. Re-searching 45 of them at the old window to remove the drift confound gives −0.074 [−0.157, +0.030] — the interval includes zero, so ADR-074's pre-stated criterion does not fire and the window stays | scripts/window_experiment.py (ADR-063/074) |
The last five rows are the point, and the third and fourth are the sharpest thing this project has measured. With drift removed from both sides, walk-forward selection subtracts about 0.12 Sharpe more on real symbols than on data built to contain nothing, while purged-CV excess is zero and does not differ from the null. The pipeline does not find an edge under either diagnostic; only the causal walk-forward measurement currently shows an additional real-versus-null selection cost. The retained per-symbol counters currently preserve 477,000+ candidate evaluations as the conservative cumulative DSR/MinTRL lower bound (ADR-062/066). The pipeline has graduated nothing that is distinguishable from best-of-N selection luck — and when the strategies it proposes are compared against a surrogate with no serial structure at all, they do not win either. Walk-forward and purged-CV out-of-sample Sharpes are read against their own measured null distributions (bootstrap p95 ≈ +0.98 walk-forward, +1.00 purged CV), never against zero. Because every strategy in the catalog trades serial structure, and the bootstrap null destroys serial structure while preserving SPY's return shape exactly, the search's advantage over an iid-normal null is distributional rather than predictive (ADR-051).
ADR-064 made that comparison formally valid for the first time, and it survived. The
report used to refuse all four comparisons: it required the pool's median history to equal the
null artifact's exactly, and a median that grows by one bar per trading day never does. The
comparison now runs on the subset of experiments whose history is within 10% of the null's — 2,427
of them, at a median 5,445 bars against nulls at 5,400, all carrying the null's own search
fingerprint. On that matched subset the pool's walk-forward median is +0.542 against a
bootstrap null median of +0.652 (p95 +0.983) and purged-CV +0.584 against +0.661
(p95 +1.003): the search does not separate from a no-edge surrogate, and against the bootstrap it
sits below the null's median. The pool-wide median would have read +0.567 — the subset is not a
formality. Those figures are against the 5,400-bar nulls. ADR-063's re-dispatch
briefly replaced them with 7,400-bar ones and the report went to 0 matched — correctly, since a
pool at 5,445 bars is 26% away from a 7,400-bar null. ADR-065 fixed the cause rather than the
symptom: a null artifact is now named for the history it measured (bootstrap_spy_5400.json,
bootstrap_spy_7400.json), both lengths sit on disk, and a pool in transition is read against each
cohort separately. What changed permanently is the mechanism, not the verdict.
ADR-068 then found where the level of that number comes from, and it is not the search. The walk-forward out-of-sample Sharpe is denominated in the drift of whatever series it was computed on. Measured: each null's finalist median equals its own generator's buy-and-hold Sharpe to within ±0.03 (iid-normal 0.397 → +0.414; bootstrap:SPY, whose source is SPY 1993→2026, 0.650 → +0.652), and the matched pool's 0.546 → +0.542. The −0.11 "gap" above is the gap between SPY's 33-year drift and the median pool symbol's over its own window. The null side is now measured pairwise, not estimated (run 33287465013, 200 symbols per mode at 7,400 bars): the median per-symbol excess of the finalist over holding the same series across the same windows is −0.006 against the bootstrap null and +0.000 against the iid-normal one, and the searched finalist beats simply holding only 18.5% and 36.0% of the time — it pays turnover costs for a signal that is not there, exactly as it should. Controlling for drift also makes the comparison an order of magnitude more sensitive: the bootstrap null's p95 falls from +0.968 raw to +0.096 in excess terms, so a real edge no longer has to out-run the spread of market drift to be visible.
The real side of that row measured for the first time on 2026-08-31, and it is negative. Across the 77 experiments (66 symbols, 7,345 bars) that now carry the paired benchmark and match the 7,400-bar nulls, the median excess is −0.125, negative in 75.3% of them, against −0.006 and +0.000 on the two nulls. What the search adds out-of-sample over holding the same series across the same windows is less than what it adds on data with no edge by construction. The single-draw verdict rule does not change on it: −0.125 sits at the bootstrap null's 11.5th percentile, inside its p5 of −0.233, so that row reads does not separate — and ADR-072 made the band two-sided so a reader can see which side of it the number is on.
ADR-075 then fixed the comparison itself, and the answer flipped. The band above is over individual null draws: it answers "could one symbol look like this", and one easily could. The question the row is actually asked is whether the pool's centre differs from the null's, and that needs an interval on the difference of medians — resampled with the real side clustered by symbol, because the 77 excesses come from 66 symbols and are not independent draws. Measured: −0.119 [−0.215, −0.061] against the bootstrap null and −0.125 [−0.218, −0.063] against the iid one, over 66 clusters. Both exclude zero. So the honest headline is not "the search does not separate from a no-edge surrogate" but "the search separates from a no-edge surrogate in the wrong direction" — the strongest evidence here that the in-sample argmax fits structure that does not persist. Three qualifications ship with that sentence and are printed beside it: the interval is a lower bound on its own width (the symbols share one calendar window, so cross-sectional correlation survives — FINDING-012 requires independent correlated-panel null replicates, because one correlated panel followed by symbol-level resampling would repeat the error); the single-draw verdict is reported unchanged rather than replaced; and ADR-075 discloses that the point estimate was known before the scheme was fixed, so what was pre-registered is the procedure, not the answer.
That is a claim about this universe and this catalog jointly, not about the thresholds. The power row is what makes it sayable: a gate that detects nothing could not distinguish "no edge here" from "no ability to see one", and this one detects a planted edge 66% of the time and clears its own deflation bar 33 times in 50.
Where it still fails is capture, and ADR-055 sharpened that considerably by charging the oracle the same transaction costs every catalog strategy pays. Measured net of costs, the catalog's best in-sample config beats a cost-paying oracle on AR(1) processes (103–128%) and reaches only 30–42% on band reversion at half-lives 1–5, where it detects 0%. The fast band cells have a higher net oracle than the AR(1) cell detected 24% of the time (+1.70 vs +1.15), and capture rises monotonically with the horizon (31% → 50%). Those ratios all fell when ADR-063 lengthened the searched history, which is a correction rather than a regression: capture's numerator is an in-sample maximum over the searched grid, so more in-sample data regresses it toward the value it estimates and the older, higher ratios carried more selection bias.
ADR-056/057/058 then took that finding apart. A strategy was added specifically to express fast reversion to a slow-moving level, the calibration was re-run, and net capture moved by at most 0.7pp — so the record now says which strategy won each of the 50 searches per cell. At half-life 1 the search picks a Trend strategy 68% of the time on a process that is by construction fast reversion, and capture tracks that recognition share almost exactly (18% reverting finalists → 32% capture; 94% → 45–56%). The gap is recognition, not expression: at fast half-lives no reverting strategy wins the in-sample comparison, so the search never gets as far as choosing between them. Splitting capture by category (ADR-059) makes the size of that plain — at half-life 1 the headline 32% is carried by trend strategies fitting the level, while the reverting strategies keep 22%.
Then ADR-061 asked what was recoverable at all. The planted process is a random-walk level plus a fast deviation, and only their sum is observable, so the oracle every one of those ratios divided by knows a state no strategy can see. A Kalman filter given the true process parameters — an upper bound on any causal price-based strategy — nets −0.08 Sharpe at half-life 1 against that oracle's +1.70, and +0.95 at half-life 5. Measured against what is actually recoverable, the catalog converts 88–98% of it at half-lives 5–20 and there is nothing to convert at half-lives 1–2. The gap was the benchmark, not the catalog — and the zero detection rate follows from the detectable-edge frontier alone: a recoverable Sharpe of at most ~0.95 against a requirement of ~1.50 even after ADR-063 lengthened the holdout (it was ~1.76 before). A gate that graduated any of those cells would have been wrong. The added strategy was removed once it failed its own pre-stated criterion — that loop, stating a criterion before the measurement and honouring it afterwards, is the point of the project.
The same message shows up on real data, in a different place. ADR-060 records, for every experiment in the pool, the best in-sample Sharpe achieved by each catalog category, so the pool report can ask whether the family the search selects is separable from the one it passed over. Across 3,255 experiments the medians are Trend +0.569 (wins 53% of searches), Breakout +0.495 (10%), Mean Reversion +0.469 (33%), Combination +0.316 (3%) — and the median lead of the winning category over the runner-up is +0.074 against a Lo (2002) Sharpe standard error of 0.215. For a typical symbol the kind of strategy the search picks is inside a third of one standard error of the kind it rejected. That is not a defect in the selection rule; it is the same statement ADR-061 makes on synthetic data — at this history length the data does not contain enough information to separate these hypotheses, which is why the honest output of the whole pipeline is still zero graduates.
End-to-end, all gates green (backend 98.02% coverage, 1,809 tests; frontend 97.57%, 349 tests):
- 18 HTTP endpoints: health, strategy catalog (single source of truth per ADR-010), ingest,
bars, backtest, validate, Monte Carlo, plus the research-lab surface — leaderboard, graduates,
pool report, paper portfolio, equity curve, cross-sectional factors, null calibration, null
comparison, window comparison, window experiment, power calibration. Sync
defper ADR-009, so blocking yfinance + DB calls go through FastAPI's threadpool. - 7 product pages: Validation Report (full statistical suite + plain-English verdicts), Data Explorer, Backtest Results (equity curve with buy-and-hold overlay, underwater drawdown, rolling Sharpe, return distribution), Compare Configs, Lab dashboard (the deflation headline, the measured gate calibration, the paper book, cross-sectional factors), Discoveries, About.
- Data layer: PriceBar / FundamentalData / quality models; yfinance adapter + OHLCV normalizer with split/dividend adjustment; 6-active-check DataQualityEngine (honest "flags potential X" wording — never "guarantees"); SEC EDGAR fundamentals; sync TimescaleDB repository on psycopg3 with Alembic migration (hypertable + index), Docker-gated integration tests.
- Research engine: vectorized pandas/numpy backtester (ADR-007 — vectorbt rejected: fails on
Python 3.12); 34 single-name strategies and 13 cross-sectional factors (16 when
fundamentals scores are available), each with its paper citation in
.claude/context/research-papers.md; adding a strategy is a single backend diff (ADR-010); benchmark comparator; Monte Carlo simulator; experiment manifest. - Validation engine: PBO via CSCV (Bailey 2015), a multiple-testing-adjusted Sharpe margin,
scored walk-forward with Pardo efficiency (ADR-038) and scored purged K-fold CV whose
embargo is sized from the grid's longest lookback (ADR-039), parameter stability, regime
analysis, universe-level deflation (ADR-018). Every financial-math invariant
(
docs/ARCHITECTURE.md§8) is a Hypothesis property test. - Autonomous research loop: sharded daily discovery, weekly cross-sectional and fundamental sweeps, and paper forward-testing run as scheduled GitHub Actions with no human in the loop; graduates are frozen into a paper book (ADR-019 — paper only, never real money) and retired on measured decay.
The engine is calibrated to be honest: a random walk yields PBO ≈ 0.9 and does not pass.
- Backend: Python 3.12, FastAPI, Pydantic v2, SQLAlchemy 2.0 (sync, psycopg3), TimescaleDB, Alembic
- Research: NumPy, SciPy, Pandas — vectorized backtesting on pandas/numpy (ADR-007)
- Frontend: React 19 + TypeScript strict, Vite, Tanstack Query 5, Zustand 5, Recharts 3, Zod 4
- Testing: pytest + Hypothesis (backend); Vitest + React Testing Library + MSW (frontend); coverage gates 85% backend / 75% frontend, currently 98.02% / 97.57%
- Tooling: uv (Python env), ruff (lint + format), mypy (strict), pre-commit, GitHub Actions CI (backend + frontend + pre-commit, gating every commit)
backend/ FastAPI app, data layer, research engine, validation engine
frontend/ React dashboard (Vite + TS strict + Vitest)
docs/ ARCHITECTURE.md, ADRs (ADR-001..043), C4 diagrams
.claude/ Codified context: constitution (CLAUDE.md), domain agents,
cold-memory docs, playbooks — drives AI-assisted sessions
# One-time
make dev # start docker-compose (TimescaleDB + Redis + backend)
make migrate # apply Alembic migrations
# Per-commit gates
make check # backend: ruff + format-check + mypy + pytest + coverage
make frontend-check # frontend: eslint + tsc + vitest + coverage
make check-all # both, before pushing
# Run the UI locally
cd backend && uv run uvicorn app.main:app --reload # one terminal
cd frontend && npm run dev # the other; Vite proxies /apiCI gates on deterministic synthetic fixtures only. Live-data tests (@pytest.mark.live)
and Docker-gated integration tests (@pytest.mark.integration) run locally via
make test-live / make test-integration.
- Validation-first (ADR-008): every Sharpe is deflated, every report carries its PBO, walk-forward, and parameter stability. A "good" Sharpe with PBO ≥ 0.5 fails.
- Honest data quality (CLAUDE.md rule 6): quality-check messages say "flags potential X" — never "prevents" or "guarantees." A gate informs review; it does not certify correctness.
- Sync DB stack (ADR-009): SQLAlchemy 2.0 sync on psycopg3, FastAPI routes are sync
defand threadpooled by the framework. Researched and ratified 2026-05-28. - Cache-aside read path:
/validateand/backtestread bars from the repository first; on miss they run the ingestion pipeline (quality-gated) and re-read. TimescaleDB is the cache today; Redis is wired in config for a future hot path. - Codified context (Vasilopoulos 2026 arXiv:2602.20478, validated across 283 dev
sessions): three-tier — always-loaded constitution (
CLAUDE.md), domain-expert agents (.claude/agents/), on-demand cold memory (.claude/context/). The repo is built to be picked up by a fresh Claude session and continued without losing rigor.
- Phase 1 (foundation) — done
- Phase 2 (data engineering) — done; TimescaleDB repo + Alembic migration built and
integration-tested (
make test-integration) - Phase 3 (research engine) — done; oracle tests pass on every invariant
- Phase 4 (validation engine) — done; ValidationReport is the MVP deliverable
- Phase 5 (product surface) — done; seven pages shipped end-to-end
- Phase 6 (autonomous research) — running; scheduled discovery, forward testing, and the gate-calibration measurements above. The open question is not "does it find strategies" but "is anything it finds distinguishable from luck" — and the answer is currently, honestly, no.