Applied mathematician (PhD). I build systems that are honest about what they don't know.
I work on LLM evaluation: evaluation harnesses with statistical uncertainty, golden datasets with documented rubrics, and pipelines whose accuracy is measured rather than assumed. Previously, I spent four years shipping production models at a Fortune 100 insurer, followed by a short ML research contract in document AI over the summer, which has ended. I'm now looking for statistical modeling work in insurance.
- eval-toolkit — evaluation library on PyPI: bootstrap confidence intervals, leakage checks, versioned result schemas, CI gates.
- prompt-injection-detection-prototype — a full detection study published with its negative result: confidence intervals, baseline contamination disclosed, failure modes analyzed. Honest evaluation, demonstrated.
- ir-eval — statistical retrieval evaluation for CI/CD, with paired tests and drift detection over golden-set results.
- temporalcv — released Python package for time-series cross-validation with gap enforcement and leakage checks.
Site: brandon-behring.dev



