Applied mathematician (PhD). I build systems that are honest about what they don't know.
I work on LLM evaluation and document AI: evaluation harnesses with statistical uncertainty, golden datasets with documented rubrics, and extraction pipelines whose accuracy is measured rather than assumed. Previously, I spent four years shipping production models at a Fortune 100 insurer, followed by a document-AI research contract. I'm available for evaluation and document-AI work.
- eval-toolkit — evaluation library on PyPI: bootstrap confidence intervals, leakage checks, versioned result schemas, CI gates.
- prompt-injection-detection-prototype — a full detection study published with its negative result: confidence intervals, baseline contamination disclosed, failure modes analyzed. Honest evaluation, demonstrated.
- ir-eval — statistical retrieval evaluation for CI/CD, with paired tests and drift detection over golden-set results.
- research-kb — a multi-thousand-source research knowledge base: PDF ingestion → hybrid retrieval (BM25 + vectors + reranking) → MCP server, with a retrieval eval suite gating the weekly rebuild.
- temporalcv — released Python package for time-series cross-validation with gap enforcement and leakage checks.
Site: brandon-behring.dev