Independent LLM-evaluation researcher. Sydney, Australia.
My work audits large language models the way clinical psychology audits human subjects, through validated measurement instruments, population ground truth, effect-size estimation, and fully released receipts rather than binary pass/fail benchmarks. The recurring finding across these studies is consistent: aggregate metrics conceal fine-grained distortion. A bias will show its direction plainly while its distribution is lost entirely. Each study runs solo and unfunded, without institutional compute or supervision, and this is deliberate. Independent evaluation grounded in its own data is the argument, and every repository here is built to let you check the numbers yourself.
Selected work
- Plausible Patients, Impossible Populations: language models simulate psychiatric patients who are individually plausible but form populations that diverge from epidemiological ground truth.
- The Granularity Gap: an LLM judge's own pass/fail verdict explains under a third of the variance in the severity scores that same judge assigns.
- Access, Not Capability (with R. H. Tai): a frontier API answering questions on your own documents mostly sells retrieval access, not a capability that a small local model lacks.
Papers on arXiv
- arXiv:2606.05183 The Granularity Gap
- arXiv:2604.17359 Plausible Patients, Impossible Populations
Research repositories
| Repository | What it holds |
|---|---|
| The-Granularity-Gap | Data, code and verification gates for a continuous-scale audit of sycophancy across 8 Gemini variants and 8,830 responses, with 10,792 per-vote judge logs. |
| plausible-patients | Four cost-tier models generate 28,800 simulated psychiatric patients across 120 demographic cohorts, scored against survey-weighted NHANES population norms. |
| psych_scope | A counterfactual audit of socioeconomic-cue bias in LLM mental-health assessment across three model families, with a sparse-autoencoder test of whether interpretability can explain it. |
| deeper-than-the-guardrails | A ten-pair base/abliterated audit of demographic bias in clinical screening, 95,920 observations graded against author-derived NHANES 2005-2018 norms. |
| telling-more-than-they-can-know | A four-model factorial audit of whether model self-explanations reveal or conceal the demographic drivers of their psychiatric-instrument scoring, across 18 intersectional cohorts. |
| access-not-capability | A controlled audit of what retrieval buys local models against frontier APIs on private-corpus QA, judged blind in both presentation orders under family-wise error control. |
| the-listening-gap | Six LLMs translate semantic audio descriptors into parametric EQ curves, graded against SAFE-DB settings from real audio engineers. No LLM judge. |
| leave-the-image-alone | 30 CORD-v2 receipts across 4 preprocessing conditions, 4 vision models and 3 repeats, 1,440 transcriptions. Raw input beats binarize, upscale and downscale. |
Systems
Mimesis_Voice_Clone is an offline MCP server for drafting in a target author's voice, using SQLite FTS5 hybrid search, stylometric profiling and local ONNX embeddings. claude_mind is a self-maintaining personal memory system: vector and graph indexes over an Obsidian vault, kept current by autonomous agent loops.
The full index of studies, findings, and receipts is in the research collection.
I also write fiction and philosophical essays, collected at my portfolio.
Open to research collaboration and funded work in LLM evaluation and safety. · pskeough@gmail.com
