CEO, Tortoise AI | 20 years getting AI into production
I build AI for environments where failure is not an option: defence, nuclear, aviation, and now live sports broadcasting.
When the frameworks I need don't exist, I build them open source.
ARMM (Agent Readiness Maturity Model)
Assessing organisational readiness to deploy AI agents in production.
Four dimensions. Five levels. Weakest-link principle. CC BY 4.0.
Decision Latency Framework (coming 2026)
A scoring methodology for knowing how much time a decision should take.
pda-platform
Open infrastructure for AI-enabled project delivery. Universal PM data parser, MCP servers for Claude integration, AI reliability tooling. MIT.
agent-task-planning
AI reliability framework with confidence extraction and outlier mining. Multi-provider. Production guardrails. MIT.
pm-data-tools
Universal parser for project management data. 8 formats + NISTA. MIT.
ARMM Assessment Tool
Interactive self-assessment across 251 criteria. AGPL-3.0.
Universal Dashboard Specification
Vendor-neutral declarative format for AI-native analytical dashboards. Apache 2.0.
inspect_ai — the UK AI Safety Institute's LLM evaluation framework.
Merged:
- Atomic log file writes — eval logs no longer corrupt when a run fills the disk or is interrupted mid-write.
aggregate(key, agg=...)metric factory — per-key aggregation for metrics computed over grouped samples.- Distinguish
math()answer-extraction failures from wrong answers viaScore.reason— an unparseable answer and an incorrect one used to score identically. - Warn when task arguments are entirely unconsumed — surfaces silently ignored arguments instead of running an evaluation that quietly differs from the one asked for.
- Correct the documented
math()scoring outcomes — the reference and Scoring Policy pages both said a malformed answer raises a scoring error, when it returnsINCORRECTwith a reason. - Fix the MMLU CLI command on the Evals page.
inspect_evals — the companion suite of published evaluations.
Merged:
- Flag model roles that can silently fall back to the model under evaluation — a lint check for graders that, left unconfigured, end up grading their own output.
- Fix a
UnicodeDecodeErrorthat stops the lint tool running on Windows.
In review:
- Make the grader configurable via the grader model role across agieval, math and frontierscience.
inspect_scout — in-depth analysis of AI agent transcripts.
Merged:
- Collect concurrent scan reads with
tg_collect— moves the remainingasyncio.gathercall sites onto the project's structured-concurrency helper, closing a path where a cancelled read could silently truncate the results.
In review:
- Keep adaptive connections when the scan is certainly single-process —
max_connectionswas stamped even when nothing could contend for them, putting adaptive sizing out of reach. - Say what to write when a filter is given as a bare string — a string is iterable, so
messages="assistant"was read as its characters: the error listed the word's own letters as invalid while namingassistanton the same line as allowed. messages=["all"]andevents=["all"]scanned nothing instead of everything —"all"means everything only as a bare filter, so inside a list it reached selection and was matched against role and event names, matching none. It passed validation and raised nothing, so the run was indistinguishable from a correct one.
pytest — the Python testing framework.
In review:
- Do not fail on undecodable subprocess output —
Pytester.run()read captured output as strict UTF-8, so a child process writing anything else took the whole run down with it, exit code included. Fixes #7623, open since 2020.
huggingface/datasets — the dataset library behind the Hugging Face Hub.
In review:
- Close the Arrow writer when
finalize()fails — a writer left open meant Windows cleanup raised from inside afinally, replacing the real error with a file-lock one. Fixes #6917. - Do not rebuild a finalized
ArrowWriter—finalize()clearspa_writerbefore closing the stream, so a close that raises leaves the writer indistinguishable from one never built and the next call rebuilds onto a closed stream. Companion to the above, which fixes the builder call sites; this fixes the writer itself.
Verified Autonomy: A Field Guide to Engineering Trust in AI Systems (May 2026)
A field guide for engineering trust into production AI systems, covering calibration, conformal prediction, audit trails, and constrained autonomy. Ant Newman, Shanti Greene, Malia Hosseini, Philip Kitchener, Rainier Potgieter, Hadley Christoffels. Companion repo: verified-autonomy.
CC BY 4.0 (content), MIT (code).
Module-Level versus Chain-Level Stabilisation in Multiparameter Persistence (Mar 2026)
Proves that module-level and chain-level ε-pruning are not equivalent, separating Bjerkevik's pruning of persistence modules from chain-level pruning of the M-invariant of spectral systems. Ant Newman.
CC BY 4.0.
From Policy to Practice: An Open Framework for AI-Ready Project Delivery (Feb 2026)
An open framework for AI-ready project delivery. Ant Newman. Companion repo: Project-Delivery-Toolkit. Also on SSRN.
CC BY 4.0.
Agent Readiness Maturity Model (ARMM) Framework v1.1 (Jan 2026)
A maturity model that scores whether an organisation is ready to deploy AI agents in production, applying a weakest-link rule across four dimensions and five levels. Ant Newman. Interactive tool: tortoiseai.co.uk/armm.
CC BY 4.0.
The Sharon Instability Theorem: Generic Instability of the M-Invariant (Dec 2025)
Proves the M-invariant in multiparameter persistence is generically unstable, resolving an open question in the field. Ant Newman. Companion repo: sharon-instability. Also on SSRN.
CC BY 4.0.
Closing the Gap: A Practical Framework for Implementing Data Analytics and AI into the Built Environment (Jun 2025)
A green paper on embedding AI and data analytics in major government and infrastructure projects, setting out target states, 90-day starter actions and 12-month milestones. Project Data Analytics Task Force, with Ant Newman, Donnie MacNicol, Nermeen Latif, Richard Morgan and Jonah Froggatt.
CC BY 4.0.
Also indexed at ORCID · Google Scholar · ResearchGate · SSRN · Academia.
I publish on AI deployment, decision-making under complexity, and what production-grade reliability actually requires. I also write Radical Productivity, a weekly Substack on productivity systems and inclusion.
Radical Productivity · LinkedIn · tortoiseai.co.uk/insights



