Electronics and control engineer · Spain
I measure AI evaluation suites against the papers they implement, and send the fixes.
I read evaluation code next to the benchmark it claims to implement. When a scorer, a sample id or a declared count doesn't match, I reproduce the gap with the real code, measure both directions (what it wrongly accepts and what it wrongly rejects) and open the fix upstream.
Seven merged across five organisations, five more open.
| project | what was wrong | status |
|---|---|---|
| inspect_evals (UK AI Security Institute) | worldsense never evaluated a single FALSE or IMPOSSIBLE answer |
merged |
| inspect_evals (UK AI Security Institute) | apps declared 5,000 samples and loaded 3,000 |
merged |
| inspect_evals (UK AI Security Institute) | bbh declared 250 samples and loads 6,509 |
merged |
| gift-eval (Salesforce AI Research) | two model-name mismatches hid iTransformer and split Zeus on the leaderboard | merged |
| evalscope (ModelScope) | BBH dyck_languages lost its brackets, so correct answers scored as wrong |
merged |
| kornia (computer vision library) | a module-level cache kept tensors made under inference_mode or torch.export, which broke later normal calls |
merged |
| mteb (Massive Text Embedding Benchmark) | NeuCLIR2023Retrieval ran on the 2022 queries and relevance judgements |
merged |
| inspect_evals (UK AI Security Institute) | bbh scored answers by suffix match |
open |
| lm-evaluation-harness (EleutherAI) | BBH few-shot answers of (Q) could never score |
open |
| AIOpsLab (Microsoft) | keyword submit() calls were graded as an empty answer |
open |
| physics-IQ (Google DeepMind) | 40 test videos reuse their first recording, which was not documented; the PR adds a README note (see #84) | open |
| RULER (NVIDIA) | five effective-length entries in the README results table did not follow the table's own threshold rule | open |
"Minimal, root-cause fix with well-scoped tests." (evalscope maintainer, reviewing #1773)
- Ninety-Three Wrong Claims: a research programme on crypto perpetual-futures microstructure that found no edge, published with the register of the 93 claims it got wrong and who or what caught each one.
- quant-system: the laboratory behind it, with 476 tests and a mutation harness that proves each guard actually fails when it should.
- ifs-aifs-siar: ECMWF IFS and AIFS solar radiation forecasts against 34 SiAR ground stations in Spain.
I preregister what counts as success before the data exists, and I build controls that have to prove they bite. I also keep a public record of every claim of mine that turned out to be false.
