Skip to content
View Lostmanu's full-sized avatar

Block or report Lostmanu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Lostmanu/README.md

Hi, I'm Manuel

Electronics and control engineer · Spain
I measure AI evaluation suites against the papers they implement, and send the fixes.


I read evaluation code next to the benchmark it claims to implement. When a scorer, a sample id or a declared count doesn't match, I reproduce the gap with the real code, measure both directions (what it wrongly accepts and what it wrongly rejects) and open the fix upstream.

Open-source fixes

Seven merged across five organisations, five more open.

project what was wrong status
inspect_evals (UK AI Security Institute) worldsense never evaluated a single FALSE or IMPOSSIBLE answer merged
inspect_evals (UK AI Security Institute) apps declared 5,000 samples and loaded 3,000 merged
inspect_evals (UK AI Security Institute) bbh declared 250 samples and loads 6,509 merged
gift-eval (Salesforce AI Research) two model-name mismatches hid iTransformer and split Zeus on the leaderboard merged
evalscope (ModelScope) BBH dyck_languages lost its brackets, so correct answers scored as wrong merged
kornia (computer vision library) a module-level cache kept tensors made under inference_mode or torch.export, which broke later normal calls merged
mteb (Massive Text Embedding Benchmark) NeuCLIR2023Retrieval ran on the 2022 queries and relevance judgements merged
inspect_evals (UK AI Security Institute) bbh scored answers by suffix match open
lm-evaluation-harness (EleutherAI) BBH few-shot answers of (Q) could never score open
AIOpsLab (Microsoft) keyword submit() calls were graded as an empty answer open
physics-IQ (Google DeepMind) 40 test videos reuse their first recording, which was not documented; the PR adds a README note (see #84) open
RULER (NVIDIA) five effective-length entries in the README results table did not follow the table's own threshold rule open

"Minimal, root-cause fix with well-scoped tests." (evalscope maintainer, reviewing #1773)

Selected work

  • Ninety-Three Wrong Claims: a research programme on crypto perpetual-futures microstructure that found no edge, published with the register of the 93 claims it got wrong and who or what caught each one.
  • quant-system: the laboratory behind it, with 476 tests and a mutation harness that proves each guard actually fails when it should.
  • ifs-aifs-siar: ECMWF IFS and AIFS solar radiation forecasts against 34 SiAR ground stations in Spain.

How I work

I preregister what counts as success before the data exists, and I build controls that have to prove they bite. I also keep a public record of every claim of mine that turned out to be false.


Languages

Python C C++ TypeScript JavaScript SQL MATLAB Simulink PLC (IEC 61131-3 Structured Text) HTML5 CSS

Tools

pytest NumPy pandas React Supabase SQLite Git GitHub Actions Linux Vercel


Get in touch

LinkedIn Gmail

Popular repositories Loading

  1. ninety-three-wrong-claims ninety-three-wrong-claims Public

    93 claims a one-person research programme made and later found false, with who or what caught each one. The register, the paper, and the tool that counts it.

    Python 1

  2. Lostmanu Lostmanu Public

    Profile

    1

  3. ifs-aifs-siar ifs-aifs-siar Public

    Radiación solar: IFS y AIFS frente a 34 estaciones SiAR. Resultados, sensibilidad de escala y reproducción sin red.

    Python 1

  4. quant-system quant-system Public

    A closed one-person research lab on crypto perpetual-futures microstructure, June to September 2026. No edge found.

    Python 1

  5. RULER RULER Public

    Forked from NVIDIA/RULER

    This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?

    Python 1

  6. mteb mteb Public

    Forked from embeddings-benchmark/mteb

    MTEB: State-of-the-art evaluation of embeddings across languages and modalities

    Python 1