Skip to content
@evaleval

EvalEval

We’re building a research coalition on evaluating evaluations (EvalEval)! Hosted by Hugging Face, University of Edinburgh, and EleutherAI.

Popular repositories Loading

  1. every_eval_ever every_eval_ever Public

    Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to loca…

    Python 126 50

  2. benchmark-saturation benchmark-saturation Public

    Repository for benchmark saturation research project.

    Python 5 3

  3. evalHub evalHub Public archive

    Python 3 4

  4. auto-benchmarkcard auto-benchmarkcard Public

    Automated generation of BenchmarkMetadataCards for AI evaluation benchmarks

    Python 3

  5. oldsite oldsite Public

    Forked from jeromelachaud/freelancer-theme

    Jekyll theme based on Freelancer Start Bootstrap theme

    JavaScript 2

  6. eval_cards_backend_pipeline eval_cards_backend_pipeline Public

    Data processing pipeline for Eval Cards data

    Python 2 3

Repositories

Showing 10 of 15 repositories
  • eval-card-registry Public

    Registry for eval-related entities

    evaleval/eval-card-registry's past year of commit activity
    Python 1 3 4 1 Updated Sep 23, 2026
  • evaleval/evaleval.github.io's past year of commit activity
    HTML 1 8 0 0 Updated Sep 22, 2026
  • eval-cards Public
    evaleval/eval-cards's past year of commit activity
    TypeScript 0 1 2 1 Updated Sep 21, 2026
  • eval_cards_backend_pipeline Public

    Data processing pipeline for Eval Cards data

    evaleval/eval_cards_backend_pipeline's past year of commit activity
    Python 2 3 9 (1 issue needs help) 1 Updated Sep 21, 2026
  • every_eval_ever Public

    Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused.

    evaleval/every_eval_ever's past year of commit activity
    Python 126 MIT 50 12 (2 issues need help) 17 Updated Sep 20, 2026
  • model-cards Public
    evaleval/model-cards's past year of commit activity
    Python 1 MIT 0 0 0 Updated Sep 9, 2026
  • evaleval/proceedings-to-paper's past year of commit activity
    Python 1 MIT 0 0 0 Updated Sep 9, 2026
  • auto-benchmarkcard Public

    Automated generation of BenchmarkMetadataCards for AI evaluation benchmarks

    evaleval/auto-benchmarkcard's past year of commit activity
    Python 3 MIT 0 2 0 Updated Sep 9, 2026
  • eee-registry Public
    evaleval/eee-registry's past year of commit activity
    0 0 0 0 Updated Jun 27, 2026
  • benchmark-saturation Public

    Repository for benchmark saturation research project.

    evaleval/benchmark-saturation's past year of commit activity
    Python 5 MIT 3 4 (2 issues need help) 1 Updated Jun 26, 2026

Most used topics

Loading…