Every Eval Ever is a unified open data format and public dataset for AI evaluation results. Developed by the EvalEval Coalition, it provides one standardized schema to collect evaluation outputs from disparate sources.
The initiative addresses fragmentation in AI evaluation where results are siloed by framework. It creates a common interchange format so outputs from HELM, EleutherAI, Inspect, and custom scripts can coexist and be compared directly. The schema captures full experimental context including prompt templates, inference parameters, and system states to make each result traceable, transparent, and reproducible. This supports rigorous research, meta-analysis, and the construction of automated leaderboards by turning scattered metrics into a structured, queryable global dataset.
Version 0.2.2 of the schema is documented with aggregate and instance examples. A granular line-by-line breakdown is available along with the dataset on GitHub. The project incorporates feedback from researchers at Inspect and was launched as an open effort to enable trust and comparability across the ecosystem.
Every Eval Ever sits in PulseGate's LLM eval & observability category. Fragmented and incomparable AI evaluation results scattered across different frameworks and formats. It is built as an open-source project for AI researchers and developers. The project is open source (MIT). It ships for the web, the command line, and API.
EvalEval Coalition builds and maintains Every Eval Ever, and it first shipped in 2025. The project is developed in the open on GitHub with 94 stars and 79 commits in the last 90 days. Among its 4 catalogued features are unified evaluation schema, standardized dataset, and provenance tracking. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match