eval-integrity is a lightweight, dependency-free Python package that performs statistical checks on AI evaluation claims. It detects issues such as multiple-comparisons problems, judge bias, insufficient resolution, and result fragility. Available as both a CLI tool and an MCP server that agents can call before accepting benchmark numbers. Licensed under MIT with source on GitHub.
eval-integrity sits in PulseGate's LLM eval & observability category. It focuses on verifying the statistical integrity of AI benchmark and evaluation claims before trusting them. It is built as an open-source project for AI researchers and engineers. eval-integrity is open source under the MIT license. eval-integrity is available on the command line and API.
It is developed by ipezygj, and the product first shipped in 2026. Development happens publicly on GitHub with 10 commits in the last 90 days. Key capabilities include Statistical Checks, Judge Bias Detection, and Multiple Comparisons. It exposes integrations via an MCP server.
Latest indexed changes and source events
eval-integrity verified by the PulseGate indexer
Other apps tracked under the same category.