inspect-evals is an open-source command-line tool for evaluating large language models against a collection of benchmarks. It helps AI researchers and engineers assess model performance using standardized metrics and workflows. The tool is MIT licensed and maintained by UKGovernmentBEIS.
inspect-evals sits in PulseGate's LLM evaluation & benchmarks category. It focuses on evaluating and benchmarking large language models efficiently using standardized metrics. It is built as an open-source project for AI researchers and ML engineers. The project is open source (MIT). It runs on the command line.
It is developed by UKGovernmentBEIS, and it first shipped in 2024. The project is developed in the open on GitHub with 561 stars and 296 commits in the last 90 days. Key capabilities include model evaluation, benchmarking, and CLI interface.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do