inspect-ai is an open-source framework designed to automate and standardize the evaluation of large language models. It provides tools for running automated tests, defining custom metrics, and integrating with Python workflows. Ideal for AI researchers and developers seeking robust LLM evaluation pipelines.
inspect-ai sits in PulseGate's LLM eval & observability category. It focuses on automating and standardizing the evaluation of large language models for developers and researchers. It is built as an open-source project for AI researchers and developers. inspect-ai is open source under the MIT license. It ships for the web and the command line, and it can be self-hosted.
It is developed by UK Government Department for Business, Energy & Industrial Strategy (United Kingdom), and it first shipped in 2024. Development happens publicly on GitHub with 2.3k stars and 1.2k commits in the last 90 days. Among its 5 catalogued features are evaluation framework, automated testing, and LLM support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do