evalint audits LLM evaluation sets by treating them like academic exams. It analyzes what the set can actually measure, identifies ineffective items, and evaluates overall test quality using psychometrics and reliability metrics. Available as a MIT-licensed Python package.
evalint sits in PulseGate's LLM eval & observability category. It focuses on identifying flaws, dead weight, and measurement gaps in LLM evaluation datasets before deployment. evalint is an open-source project aimed at developers. The project is open source (MIT). evalint is available on the command line.
Behind evalint is CAOShurong, and it first shipped in 2026. Development happens publicly on GitHub with 13 commits in the last 90 days. Key capabilities include Item Analysis, Reliability Metrics, and Test Quality Audit.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do