judge-bench is an open-source package that provides synthetic bias and calibration probes for evaluating the reliability of LLMs acting as judges. It supports researchers in testing and analyzing model bias and calibration.
In the Model leaderboards space, judge-bench takes a focused approach. It enables systematic testing of LLM-as-judge reliability by generating synthetic bias and calibration probes. judge-bench is an open-source project aimed at AI evaluation researchers and developers. The project is open source (MIT). It ships for the command line, and it can be self-hosted.
It is developed by auraoneai, and it first shipped in 2026. The project is developed in the open on GitHub with 13 commits in the last 90 days. Key capabilities include bias probes, calibration probes, and LLM-as-judge testing.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do