judge-bench is an open-source package that provides synthetic bias and calibration probes for evaluating the reliability of LLMs acting as judges. It supports researchers in testing and analyzing model bias and calibration.
judge-bench is a LLM eval & observability product. It enables systematic testing of LLM-as-judge reliability by generating synthetic bias and calibration probes. It is built as an open-source project for AI evaluation researchers and developers. judge-bench is open source under the MIT license. The product ships for the command line, and it can be self-hosted.
auraoneai builds and maintains judge-bench, and the product first shipped in 2026. Development happens publicly on GitHub with 13 commits in the last 90 days. Key capabilities include bias probes, calibration probes, and LLM-as-judge testing.
Latest indexed changes and source events
Other apps tracked under the same category.