judgegate is an open-source CLI tool for evaluating LLM judges in continuous integration workflows. It provides agreement statistics against human labels, bias probes, and a label budget calculator, supporting robust AI evaluation and monitoring for developers and researchers.
judgegate is a LLM evaluation & benchmarks project. It helps developers evaluate and monitor the reliability and bias of LLM judges in CI pipelines. judgegate is an open-source project aimed at AI evaluation engineers. The project is open source (MIT). It runs on the command line, and it can be self-hosted.
It is developed by yashchimata, and it first shipped in 2026. The project is developed in the open on GitHub with 1 commit in the last 90 days. Among its 5 catalogued features are agreement statistics, bias probes, and label budget calculator.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do