judge-drift-sentinel is a Python package that monitors LLM evaluation pipelines. It uses a frozen set of human-labeled anchor examples to distinguish between changes in the underlying system and drift in the LLM judge's scoring behavior. This helps AI engineering teams maintain reliable evaluation signals over time.
judge-drift-sentinel is a LLM eval & observability project. It focuses on detecting whether shifts in LLM evaluation scores are caused by changes to the system under test or by drift in the LLM judge itself. It is built as an open-source project for AI engineers and LLM application developers. judge-drift-sentinel is open source under the MIT license. It runs on the command line and API.
It is developed by Homayoun Safarpour, and it first shipped in 2026. The project is developed in the open on GitHub with 13 commits in the last 90 days. Among its 4 catalogued features are Drift Detection, Anchor Set Analysis, and LLM Judge Monitoring.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match