trajectory-judge is an open-source Python library that evaluates LLM agent trajectories. It measures cases where an agent arrives at the correct answer but through flawed or unintended reasoning that standard LLM judges often miss. The package helps researchers and developers better calibrate and debug autonomous agents by providing more nuanced evaluation metrics beyond final-answer correctness.
trajectory-judge is a LLM eval & observability project. It focuses on evaluating whether LLM agents followed the correct reasoning path even when they reach the right final answer. trajectory-judge is an open-source project aimed at AI researchers and developers. The project is open source (MIT). It ships for the command line and API.
Behind trajectory-judge is Hadi Mohammadi, and it first shipped in 2026. The project is developed in the open on GitHub with 38 commits in the last 90 days. Among its 4 catalogued features are Trajectory Evaluation, LLM-as-Judge, and Calibration Metrics.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do