journeyman-bench is an open-source benchmark for evaluating the process quality of LLM agents. It assesses agent workflows and tool use, and uses calibrated LLM-as-a-judge evaluations for researchers and developers.
In the Agent evaluation & testing space, journeyman-bench takes a focused approach. It focuses on evaluating how LLM agents work, including their tool use and decision-making process, rather than only checking final outcomes. It is built as an open-source project for AI researchers and developers evaluating LLM agents. The project is open source (Apache-2.0). It ships for the command line, and it can be self-hosted.
It is developed by codechu, and it first shipped in 2026. The project is developed in the open on GitHub with 87 commits in the last 90 days. Among its 6 catalogued features are agent benchmarking, process evaluation, and tool-use grading.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do