Evaluation harness for real headless CLI agents - reproducible multi-turn scenarios, rule + LLM scoring, cross-agent comparison
agent-belt sits in PulseGate's LLM eval & observability category. It focuses on benchmarking and evaluating headless CLI agents in reproducible scenarios. It is built as an open-source project for AI agent developers and researchers. agent-belt is open source under the Apache-2.0 license. It ships for the web and the command line.
It is developed by jfrog, and it first shipped in 2026. The project is developed in the open on GitHub with 16 stars and 8 commits in the last 90 days. Among its 4 catalogued features are agent benchmarking, multi-turn scenario evaluation, and rule and LLM scoring.
What PulseGate has recorded for this listing
Closest matches by what these projects do