agentaudit-eval is an open-source evaluation framework for multi-agent AI workflows. It provides tools for handoff quality scoring, failure attribution, loop detection, and cost guardrails, helping developers monitor and improve the reliability and efficiency of complex AI agent systems.
In the LLM eval & observability space, agentaudit-eval takes a focused approach. It focuses on evaluating and monitoring the performance and reliability of multi-agent AI workflows. It is built as an open-source project for AI workflow developers and researchers. agentaudit-eval is open source under the MIT license. It runs on the command line, and it can be self-hosted.
It is developed by vineetha00, and it first shipped in 2026. Development happens publicly on GitHub with 4 commits in the last 90 days. Key capabilities include quality scoring, failure attribution, and loop detection.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do