Skip to content
Alternatives
Software like AgentEval
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- agent-evalpypi.orgAgent evaluation toolkit
- AgentEvalsaevals.aiAgentEvals is an open-source tool designed to evaluate and score the behavior of AI agents using telemetry data captured from real production or test environments. By analyzing OpenTelemetry Protocol (OTLP) streams and Jaeger JSON traces, it enables users to assess agent performance and inference quality without the need to rerun or replay expensive large language model (LLM) calls. This approach allows for benchmarking agents before deployment and provides insights based on actual agent traces rather than synthetic replays. The platform offers several evaluation features, including the ability to define golden evaluation sets that describe expected agent behaviors, tool calls, and trajectories. AgentEvals supports flexible trajectory matching with strict, unordered, subset, or superset modes, enabling nuanced comparisons between expected and observed agent actions. Users can also create custom evaluators in Python, JavaScript, or any language of their choice and share them through a community registry. AgentEvals is accessible through both a command-line interface (CLI) and a web user interface (Web UI). The CLI is tailored for automation and integration into CI/CD pipelines, enabling teams to gate deployments based on agent behavior quality scores. The Web UI provides interactive capabilities for visually inspecting traces, browsing results, comparing runs, and drilling into detailed evaluations. Installation is available via Python wheel, and evaluations can be run directly against trace files. 0 license. Its focus on trace-driven evaluation and support for both automated and interactive workflows make it suitable for developers and teams seeking to ensure the reliability and quality of AI agent behavior before production deployment.
- evalitepypi.orgevalite is a lightweight, model-agnostic framework for evaluating AI agents. It provides developers with tools for defining evaluation tests and measuring agent behavior across models.
- openagent-evalpypi.orgopenagent-eval is an open-source command-line framework designed for evaluating Retrieval-Augmented Generation (RAG) systems and AI agents. It provides tools and metrics for assessing LLM-based workflows, making it useful for AI researchers and developers who need to benchmark and analyze agent performance.
- agentaudit-evalpypi.orgagentaudit-eval is an open-source evaluation framework for multi-agent AI workflows. It provides tools for handoff quality scoring, failure attribution, loop detection, and cost guardrails, helping developers monitor and improve the reliability and efficiency of complex AI agent systems.
- Evalgentevalgent.comEvalgent is an AI voice agents testing and evaluation platform. It is built for voice AI teams that want to validate agents before going live and improve them across development cycles. The site also says it is meant to help teams ship production-ready agents faster and with confidence. Its main workflow centers on scenario-driven testing. Users define real-world test conversations, configure caller personas and behaviors, and run automated evaluation batches. The platform describes behavioral testing with human interaction profiles such as interruptions, background noise, fast speech, and an impatient user. It also includes limit testing, which pushes behavior and conditions until reliability drops below acceptable thresholds. Evalgent says each test can be run multiple times to measure consistency rather than luck. Results are shown in reviews where users can inspect and analyze evaluation outcomes. The page presents custom metrics, including task completion, tone adherence, hallucination detection, policy compliance, average duration, and average turns. It also says every success or failure is evidence-backed and auditable, with conversation transcripts and step-by-step failure analysis. A sample interface shown on the page includes scenarios, profiles, metrics, runs, and reviews. The product also mentions a planned monitoring feature for tracking production quality continuously. Evalgent lists use cases for healthcare, financial services, insurance, e-commerce, hospitality, logistics, automotive, real estate, recruiting, and education, with examples such as bookings, payments, claims, orders, dispatch, service, screening, admissions, and student support. It is delivered through a platform on evalgent.com, and the page includes a contact option and a schedule-a-call link. The company states that it is a pre-deployment testing layer and foundational infrastructure for shipping reliable voice agents.
- agent-genesisagent-genesis-ai.comagent-genesis is an open-source SDK and CLI/API toolkit for evaluating and testing AI agents. It provides developers with tools to benchmark, analyze, and improve agent performance during the development lifecycle.
- agent-examreadthedocs.ioagent-exam is an Apache-2.0 evaluation framework for testing agent skills across Claude Code, Codex CLI, Copilot CLI, and OpenCode. It is intended for developers building and benchmarking agent workflows.
- agent-skill-evalpypi.orgagent-skill-eval is an open-source CLI framework for evaluating the skills of code-generating agents across models like OpenCode, Claude Code, and Codex. It enables researchers and developers to benchmark agent performance using standardized tests.
- agenteval-debuggerpypi.orgagenteval-debugger is an MIT-licensed Python package for debugging AI agent executions. It provides tracing and observability capabilities for learners and developers working with agent workflows such as LangGraph.
- robotframework-agentevalpypi.orgrobotframework-agenteval is an open-source library for Robot Framework that enables automated testing of agentic stacks, including MCP servers, Agent Skills, SubAgents, and Hooks. It supports deterministic, LLM-judged, and coding agent-based evaluation, making it suitable for developers and QA engineers working with AI agents.
- AI Evaluatoraievaluator.devAI Evaluator is an MIT-licensed command-line tool for evaluating LLM agents. It supports agent testing and evaluation workflows that can run locally or in CI/CD environments, targeting developers building AI applications.
- coder-evalpypi.orgcoder-eval is an open-source command-line tool for evaluating, benchmarking, and A/B testing AI coding agents. It uses sandboxed, reproducible YAML task suites, making it suitable for AI researchers and developers assessing agent performance.
- agentanvilgithub.comContract-based testing framework for LLM agents — hybrid metrics (objective + LLM-as-judge + human), multi-agent and A2A protocol support, deterministic record/replay envelope.
- agentveritypypi.orgAgentverity is a Python package that implements decision stability and coverage checks for testing AI agents. It addresses challenges in testing non-deterministic LLM-based agents by providing metrics for verdict stochasticity, metamorphic testing, and overall suite quality. It is designed for developers building and validating reliable autonomous AI systems.
- ai-eval-forgegithub.comZero-dependency eval harness for LLM and agent regression testing. Scores outputs with exact, contains, regex, JSON, citation, and token-F1 checks. Compares two runs to flag regressions.
- agentbench-clipypi.orgagentbench-cli is an open-source command-line tool that allows AI developers and researchers to evaluate, test, and scan the behavior and safety of AI agents. It provides automated checks and reporting for agent performance and compliance.
- EvalSurferpypi.orgEvalSurfer is an open-source CLI tool that provides a skill-first, agent-native evaluation protocol for AI applications. It supports operational metrics and integrates with the Model Context Protocol (MCP), making it suitable for developers and researchers who need robust evaluation of LLMs and agentic systems.
Ranked by how close each one sits to AgentEval in the index, not by popularity. Back to AgentEval →