Skip to content
Alternatives
Software like Evalgent
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- AgentEvalagenteval.devAI. NET ecosystem. The platform provides features such as tool usage validation, which allows users to assert on tool chains and verify that specific tools are called in the correct order with appropriate arguments. Stochastic evaluation is supported, enabling repeated runs of agent tasks to assess actual success rates and standard deviations, reflecting the non-deterministic nature of large language models. Workflow evaluation capabilities allow for the testing of multi-agent flows, including validation of executor order, edge traversal, and per-graph tool calls. Performance evaluation tools enable users to set and assert on service level agreements (SLAs) related to response times, total duration, and estimated costs. AgentEval includes model comparison functionality, letting users benchmark multiple models against defined metrics such as tool accuracy, relevance, and cost per request. The toolkit also supports recording and replaying agent interactions, which allows for consistent, repeatable evaluations without incurring additional API costs. Security evaluation is addressed through a Red Team module that tests agents against 258 attack probes across all 10 OWASP LLM Top 10 vulnerabilities, with MITRE ATLAS technique mapping. This module covers a wide range of attack types, including prompt injection, jailbreaks, PII leakage, and more, and supports both quick scans and advanced, customizable attack pipelines. Security compliance reports can be exported in PDF format. Memory evaluation is another key feature, with tools for benchmarking agent memory retention, recall depth, temporal reasoning, fact updates, cross-session persistence, and noise resistance. Results can be exported as interactive HTML reports. NET developers seeking to rigorously test, benchmark, and ensure the reliability, security, and performance of their AI agents before production use.
- Evalta AIevaltaai.comEvalta AI is a web-based tool designed for agencies to audit websites, detect SEO and technical issues, and provide step-by-step AI-guided fixes. It tracks site health, performance, and structured data, helping users resolve problems quickly and improve search visibility.
- evalitepypi.orgevalite is a lightweight, model-agnostic framework for evaluating AI agents. It provides developers with tools for defining evaluation tests and measuring agent behavior across models.
- agent-evalpypi.orgAgent evaluation toolkit
- AgentEvalsaevals.aiAgentEvals is an open-source tool designed to evaluate and score the behavior of AI agents using telemetry data captured from real production or test environments. By analyzing OpenTelemetry Protocol (OTLP) streams and Jaeger JSON traces, it enables users to assess agent performance and inference quality without the need to rerun or replay expensive large language model (LLM) calls. This approach allows for benchmarking agents before deployment and provides insights based on actual agent traces rather than synthetic replays. The platform offers several evaluation features, including the ability to define golden evaluation sets that describe expected agent behaviors, tool calls, and trajectories. AgentEvals supports flexible trajectory matching with strict, unordered, subset, or superset modes, enabling nuanced comparisons between expected and observed agent actions. Users can also create custom evaluators in Python, JavaScript, or any language of their choice and share them through a community registry. AgentEvals is accessible through both a command-line interface (CLI) and a web user interface (Web UI). The CLI is tailored for automation and integration into CI/CD pipelines, enabling teams to gate deployments based on agent behavior quality scores. The Web UI provides interactive capabilities for visually inspecting traces, browsing results, comparing runs, and drilling into detailed evaluations. Installation is available via Python wheel, and evaluations can be run directly against trace files. 0 license. Its focus on trace-driven evaluation and support for both automated and interactive workflows make it suitable for developers and teams seeking to ensure the reliability and quality of AI agent behavior before production deployment.
- openagent-evalpypi.orgopenagent-eval is an open-source command-line framework designed for evaluating Retrieval-Augmented Generation (RAG) systems and AI agents. It provides tools and metrics for assessing LLM-based workflows, making it useful for AI researchers and developers who need to benchmark and analyze agent performance.
Ranked by how close each one sits to Evalgent in the index, not by popularity. Back to Evalgent →