AgentEvals
PulseGate's liveness check found it on 13 Sep 2026; it is registered on GitHub and has been in the index since 26 Jun 2026. How this is checked
AgentEvals is an open-source tool designed to evaluate and score the behavior of AI agents using telemetry data captured from real production or test environments. By analyzing OpenTelemetry Protocol (OTLP) streams and Jaeger JSON traces, it enables users to assess agent performance and inference quality without the need to rerun or replay expensive large language model (LLM) calls. This approach allows for benchmarking agents before deployment and provides insights based…
Inferred · not functionally tested
Overview
6 featuresPurpose: Evaluating and benchmarking AI agent behavior from real production traces without rerunning agents.
Inferred · not functionally tested
Audience: AI developers and ML engineers
Inferred · not functionally tested
Functions: analytics, monitoring
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI, WEB · deployment: browser, cli, cloud_managed
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: aevals.ai · github.com. These links do not verify the individual claims.
AgentEvals sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on evaluating and benchmarking AI agent behavior from real production traces without rerunning agents. Inferred · not functionally tested: It is built as an open-source project for AI developers and ML engineers. Basis unknown · not verified: The project is open source (Apache-2.0). Basis unknown · not verified: AgentEvals is available on the web and the command line.
Behind AgentEvals is AgentEvals Maintainers, and it first shipped in 2026. The project is developed in the open on GitHub with 141 stars and 159 commits in the last 90 days. Inferred · not functionally tested: Among its 6 catalogued features are trace-based evaluation, LLM-powered scoring, and custom evaluators. Inferred · not functionally tested: Catalogued interfaces include a public API.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Trace-based evaluation
- LLM-powered scoring
- Custom evaluators
- CI/CD integration
- Trajectory matching
- Golden eval sets
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about AgentEvals
- What is AgentEvals?
- Inferred · not functionally tested: AgentEvals focuses on evaluating and benchmarking AI agent behavior from real production traces without rerunning agents. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is AgentEvals for?
- Inferred · not functionally tested: AgentEvals is an open-source project built for AI developers and ML engineers.
- Does AgentEvals have a free plan?
- Basis unknown · not verified: Yes — AgentEvals is open source under the Apache-2.0 license and free to use.
- What platforms does AgentEvals run on?
- Basis unknown · not verified: AgentEvals runs on the web and the command line.
- Is AgentEvals still maintained?
- PulseGate's liveness check found it on 13 Sep 2026. Its GitHub repository shows 159 commits in the last 90 days.
- What are alternatives to AgentEvals?
- Similar projects tracked by PulseGate include AgentEval, agent-eval, and Aeval.AgentEvalagent-evalAeval
- Who makes AgentEvals?
- AgentEvals is developed by AgentEvals Maintainers.
- How long has AgentEvals been around?
- AgentEvals first shipped in 2026.
Similar projects
Closest matches by what these projects do