PulseGateCategoriesMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets90 categories · 139 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesIndustry UpdatesMethodologySupply IndexData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGateAGAgentEvals
Visit↗
Skip to content
  1. Index›
  2. LLM eval & observability›
  3. AgentEvals
← Back to the index
AG

AgentEvals

aevals.ai·Infrastructure

AgentEvals is an open-source tool designed to evaluate and score the behavior of AI agents using telemetry data captured from real production or test environments. By analyzing OpenTelemetry Protocol (OTLP) streams and Jaeger JSON traces, it enables users to assess agent performance and inference quality without the need to rerun or replay expensive large language model (LLM) calls. This approach allows for benchmarking agents before deployment and provides insights based on actual agent traces rather than synthetic replays.

The platform offers several evaluation features, including the ability to define golden evaluation sets that describe expected agent behaviors, tool calls, and trajectories. AgentEvals supports flexible trajectory matching with strict, unordered, subset, or superset modes, enabling nuanced comparisons between expected and observed agent actions. Users can also create custom evaluators in Python, JavaScript, or any language of their choice and share them through a community registry.

AgentEvals is accessible through both a command-line interface (CLI) and a web user interface (Web UI). The CLI is tailored for automation and integration into CI/CD pipelines, enabling teams to gate deployments based on agent behavior quality scores. The Web UI provides interactive capabilities for visually inspecting traces, browsing results, comparing runs, and drilling into detailed evaluations. Installation is available via Python wheel, and evaluations can be run directly against trace files.

0 license. Its focus on trace-driven evaluation and support for both automated and interactive workflows make it suitable for developers and teams seeking to ensure the reliability and quality of AI agent behavior before production deployment.

Open SourceApache-2.0WebCLICloud-managed
AAgentEvals preview
Visit aevals.ai↗
141stars
18forks
6features
2026since

Overview

6 features

AgentEvals sits in PulseGate's LLM eval & observability category. It focuses on evaluating and benchmarking AI agent behavior from real production traces without rerunning agents. It is built as an open-source project for AI developers and ML engineers. The project is open source (Apache-2.0). AgentEvals is available on the web and the command line.

Behind AgentEvals is AgentEvals Maintainers, and it first shipped in 2026. The project is developed in the open on GitHub with 141 stars and 159 commits in the last 90 days. Among its 6 catalogued features are trace-based evaluation, LLM-powered scoring, and custom evaluators.

Summary written by a language model from the project’s public pages.

  • ✓Trace-based evaluation
  • ✓LLM-powered scoring
  • ✓Custom evaluators
  • ✓CI/CD integration
  • ✓Trajectory matching
  • ✓Golden eval sets
Tags
agent-evaluationopentelemetrytrace-analysis
AI capabilities
Structured
Weights: Open

Built with & integrations

Framework
FastAPI
AI providers
openai
Runs on
BrowserCLICloud-managed

Trust & compliance

License
Apache-2.0
Verified signals
✓HTTPS✓Open Source✓Free tier✓GitHub · ★ 141✓Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed26 Jun · 23:51 UTC
    Listing verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about AgentEvals

What is AgentEvals?
AgentEvals focuses on evaluating and benchmarking AI agent behavior from real production traces without rerunning agents. It is catalogued under LLM eval & observability on PulseGate.
Who is AgentEvals for?
AgentEvals is an open-source project built for AI developers and ML engineers.
Does AgentEvals have a free plan?
Yes — AgentEvals is open source under the Apache-2.0 license and free to use.
What platforms does AgentEvals run on?
AgentEvals runs on the web and the command line.
Is AgentEvals still maintained?
The GitHub repository shows 159 commits in the last 90 days.
What are alternatives to AgentEvals?
Similar projects tracked by PulseGate include AgentEval, agent-eval, and openagent-eval.AgentEvalagent-evalopenagent-eval
Who makes AgentEvals?
AgentEvals is developed by AgentEvals Maintainers.
How long has AgentEvals been around?
AgentEvals first shipped in 2026.

At a glance

Pricing
Open Source
Platforms
Cli · Web
Languages
English
Open source
Yes · ★ 141
License
Apache-2.0
Built for
AI developers and ML engineers
Model
Open source
Solves
Evaluating and benchmarking AI agent behavior from real production traces without rerunning agents.

Registered as

GitHub
agentevals-dev/agentevals

Developer

AgentEvals Maintainers
Team
↗ GitHub

Open source

View on GitHub →
Stars
141
Forks
18
Open issues
24
Last commit
25 Jun 2026
Commits 90d
159
Contributors
12
Authorship
Team
Default branch
main
Latest release
v0.9.5 · 25 Jun 2026

Index record

Identity confidence
High · 95.6
Indexed
26 Jun 2026
Lifecycle
Alive
Last seen
26 Jun 2026
Identity audit (12)
Slug
agentevals-score-agent-behavior-from-opentelemetry-traces-aevals-ai
Lifecycle last checked
8 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
26 Jun 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Derived from the project's own page and URL.
Category from
Assigned by a language model.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model from page content.
Canonical URL
https://aevals.ai/

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • AGAgentEvalagenteval.dev
  • AGagent-evalpypi.org
  • OPopenagent-evalpypi.org
  • AGagentaudit-evalpypi.org
  • EVevalitepypi.org
  • AIAI-Evalsai-evals.io