PulseGateCategoriesMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets90 categories · 139 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesIndustry UpdatesMethodologySupply IndexData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGateAGAgentEval
Visit↗
Skip to content
  1. Index›
  2. LLM eval & observability›
  3. AgentEval
← Back to the index
AG

AgentEval

agenteval.dev·Infrastructure

AI. NET ecosystem.

The platform provides features such as tool usage validation, which allows users to assert on tool chains and verify that specific tools are called in the correct order with appropriate arguments. Stochastic evaluation is supported, enabling repeated runs of agent tasks to assess actual success rates and standard deviations, reflecting the non-deterministic nature of large language models. Workflow evaluation capabilities allow for the testing of multi-agent flows, including validation of executor order, edge traversal, and per-graph tool calls. Performance evaluation tools enable users to set and assert on service level agreements (SLAs) related to response times, total duration, and estimated costs.

AgentEval includes model comparison functionality, letting users benchmark multiple models against defined metrics such as tool accuracy, relevance, and cost per request. The toolkit also supports recording and replaying agent interactions, which allows for consistent, repeatable evaluations without incurring additional API costs. Security evaluation is addressed through a Red Team module that tests agents against 258 attack probes across all 10 OWASP LLM Top 10 vulnerabilities, with MITRE ATLAS technique mapping. This module covers a wide range of attack types, including prompt injection, jailbreaks, PII leakage, and more, and supports both quick scans and advanced, customizable attack pipelines. Security compliance reports can be exported in PDF format.

Memory evaluation is another key feature, with tools for benchmarking agent memory retention, recall depth, temporal reasoning, fact updates, cross-session persistence, and noise resistance. Results can be exported as interactive HTML reports. NET developers seeking to rigorously test, benchmark, and ensure the reliability, security, and performance of their AI agents before production use.

Open SourceMITCLIAPI
AAgentEval preview
Visit agenteval.dev↗
124stars
10forks
1alternative
5features
2026since

Overview

5 features

AgentEval sits in PulseGate's LLM eval & observability category. It focuses on evaluating and benchmarking the performance and reliability of AI agents in .NET environments. It is built as an open-source project for .NET developers. The project is open source (MIT). It runs on the command line and API.

AgentEval first shipped in 2026. Development happens publicly on GitHub with 124 stars and 38 commits in the last 90 days. Among its 5 catalogued features are tool usage validation, RAG quality metrics, and stochastic evaluation.

Summary written by a language model from the project’s public pages.

  • ✓Tool usage validation
  • ✓RAG quality metrics
  • ✓Stochastic evaluation
  • ✓Model comparison
  • ✓Memory benchmarks
Tags
agent-evaluationrag-metricsdotnet-toolkitai-benchmarking
AI capabilities
CodeStructured

Built with & integrations

AI providers
openai
Runs on
CLIAPI-only
Detected from
openai
bgpt- in the HTML

Trust & compliance

License
MIT
Verified signals
✓HTTPS✓Open Source✓Free tier✓GitHub · ★ 124✓Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed26 Jun · 00:44 UTC
    Listing verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about AgentEval

What is AgentEval?
AgentEval focuses on evaluating and benchmarking the performance and reliability of AI agents in .NET environments. It is catalogued under LLM eval & observability on PulseGate.
Who should use AgentEval?
AgentEval is an open-source project built for .NET developers.
Does AgentEval have a free plan?
Yes — AgentEval is open source under the MIT license and free to use.
What platforms does AgentEval run on?
AgentEval runs on the command line and API.
Is AgentEval still maintained?
The GitHub repository shows 38 commits in the last 90 days.
What are alternatives to AgentEval?
Similar projects tracked by PulseGate include agent-eval, AgentEvals, and evalite.agent-evalAgentEvalsevaliteSee all 18 alternatives
How long has AgentEval been around?
AgentEval first shipped in 2026.
Is AgentEval open source?
Yes — AgentEval is open source under the MIT license, developed on GitHub.

At a glance

Pricing
Open Source
Platforms
Api · Cli
Languages
English
Open source
Yes · ★ 124
License
MIT
Built for
.NET developers
Model
Open source
Solves
Evaluating and benchmarking the performance and reliability of AI agents in .NET environments.

Registered as

GitHub
AgentEvalHQ/AgentEval

Developer

Agentevalhq
Small team
↗ GitHub

Open source

View on GitHub →
Stars
124
Forks
10
Open issues
1
Last commit
25 Jun 2026
Commits 90d
38
Contributors
4
Authorship
Small team
Default branch
main
Latest release
v0.10.1-beta · 18 May 2026

Index record

Identity confidence
High · 95
Indexed
26 Jun 2026
Lifecycle
Alive
Last seen
26 Jun 2026
Identity audit (12)
Slug
agenteval-agenteval-agenteval-dev
Lifecycle last checked
8 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
26 Jun 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Derived from the project's own page and URL.
Category from
Assigned by a language model.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model from page content.
Canonical URL
https://agenteval.dev/

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • AGagent-evalpypi.org
  • AGAgentEvalsaevals.ai
  • EVevalitepypi.org
  • OPopenagent-evalpypi.org
  • AGagentaudit-evalpypi.org
  • EVEvalgentevalgent.com

See 18 alternatives to AgentEval →