evalgrid
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 29 Jun 2026. How this is checked
evalgrid is an open-source framework designed for evaluating, scoring, and tracking the correctness of AI agents at scale. It provides tools for running large-scale agent tests, collecting metrics, and analyzing results, making it ideal for AI researchers and developers who need robust evaluation pipelines.
Inferred · not functionally tested
Overview
5 featuresPurpose: Evaluating and tracking the correctness and performance of AI agents at scale for developers and researchers.
Inferred · not functionally tested
Audience: AI researchers and developers
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
evalgrid sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on evaluating and tracking the correctness and performance of AI agents at scale for developers and researchers. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers. Basis unknown · not verified: evalgrid is open source under the MIT license. Basis unknown · not verified: It ships for the command line, and it can be self-hosted.
evalgrid first shipped in 2026. Inferred · not functionally tested: Key capabilities include agent evaluation, scoring, and tracking.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- agent evaluation
- scoring
- tracking
- scalability
- AI agent support
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about evalgrid
- What does evalgrid do?
- Inferred · not functionally tested: Evalgrid focuses on evaluating and tracking the correctness and performance of AI agents at scale for developers and researchers. It is catalogued under Agent evaluation & testing on PulseGate.
- Who should use evalgrid?
- Inferred · not functionally tested: evalgrid is an open-source project built for AI researchers and developers.
- Is evalgrid free?
- Basis unknown · not verified: Yes — evalgrid is open source under the MIT license and free to use.
- What platforms does evalgrid run on?
- Basis unknown · not verified: evalgrid runs on the command line. It can also be self-hosted.
- Is evalgrid still active?
- PulseGate's liveness check found it on 14 Sep 2026.
- What projects are similar to evalgrid?
- Similar projects tracked by PulseGate include agent-eval, evalite, and agentsec-eval.agent-evalevaliteagentsec-eval
- When did evalgrid launch?
- evalgrid first shipped in 2026.
- Is evalgrid open source?
- Basis unknown · not verified: Yes — evalgrid is open source under the MIT license, developed on GitHub.
Similar projects
Closest matches by what these projects do