Skip to content
Back to the index

evalgrid

PyPIInfrastructure

PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 29 Jun 2026. How this is checked

evalgrid is an open-source framework designed for evaluating, scoring, and tracking the correctness of AI agents at scale. It provides tools for running large-scale agent tests, collecting metrics, and analyzing results, making it ideal for AI researchers and developers who need robust evaluation pipelines.

Inferred · not functionally tested

Open SourceMITCLISelf-hosted
Visit PyPI

Overview

5 features

Purpose: Evaluating and tracking the correctness and performance of AI agents at scale for developers and researchers.

Inferred · not functionally tested

Audience: AI researchers and developers

Inferred · not functionally tested

Functions: analytics

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org · github.com. These links do not verify the individual claims.

evalgrid sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on evaluating and tracking the correctness and performance of AI agents at scale for developers and researchers. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers. Basis unknown · not verified: evalgrid is open source under the MIT license. Basis unknown · not verified: It ships for the command line, and it can be self-hosted.

evalgrid first shipped in 2026. Inferred · not functionally tested: Key capabilities include agent evaluation, scoring, and tracking.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • agent evaluation
  • scoring
  • tracking
  • scalability
  • AI agent support

Topics: Inferred · not functionally tested

Tags
agent-evaluationllm-testingai-benchmarking
AI capabilities
CodeStructured

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
CLISelf-hosted

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tierGitHub

Indexing history

What PulseGate has recorded for this listing

Nothing recorded for this listing in this window.

Frequently asked questions about evalgrid

What does evalgrid do?
Inferred · not functionally tested: Evalgrid focuses on evaluating and tracking the correctness and performance of AI agents at scale for developers and researchers. It is catalogued under Agent evaluation & testing on PulseGate.
Who should use evalgrid?
Inferred · not functionally tested: evalgrid is an open-source project built for AI researchers and developers.
Is evalgrid free?
Basis unknown · not verified: Yes — evalgrid is open source under the MIT license and free to use.
What platforms does evalgrid run on?
Basis unknown · not verified: evalgrid runs on the command line. It can also be self-hosted.
Is evalgrid still active?
PulseGate's liveness check found it on 14 Sep 2026.
What projects are similar to evalgrid?
Similar projects tracked by PulseGate include agent-eval, evalite, and agentsec-eval.agent-evalevaliteagentsec-eval
When did evalgrid launch?
evalgrid first shipped in 2026.
Is evalgrid open source?
Basis unknown · not verified: Yes — evalgrid is open source under the MIT license, developed on GitHub.

Similar projects

Closest matches by what these projects do