Skip to content
Back to the index

AI Evaluator

aievaluator.devInfrastructure

PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 7 Aug 2026. How this is checked

AI Evaluator is an MIT-licensed command-line tool for evaluating LLM agents. It supports agent testing and evaluation workflows that can run locally or in CI/CD environments, targeting developers building AI applications.

Inferred · not functionally tested

Open SourceMITWebCLISelf-hosted
AI Evaluator preview
Visit aievaluator.dev

Overview

5 features

Purpose: Testing and measuring the quality of LLM agents from the command line and in CI/CD pipelines.

Inferred · not functionally tested

Audience: AI developers building LLM agents

Inferred · not functionally tested

Functions: analytics, monitoring

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI, WEB · deployment: browser, cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: aievaluator.dev · github.com. These links do not verify the individual claims.

AI Evaluator is an Agent evaluation & testing project. Inferred · not functionally tested: It focuses on testing and measuring the quality of LLM agents from the command line and in CI/CD pipelines. Inferred · not functionally tested: AI Evaluator is an open-source project aimed at AI developers building LLM agents. Basis unknown · not verified: The project is open source (MIT). Basis unknown · not verified: It ships for the web and the command line, and it can be self-hosted.

aievaluator-dev builds and maintains AI Evaluator, and it first shipped in 2026. The project is developed in the open on GitHub with 111 commits in the last 90 days. Inferred · not functionally tested: Among its 5 catalogued features are LLM evaluation, agent testing, and CLI workflows.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • LLM evaluation
  • Agent testing
  • CLI workflows
  • CI/CD integration
  • Evaluation reports

Topics: Inferred · not functionally tested

Tags
llm-evaluationagent-testingai-qualityci-cd-testing
AI capabilities
TextStructured
Inference: Cloud API

JSON profile · Text profile · Access guide

Built with & integrations

Connectors
github
Runs on
BrowserCLISelf-hosted

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tierGitHubActive maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed7 Aug · 15:52 UTC
    aievaluator seen via PyPI Fresh Feed
    Source: PyPI Fresh Feed · Open

Frequently asked questions about AI Evaluator

What does AI Evaluator do?
Inferred · not functionally tested: AI Evaluator focuses on testing and measuring the quality of LLM agents from the command line and in CI/CD pipelines. It is catalogued under Agent evaluation & testing on PulseGate.
Who should use AI Evaluator?
Inferred · not functionally tested: AI Evaluator is an open-source project built for AI developers building LLM agents.
Is AI Evaluator free?
Basis unknown · not verified: Yes — AI Evaluator is open source under the MIT license and free to use.
What platforms does AI Evaluator run on?
Basis unknown · not verified: AI Evaluator runs on the web and the command line. It can also be self-hosted.
Is AI Evaluator still maintained?
PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 111 commits in the last 90 days.
What are alternatives to AI Evaluator?
Similar projects tracked by PulseGate include AI-Evals, ai-eval-forge, and aihr.AI-Evalsai-eval-forgeaihr
Who makes AI Evaluator?
AI Evaluator is developed by aievaluator-dev.
When did AI Evaluator launch?
AI Evaluator first shipped in 2026.

Similar projects

Closest matches by what these projects do