Skip to content
Back to the index

agent-evaluation-lab

PyPIInfrastructure

PulseGate's liveness check found it on 19 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 19 Sep 2026. How this is checked

Agent Evaluation Lab is an open-source benchmark harness for evaluating command-line coding agents with deterministic tests. It is intended for developers and researchers building or comparing AI coding agents.

Inferred · not functionally tested

Open SourceMITCLISelf-hosted
Visit PyPI
2stars
5features
2026since

Overview

5 features

Purpose: Evaluating command-line coding agents consistently and reproducibly.

Inferred · not functionally tested

Audience: AI agent developers and researchers

Inferred · not functionally tested

Functions: Unknown

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org · github.com. These links do not verify the individual claims.

agent-evaluation-lab is an Agent evaluation & testing project. Inferred · not functionally tested: It focuses on evaluating command-line coding agents consistently and reproducibly. Inferred · not functionally tested: agent-evaluation-lab is an open-source project aimed at AI agent developers and researchers. Basis unknown · not verified: The project is open source (MIT). Basis unknown · not verified: It runs on the command line, and it can be self-hosted.

Aditya Supag1 builds and maintains agent-evaluation-lab, and it first shipped in 2026. The project is developed in the open on GitHub with 37 commits in the last 90 days. Inferred · not functionally tested: Among its 5 catalogued features are Deterministic Benchmarks, Agent Testing, and CLI Harness.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Deterministic Benchmarks
  • Agent Testing
  • CLI Harness
  • Coding Agent Evaluation
  • Benchmark Reporting

Topics: Inferred · not functionally tested

Tags
agent-benchmarkscoding-agent-evaluationdeterministic-testingcli-harness
AI capabilities
Code

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
CLISelf-hosted

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tierGitHub · ★ 2Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed19 Sep · 14:30 UTC
    agent-evaluation-lab seen via PyPI Bulk Enumerator
    Source: PyPI Bulk Enumerator · Open

Frequently asked questions about agent-evaluation-lab

What is agent-evaluation-lab?
Inferred · not functionally tested: Agent-evaluation-lab focuses on evaluating command-line coding agents consistently and reproducibly. It is catalogued under Agent evaluation & testing on PulseGate.
Who is agent-evaluation-lab for?
Inferred · not functionally tested: agent-evaluation-lab is an open-source project built for AI agent developers and researchers.
Is agent-evaluation-lab free?
Basis unknown · not verified: Yes — agent-evaluation-lab is open source under the MIT license and free to use.
What platforms does agent-evaluation-lab run on?
Basis unknown · not verified: agent-evaluation-lab runs on the command line. It can also be self-hosted.
Is agent-evaluation-lab still maintained?
PulseGate's liveness check found it on 19 Sep 2026. Its GitHub repository shows 37 commits in the last 90 days.
What are alternatives to agent-evaluation-lab?
Similar projects tracked by PulseGate include AgentEval, agent-exam, and agent-eval.AgentEvalagent-examagent-eval
Who makes agent-evaluation-lab?
agent-evaluation-lab is developed by Aditya Supag1.
How long has agent-evaluation-lab been around?
agent-evaluation-lab first shipped in 2026.

Similar projects

Closest matches by what these projects do