Skip to content
Back to the index

proofbench

PyPIInfrastructure

PulseGate's liveness check found it on 14 Sep 2026; it is registered on PyPI and has been in the index since 1 Jul 2026. How this is checked

proofbench is an open-source CLI tool for evaluating and improving the performance of headless AI agents. It uses configuration files to benchmark agents against ground-truth corpora and supports prompt optimization and self-improvement workflows. Ideal for AI researchers and developers working on agent evaluation.

Inferred · not functionally tested

Open SourceMITCLI
Visit PyPI

Overview

5 features

Purpose: Automating the evaluation and improvement of headless agent skills against ground-truth datasets.

Inferred · not functionally tested

Audience: AI researchers and developers

Inferred · not functionally tested

Functions: analytics

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org. These links do not verify the individual claims.

proofbench is an Agent evaluation & testing project. Inferred · not functionally tested: It focuses on automating the evaluation and improvement of headless agent skills against ground-truth datasets. Inferred · not functionally tested: proofbench is an open-source project aimed at AI researchers and developers. Basis unknown · not verified: proofbench is open source under the MIT license. Basis unknown · not verified: It ships for the command line.

proofbench first shipped in 2026. Inferred · not functionally tested: Key capabilities include agent benchmarking, config-driven, and self-improvement.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Agent benchmarking
  • Config-driven
  • Self-improvement
  • Ground-truth evaluation
  • Prompt optimization

Topics: Inferred · not functionally tested

Tags
agent-benchmarkingeval-harnessprompt-optimization
AI capabilities
Code
Weights: Open

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
CLI

Trust & compliance

License
MIT
Public signals
HTTPSOpen Source

Indexing history

What PulseGate has recorded for this listing

Nothing recorded for this listing in this window.

Frequently asked questions about proofbench

What is proofbench?
Inferred · not functionally tested: Proofbench focuses on automating the evaluation and improvement of headless agent skills against ground-truth datasets. It is catalogued under Agent evaluation & testing on PulseGate.
Who should use proofbench?
Inferred · not functionally tested: proofbench is an open-source project built for AI researchers and developers.
Does proofbench have a free plan?
Basis unknown · not verified: Yes — proofbench is open source under the MIT license and free to use.
What platforms does proofbench run on?
Basis unknown · not verified: proofbench runs on the command line.
Is proofbench still active?
PulseGate's liveness check found it on 14 Sep 2026.
What are alternatives to proofbench?
Similar projects tracked by PulseGate include benchskills, benchspec, and agentbench-cli.benchskillsbenchspecagentbench-cli
When did proofbench launch?
proofbench first shipped in 2026.
Is proofbench open source?
Basis unknown · not verified: Yes — proofbench is open source under the MIT license.

Similar projects

Closest matches by what these projects do