Skip to content
Back to the index

CatchBench

PyPIInfrastructure

PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 14 Sep 2026. How this is checked

CatchBench is an MIT-licensed benchmark for auditing failures in AI agents across the full lifecycle. It provides developers and researchers with tooling and evaluation scenarios for assessing agent reliability.

Inferred · not functionally tested

Open SourceMITCLISelf-hosted
Visit PyPI
9stars
4features
2026since

Overview

4 features

Purpose: Auditing and measuring failures in AI agents across development and operational lifecycles.

Inferred · not functionally tested

Audience: AI researchers and developers evaluating agent reliability

Inferred · not functionally tested

Functions: analytics

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org · github.com. These links do not verify the individual claims.

CatchBench sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on auditing and measuring failures in AI agents across development and operational lifecycles. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers evaluating agent reliability. Basis unknown · not verified: CatchBench is open source under the MIT license. Basis unknown · not verified: CatchBench is available on the command line, and it can be self-hosted.

Yizhao Zhao builds and maintains CatchBench, and it first shipped in 2026. The project is developed in the open on GitHub with 74 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, failure auditing, and lifecycle evaluation.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Agent benchmarking
  • Failure auditing
  • Lifecycle evaluation
  • Reliability assessment

Topics: Inferred · not functionally tested

Tags
agent-benchmarkingfailure-analysisagent-reliabilitylifecycle-auditing
AI capabilities
Text
Inference: Local

JSON profile · Text profile · Access guide

Built with & integrations

Written with
Unspecified agent
Runs on
CLISelf-hosted
Written with — evidence
Unspecified agent
AGENTS.md

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tierGitHub · ★ 9Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed14 Sep · 23:36 UTC
    catchbench seen via PyPI Fresh Feed
    Source: PyPI Fresh Feed · Open

Frequently asked questions about CatchBench

What is CatchBench?
Inferred · not functionally tested: CatchBench focuses on auditing and measuring failures in AI agents across development and operational lifecycles. It is catalogued under Agent evaluation & testing on PulseGate.
Who should use CatchBench?
Inferred · not functionally tested: CatchBench is an open-source project built for AI researchers and developers evaluating agent reliability.
Is CatchBench free?
Basis unknown · not verified: Yes — CatchBench is open source under the MIT license and free to use.
What platforms does CatchBench run on?
Basis unknown · not verified: CatchBench runs on the command line. It can also be self-hosted.
Is CatchBench still active?
PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 74 commits in the last 90 days.
What are alternatives to CatchBench?
Similar projects tracked by PulseGate include CheatBench, benchspec, and Bench'd.CheatBenchbenchspecBench'd
Who develops CatchBench?
CatchBench is developed by Yizhao Zhao.
How long has CatchBench been around?
CatchBench first shipped in 2026.

Similar projects

Closest matches by what these projects do