CatchBench
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 14 Sep 2026. How this is checked
CatchBench is an MIT-licensed benchmark for auditing failures in AI agents across the full lifecycle. It provides developers and researchers with tooling and evaluation scenarios for assessing agent reliability.
Inferred · not functionally tested
Overview
4 featuresPurpose: Auditing and measuring failures in AI agents across development and operational lifecycles.
Inferred · not functionally tested
Audience: AI researchers and developers evaluating agent reliability
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
CatchBench sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on auditing and measuring failures in AI agents across development and operational lifecycles. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers evaluating agent reliability. Basis unknown · not verified: CatchBench is open source under the MIT license. Basis unknown · not verified: CatchBench is available on the command line, and it can be self-hosted.
Yizhao Zhao builds and maintains CatchBench, and it first shipped in 2026. The project is developed in the open on GitHub with 74 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, failure auditing, and lifecycle evaluation.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Agent benchmarking
- Failure auditing
- Lifecycle evaluation
- Reliability assessment
Topics: Inferred · not functionally tested
Built with & integrations
- Unspecified agent
- AGENTS.md
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
Frequently asked questions about CatchBench
- What is CatchBench?
- Inferred · not functionally tested: CatchBench focuses on auditing and measuring failures in AI agents across development and operational lifecycles. It is catalogued under Agent evaluation & testing on PulseGate.
- Who should use CatchBench?
- Inferred · not functionally tested: CatchBench is an open-source project built for AI researchers and developers evaluating agent reliability.
- Is CatchBench free?
- Basis unknown · not verified: Yes — CatchBench is open source under the MIT license and free to use.
- What platforms does CatchBench run on?
- Basis unknown · not verified: CatchBench runs on the command line. It can also be self-hosted.
- Is CatchBench still active?
- PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 74 commits in the last 90 days.
- What are alternatives to CatchBench?
- Similar projects tracked by PulseGate include CheatBench, benchspec, and Bench'd.CheatBenchbenchspecBench'd
- Who develops CatchBench?
- CatchBench is developed by Yizhao Zhao.
- How long has CatchBench been around?
- CatchBench first shipped in 2026.
Similar projects
Closest matches by what these projects do