rerun-bench
PulseGate's liveness check found it on 3 Oct 2026; it is registered on PyPI and has been in the index since 3 Oct 2026. How this is checked
rerun-bench is an open-source benchmarking tool for running the same coding task multiple times with an agent. It measures consistency, reliability, and execution cost to help developers evaluate coding-agent performance.
Inferred · not functionally tested
Overview
6 featuresPurpose: Measuring the consistency, reliability, and cost of coding agents across repeated task runs.
Inferred · not functionally tested
Audience: developers and AI agent evaluators
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org. These links do not verify the individual claims.
In the LLM evaluation & benchmarks space, rerun-bench takes a focused approach. Inferred · not functionally tested: It focuses on measuring the consistency, reliability, and cost of coding agents across repeated task runs. Inferred · not functionally tested: rerun-bench is an open-source project aimed at developers and AI agent evaluators. Basis unknown · not verified: rerun-bench is open source under the MIT license. Basis unknown · not verified: It ships for the command line, and it can be self-hosted.
rerun-bench first shipped in 2026. Inferred · not functionally tested: Key capabilities include repeated task runs, agent benchmarking, and consistency measurement.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Repeated task runs
- Agent benchmarking
- Consistency measurement
- Cost measurement
- Reliability evaluation
- Coding task evaluation
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
3What PulseGate has recorded for this listing
- Indexed3 Oct · 17:58 UTCrerun-bench seen via PyPI Bulk EnumeratorSource: PyPI Bulk Enumerator · Open
Frequently asked questions about rerun-bench
- What does rerun-bench do?
- Inferred · not functionally tested: Rerun-bench focuses on measuring the consistency, reliability, and cost of coding agents across repeated task runs. It is catalogued under LLM evaluation & benchmarks on PulseGate.
- Who is rerun-bench for?
- Inferred · not functionally tested: rerun-bench is an open-source project built for developers and AI agent evaluators.
- Is rerun-bench free?
- Basis unknown · not verified: Yes — rerun-bench is open source under the MIT license and free to use.
- What platforms does rerun-bench run on?
- Basis unknown · not verified: rerun-bench runs on the command line. It can also be self-hosted.
- Is rerun-bench still active?
- PulseGate's liveness check found it on 3 Oct 2026.
- What are alternatives to rerun-bench?
- Similar projects tracked by PulseGate include contamcheck, quantdiff, and toolkit-eval-harness.contamcheckquantdifftoolkit-eval-harness
- How long has rerun-bench been around?
- rerun-bench first shipped in 2026.
- Is rerun-bench open source?
- Basis unknown · not verified: Yes — rerun-bench is open source under the MIT license.
Also in LLM evaluation & benchmarks
Same category — not a similarity match