Skip to content
Back to the index

rerun-bench

PyPIInfrastructure

PulseGate's liveness check found it on 3 Oct 2026; it is registered on PyPI and has been in the index since 3 Oct 2026. How this is checked

rerun-bench is an open-source benchmarking tool for running the same coding task multiple times with an agent. It measures consistency, reliability, and execution cost to help developers evaluate coding-agent performance.

Inferred · not functionally tested

Open SourceMITCLISelf-hosted
Visit PyPI

Overview

6 features

Purpose: Measuring the consistency, reliability, and cost of coding agents across repeated task runs.

Inferred · not functionally tested

Audience: developers and AI agent evaluators

Inferred · not functionally tested

Functions: analytics

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org. These links do not verify the individual claims.

In the LLM evaluation & benchmarks space, rerun-bench takes a focused approach. Inferred · not functionally tested: It focuses on measuring the consistency, reliability, and cost of coding agents across repeated task runs. Inferred · not functionally tested: rerun-bench is an open-source project aimed at developers and AI agent evaluators. Basis unknown · not verified: rerun-bench is open source under the MIT license. Basis unknown · not verified: It ships for the command line, and it can be self-hosted.

rerun-bench first shipped in 2026. Inferred · not functionally tested: Key capabilities include repeated task runs, agent benchmarking, and consistency measurement.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Repeated task runs
  • Agent benchmarking
  • Consistency measurement
  • Cost measurement
  • Reliability evaluation
  • Coding task evaluation

Topics: Inferred · not functionally tested

Tags
coding-agent-benchmarksagent-consistencyllm-evaluationcost-measurement
AI capabilities
CodeStructured
Inference: Local

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
CLISelf-hosted

Trust & compliance

License
MIT
Public signals
HTTPSOpen Source

Indexing history

3

What PulseGate has recorded for this listing

  1. Indexed5 Oct · 08:07 UTC
    rerun-bench seen via PyPI Fresh Feed
    Source: PyPI Fresh Feed · Open
  2. Indexed4 Oct · 04:25 UTC
    rerun-bench seen via PyPI Fresh Feed
    Source: PyPI Fresh Feed · Open
  3. Indexed3 Oct · 17:58 UTC
    rerun-bench seen via PyPI Bulk Enumerator
    Source: PyPI Bulk Enumerator · Open

Frequently asked questions about rerun-bench

What does rerun-bench do?
Inferred · not functionally tested: Rerun-bench focuses on measuring the consistency, reliability, and cost of coding agents across repeated task runs. It is catalogued under LLM evaluation & benchmarks on PulseGate.
Who is rerun-bench for?
Inferred · not functionally tested: rerun-bench is an open-source project built for developers and AI agent evaluators.
Is rerun-bench free?
Basis unknown · not verified: Yes — rerun-bench is open source under the MIT license and free to use.
What platforms does rerun-bench run on?
Basis unknown · not verified: rerun-bench runs on the command line. It can also be self-hosted.
Is rerun-bench still active?
PulseGate's liveness check found it on 3 Oct 2026.
What are alternatives to rerun-bench?
Similar projects tracked by PulseGate include contamcheck, quantdiff, and toolkit-eval-harness.contamcheckquantdifftoolkit-eval-harness
How long has rerun-bench been around?
rerun-bench first shipped in 2026.
Is rerun-bench open source?
Basis unknown · not verified: Yes — rerun-bench is open source under the MIT license.

Also in LLM evaluation & benchmarks

Same category — not a similarity match