Skip to content
Back to the index

llm-agent-bench

PyPIInfrastructure

PulseGate's liveness check found it on 13 Sep 2026; it is registered on PyPI and has been in the index since 25 Jun 2026. How this is checked

llm-agent-bench is an open-source CLI tool for benchmarking autonomous AI agents on task completion, tool use, goal adherence, and safety. It works with any agent by providing a callable interface, supporting AI researchers and developers.

Inferred · not functionally tested

Open SourceMITWebCLILinuxmacOSWindows
Visit PyPI

Overview

4 features

Purpose: Evaluating and benchmarking the performance and safety of autonomous AI agents.

Inferred · not functionally tested

Audience: AI researchers and developers

Inferred · not functionally tested

Functions: analytics

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: browser, cli, linux, macos, windows

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org. These links do not verify the individual claims.

llm-agent-bench sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on evaluating and benchmarking the performance and safety of autonomous AI agents. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers. Basis unknown · not verified: llm-agent-bench is open source under the MIT license. Basis unknown · not verified: It runs on the web, the command line, Linux, macOS, and Windows.

llm-agent-bench first shipped in 2026. Inferred · not functionally tested: Key capabilities include agent benchmarking, task evaluation, and tool use analysis.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Agent benchmarking
  • Task evaluation
  • Tool use analysis
  • Safety checks

Topics: Inferred · not functionally tested

Tags
agent-benchmarkingai-evaluationautonomous-agents
AI capabilities
CodeStructured

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
BrowserCLILinuxmacOSWindows

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tier

Indexing history

What PulseGate has recorded for this listing

Nothing recorded for this listing in this window.

Frequently asked questions about llm-agent-bench

What does llm-agent-bench do?
Inferred · not functionally tested: Llm-agent-bench focuses on evaluating and benchmarking the performance and safety of autonomous AI agents. It is catalogued under Agent evaluation & testing on PulseGate.
Who should use llm-agent-bench?
Inferred · not functionally tested: llm-agent-bench is an open-source project built for AI researchers and developers.
Does llm-agent-bench have a free plan?
Basis unknown · not verified: Yes — llm-agent-bench is open source under the MIT license and free to use.
What platforms does llm-agent-bench run on?
Basis unknown · not verified: llm-agent-bench runs on the web, the command line, Linux, macOS, and Windows.
Is llm-agent-bench still active?
PulseGate's liveness check found it on 13 Sep 2026.
What are alternatives to llm-agent-bench?
Similar projects tracked by PulseGate include litebench, agentbench-cli, and bench-my-llm.litebenchagentbench-clibench-my-llm
When did llm-agent-bench launch?
llm-agent-bench first shipped in 2026.
Is llm-agent-bench open source?
Basis unknown · not verified: Yes — llm-agent-bench is open source under the MIT license.

Similar projects

Closest matches by what these projects do