Skip to content
Alternatives
Software like benchflow
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- agentbench-clipypi.orgagentbench-cli is an open-source command-line tool that allows AI developers and researchers to evaluate, test, and scan the behavior and safety of AI agents. It provides automated checks and reporting for agent performance and compliance.
- benchskillspypi.orgbenchskills is an open-source framework for benchmarking and evaluating the skills of AI agents. It provides tools for researchers and developers to assess agent performance across various tasks and metrics.
- benchstatsgithub.comStatistical Testing for Benchmark Results Comparison
- litebenchgithub.comlitebench is an open-source CLI tool that enables developers and researchers to benchmark large language models and AI agents. It supports quick setup and evaluation workflows, including popular benchmarks like GSM8K and HumanEval.
- llm-agent-benchpypi.orgllm-agent-bench is an open-source CLI tool for benchmarking autonomous AI agents on task completion, tool use, goal adherence, and safety. It works with any agent by providing a callable interface, supporting AI researchers and developers.
- benchstonegithub.comTrustworthy benchmark gating for code-optimization loops: append-only results, harness-owned reference artifacts, statistical PROMOTE/REJECT verdicts.
- InferenceBenchpypi.orgInferenceBench is an open-source CLI suite designed for AI engineers and ML researchers to benchmark inference performance across multiple AI vendors. It provides vendor-neutral, signed-envelope benchmarks to ensure reliable and reproducible results.
- pawbenchgithub.com4-dimensional LLM inference benchmark — multi-turn, multi-agent, parallel dispatch with tool calling
- Terminal-Benchtbench.aiTerminal-Bench provides a suite of benchmarks for evaluating the capabilities of AI agents in terminal environments. It offers standardized tasks, leaderboards, and performance metrics to help researchers and developers assess and compare agent performance. The platform is open source and designed for the AI research community.
- agent-beltgithub.comEvaluation harness for real headless CLI agents - reproducible multi-turn scenarios, rule + LLM scoring, cross-agent comparison
- turbobench-clipypi.orgturbobench-cli is an open-source command-line tool for benchmarking reinforcement-learning environments with correctness-gated evaluations. It supports provider-neutral benchmark workflows for developers and researchers.
- local-bench-ailocal-bench.ailocal-bench-ai is a local AI benchmarking suite for open-weight models running on local hardware. It compares model quality against the amount of VRAM required to run them, and it presents results as a Local Intelligence Index with scored variants on a board and leaderboard. The suite measures tokens per second, VRAM, latency, accuracy, and the reproducible run hash attached to each artifact. It supports a local, open-weight, judge-free runtime with deterministic warmup prompts, answer-key evaluation, and separate public and full execution paths. The public command runs five non-agentic axes with a static-only mode, while full six-axis execution uses a managed AppWorld harness. The leaderboard shows ranked model variants, including per-axis scores for agentic ability, knowledge, instruction following, tool calling, coding, and math, along with bench time and fit by VRAM tier. Users can pick a model, select VRAM and runtime settings, and get exact commands for benchmarking. The page also shows a catalog browse view and a way to paste a Hugging Face repository. Runtime options named on the page include llama.cpp, LM Studio, and vLLM, and it notes support for bringing a custom OpenAI-compatible server or custom hardware setups. The commands verify downloads against pinned hashes, cache tokenizers, check publishability before starting, and can ask before submitting results. The page also records tokenizer and chat-template digests for some runs. Installation is shown with pip using the package name local-bench-ai, including an extra for Hugging Face support. The page warns that plain pip install localbench installs an unrelated third-party package. It also states that Python 3.11+ is required for the public static path, and that the full six-axis lane additionally needs Windows with WSL2. Hugging Face gated repositories are supported through hf auth login after license acceptance, and llama-server can read its token from the HF_TOKEN environment variable. The package is described as a catalog-pinned one-command flow for benchmark, score, and optional submit workflows.
- proofbenchpypi.orgproofbench is an open-source CLI tool for evaluating and improving the performance of headless AI agents. It uses configuration files to benchmark agents against ground-truth corpora and supports prompt optimization and self-improvement workflows. Ideal for AI researchers and developers working on agent evaluation.
- benchcaddygithub.combenchcaddy is an open-source command-line tool designed for developers and performance engineers to run benchmark sweeps and analyze results. It captures environment details for reproducibility and provides lightweight profiling and analysis capabilities. The tool is suitable for optimizing and validating software performance.
- benchdiffpypi.orgbenchdiff is an open-source CLI tool for benchmarking and comparing performance results. It provides rich terminal output, supports markdown export, and is designed for developers who need to analyze and share benchmarking data efficiently.
- janus-labsgithub.com3DMark for AI Agents - Profile AI coding agent capabilities across code quality, error resilience, and instruction resilience
- benchgeckobenchgecko.aibenchgecko is an open-source Python SDK and API for querying AI model benchmarks, pricing, and performance data across hundreds of providers. It enables developers and researchers to make informed decisions about AI model selection and deployment.
- flowwgithub.comAgent reliability simulator — chaos engineering for AI agents
Ranked by how close each one sits to benchflow in the index, not by popularity. Back to benchflow →