Skip to content
Alternatives
Software like litebench
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- llm-agent-benchpypi.orgllm-agent-bench is an open-source CLI tool for benchmarking autonomous AI agents on task completion, tool use, goal adherence, and safety. It works with any agent by providing a callable interface, supporting AI researchers and developers.
- labbench-cligithub.comlabbench-cli is an open-source terminal AI assistant designed to help users with Python, Jupyter notebooks, and data workflows. It provides AI-powered assistance and automation for data scientists and developers working in terminal environments.
- local-bench-ailocal-bench.ailocal-bench-ai is a local AI benchmarking suite for open-weight models running on local hardware. It compares model quality against the amount of VRAM required to run them, and it presents results as a Local Intelligence Index with scored variants on a board and leaderboard. The suite measures tokens per second, VRAM, latency, accuracy, and the reproducible run hash attached to each artifact. It supports a local, open-weight, judge-free runtime with deterministic warmup prompts, answer-key evaluation, and separate public and full execution paths. The public command runs five non-agentic axes with a static-only mode, while full six-axis execution uses a managed AppWorld harness. The leaderboard shows ranked model variants, including per-axis scores for agentic ability, knowledge, instruction following, tool calling, coding, and math, along with bench time and fit by VRAM tier. Users can pick a model, select VRAM and runtime settings, and get exact commands for benchmarking. The page also shows a catalog browse view and a way to paste a Hugging Face repository. Runtime options named on the page include llama.cpp, LM Studio, and vLLM, and it notes support for bringing a custom OpenAI-compatible server or custom hardware setups. The commands verify downloads against pinned hashes, cache tokenizers, check publishability before starting, and can ask before submitting results. The page also records tokenizer and chat-template digests for some runs. Installation is shown with pip using the package name local-bench-ai, including an extra for Hugging Face support. The page warns that plain pip install localbench installs an unrelated third-party package. It also states that Python 3.11+ is required for the public static path, and that the full six-axis lane additionally needs Windows with WSL2. Hugging Face gated repositories are supported through hf auth login after license acceptance, and llama-server can read its token from the HF_TOKEN environment variable. The package is described as a catalog-pinned one-command flow for benchmark, score, and optional submit workflows.
- agentbench-clipypi.orgagentbench-cli is an open-source command-line tool that allows AI developers and researchers to evaluate, test, and scan the behavior and safety of AI agents. It provides automated checks and reporting for agent performance and compliance.
- benchflowgithub.comMulti-turn agent benchmarking with ACP — run any agent, any model, any provider.
- BenchLoopbench-loop.comBenchLoop is a benchmarking tool for local large language models, providing quality, speed, and reliability scores through both a web app and CLI. It supports various LLM runtimes like Ollama and OpenAI-compatible endpoints, helping developers and researchers evaluate model performance.
- llm-benchmark-runnergithub.comllm-benchmark-runner is an open-source CLI tool that allows developers and researchers to benchmark the inference latency and throughput of large language model APIs. It supports OpenAI-compatible endpoints and provides detailed performance metrics for model evaluation.
- ToolBenchpypi.orgToolBench provides a platform and command-line interface for creating and running benchmarks that test agentic tools and LLM-powered harnesses. It helps developers systematically evaluate tool-calling capabilities, orchestration logic, and overall agent performance across different models and scenarios. Primarily used by AI researchers and engineers building autonomous agents.
- porchbenchpypi.orgporchbench is an open-source CLI tool designed for rigorous benchmarking and evaluation of local large language models (LLMs). It provides paired statistics, LLM-as-judge scoring, and ensures reproducible runs, making it ideal for AI researchers and developers who need to assess the quality and performance of local AI models. The tool supports integration with local inference engines such as Ollama and quantized models.
- llm-bench-costgithub.comCLI tool for comparing LLM API pricing, ranked by cost-effectiveness against LMSYS Arena scores
- homebenchpypi.orgHomebench provides a simple terminal-based interface to benchmark locally installed large language models. It measures tokens-per-second, memory consumption, and output quality, then presents results in a clear leaderboard format. Built for users of tools like Ollama, it helps developers and enthusiasts quickly evaluate and compare different local LLMs on their own laptops without complex setup.
- Terminal-Benchtbench.aiTerminal-Bench provides a suite of benchmarks for evaluating the capabilities of AI agents in terminal environments. It offers standardized tasks, leaderboards, and performance metrics to help researchers and developers assess and compare agent performance. The platform is open source and designed for the AI research community.
- benchcaddygithub.combenchcaddy is an open-source command-line tool designed for developers and performance engineers to run benchmark sweeps and analyze results. It captures environment details for reproducibility and provides lightweight profiling and analysis capabilities. The tool is suitable for optimizing and validating software performance.
- leanlabpypi.orgleanlab is an open-source CLI tool designed for evolving and evaluating AI agent experiments and coding tasks. It supports experiment evolution against fixed metrics and automates the spec-to-merge workflow with locked acceptance tests, helping AI researchers and developers improve agent reliability and performance.
- InferenceBenchpypi.orgInferenceBench is an open-source CLI suite designed for AI engineers and ML researchers to benchmark inference performance across multiple AI vendors. It provides vendor-neutral, signed-envelope benchmarks to ensure reliable and reproducible results.
- benchdiffpypi.orgbenchdiff is an open-source CLI tool for benchmarking and comparing performance results. It provides rich terminal output, supports markdown export, and is designed for developers who need to analyze and share benchmarking data efficiently.
- BenchLLMbenchllm.comBenchLLM is a platform designed for evaluating large language model (LLM) applications. It enables developers to build test suites, generate quality reports, and choose between automated, interactive, or custom evaluation strategies. BenchLLM supports both API and CLI usage, making it suitable for AI developers and ML engineers seeking robust model evaluation tools.
- benchskillspypi.orgbenchskills is an open-source framework for benchmarking and evaluating the skills of AI agents. It provides tools for researchers and developers to assess agent performance across various tasks and metrics.
- llmetergithub.comllmeter is an open-source CLI tool that enables developers and researchers to profile the latency and throughput of large language models (LLMs). It provides cross-platform support for performance testing and optimization of AI workloads.
- LitigationBenchlitco.aiLitigationBench is a benchmark developed by Litco that measures language models on litigation tasks. It runs each model on identical tasks twice, once without safeguards and once with Litco’s safeguards active inside its production agent. The benchmark publishes both scores, the performance gap, fabricated authorities, false premises, and costs to support evaluation of model behavior in legal contexts. A composite quality score appears for each setting, shown alongside the score before penalties so the effect of penalties is visible. Columns for fabricated authorities and false premises use color coding: green for zero instances, amber for an adopted false premise, and red for a fabrication. Full flag details are available in expanded row panels and a candor matrix. Cost reflects the metered provider bill for the tasks, while the self-hosted row runs on Litco’s own hardware and lists price in its cost column. The leaderboard allows sorting by column, with arrows indicating the preferred direction. Clicking a row displays both settings side by side along with serving pins. Models missing a setting appear below scored rows with an explanatory note. The benchmark focuses on frontier and self-hosted models and includes every failure observed, even those occurring inside the product.
Ranked by how close each one sits to litebench in the index, not by popularity. Back to litebench →