Terminal-Bench
PulseGate's liveness check found it on 3 Oct 2026; it is registered on GitHub and has been in the index since 25 Jun 2026. How this is checked
Terminal-Bench provides a suite of benchmarks for evaluating the capabilities of AI agents in terminal environments. It offers standardized tasks, leaderboards, and performance metrics to help researchers and developers assess and compare agent performance. The platform is open source and designed for the AI research community.
Inferred · not functionally tested
Overview
5 featuresPurpose: Measuring and comparing the performance of AI agents in terminal-based tasks.
Inferred · not functionally tested
Audience: AI researchers and agent developers
Inferred · not functionally tested
Functions: analytics, monitoring
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI, WEB · deployment: browser, cli, cloud_managed
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: tbench.ai · github.com. These links do not verify the individual claims.
In the Agent evaluation & testing space, Terminal-Bench takes a focused approach. Inferred · not functionally tested: It focuses on measuring and comparing the performance of AI agents in terminal-based tasks. Inferred · not functionally tested: Terminal-Bench is an open-source project aimed at AI researchers and agent developers. Basis unknown · not verified: The project is open source (Apache-2.0). Basis unknown · not verified: It ships for the web and the command line.
Behind Terminal-Bench is Nicholas Carlini, and it first shipped in 2025. Development happens publicly on GitHub with 2.7k stars and 496 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, leaderboard, and task examples.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Agent benchmarking
- Leaderboard
- Task examples
- Terminal challenges
- Performance metrics
Topics: Inferred · not functionally tested
Built with & integrations
- Next.js
- x-nextjs-prerender header · x-powered-by header · /_next/static/ in the HTML
- openai
- bgpt- in the HTML
- Netlify
- x-nf-request-id header
- multiple
- bLangChain in the HTML
- anthropic
- bclaude in the HTML · bclaude- in the HTML
- google_gemini
- bgemini- in the HTML
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about Terminal-Bench
- What does Terminal-Bench do?
- Inferred · not functionally tested: Terminal-Bench focuses on measuring and comparing the performance of AI agents in terminal-based tasks. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is Terminal-Bench for?
- Inferred · not functionally tested: Terminal-Bench is an open-source project built for AI researchers and agent developers.
- Does Terminal-Bench have a free plan?
- Basis unknown · not verified: Yes — Terminal-Bench is open source under the Apache-2.0 license and free to use.
- What platforms does Terminal-Bench run on?
- Basis unknown · not verified: Terminal-Bench runs on the web and the command line.
- Is Terminal-Bench still maintained?
- PulseGate's liveness check found it on 3 Oct 2026. Its GitHub repository shows 496 commits in the last 90 days.
- What are alternatives to Terminal-Bench?
- Similar projects tracked by PulseGate include Terminal-Bench-Science 0.1, agentbench-cli, and CatchBench.Terminal-Bench-Science 0.1agentbench-cliCatchBench
- Who makes Terminal-Bench?
- Terminal-Bench is developed by Nicholas Carlini.
- When did Terminal-Bench launch?
- Terminal-Bench first shipped in 2025.
Similar projects
Closest matches by what these projects do