ClawBench
PulseGate's liveness check found it on 3 Oct 2026; it is registered on GitHub and has been in the index since 26 Aug 2026. How this is checked
ClawBench is an open benchmark for evaluating AI browser agents on live websites and everyday online tasks such as booking travel, ordering food, and applying for jobs. It uses HTTP-request interception and LLM judging to score task completion, with leaderboard results, traces, and a downloadable dataset for researchers and developers.
Inferred · not functionally tested
Overview
6 featuresPurpose: Measuring whether AI browser agents can complete and submit real-world online tasks correctly.
Inferred · not functionally tested
Audience: AI researchers and agent developers
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI, WEB · deployment: browser, cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: claw-bench.com · github.com. These links do not verify the individual claims.
ClawBench sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on measuring whether AI browser agents can complete and submit real-world online tasks correctly. Inferred · not functionally tested: ClawBench is an open-source project aimed at AI researchers and agent developers. Basis unknown · not verified: ClawBench is open source under the Apache-2.0 license. Basis unknown · not verified: It ships for the web and the command line, and it can be self-hosted.
It is developed by TIGER-AI-Lab, and it first shipped in 2026. The project is developed in the open on GitHub with 589 stars and 138 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include Live Leaderboard, Task Benchmarking, and HTTP Interception.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Live Leaderboard
- Task Benchmarking
- HTTP Interception
- LLM Judging
- Agent Traces
- Model Comparison
Topics: Inferred · not functionally tested
Built with & integrations
- openai
- bgpt- in the HTML
- multiple
- bopenrouter in the HTML
- anthropic
- bclaude in the HTML · bclaude- in the HTML
- local_oss
- huggingface in the HTML
- Cloudflare
- cf-ray header · cf-cache-status header
- google_gemini
- bgemini- in the HTML
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed26 Aug · 17:57 UTCTIGER-AI-Lab/ClawBench seen via GitHub Search (Smart)Source: GitHub Search (Smart) · Open
Frequently asked questions about ClawBench
- What does ClawBench do?
- Inferred · not functionally tested: ClawBench focuses on measuring whether AI browser agents can complete and submit real-world online tasks correctly. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is ClawBench for?
- Inferred · not functionally tested: ClawBench is an open-source project built for AI researchers and agent developers.
- Is ClawBench free?
- Basis unknown · not verified: Yes — ClawBench is open source under the Apache-2.0 license and free to use.
- What platforms does ClawBench run on?
- Basis unknown · not verified: ClawBench runs on the web and the command line. It can also be self-hosted.
- Is ClawBench still maintained?
- PulseGate's liveness check found it on 3 Oct 2026. Its GitHub repository shows 138 commits in the last 90 days.
- What are alternatives to ClawBench?
- Similar projects tracked by PulseGate include Clawbotomy, Terminal-Bench, and ClawResearch.ClawbotomyTerminal-BenchClawResearch
- Who develops ClawBench?
- ClawBench is developed by TIGER-AI-Lab.
- How long has ClawBench been around?
- ClawBench first shipped in 2026.
Similar projects
Closest matches by what these projects do