Agents' Last Exam
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and has been in the index since 1 Jul 2026. How this is checked
Agents' Last Exam is a large-scale benchmark designed to evaluate AI agents on real-world, economically valuable professional workflows. The platform focuses on assessing agent performance on long-horizon tasks with verifiable outcomes, aiming to provide objective and comparable scores across a broad range of domains. It is led by Berkeley RDI in collaboration with over 300 industry experts and spans 55 sub-industries that encompass most major fields of professional work…
Inferred · not functionally tested
Overview
5 featuresPurpose: Provides a standardized way to evaluate and compare AI agents on complex, real-world tasks.
Inferred · not functionally tested
Audience: AI researchers and developers
Inferred · not functionally tested
Functions: analytics, monitoring
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: unknown · Self-hosting: unknown
Recorded constraints: pricing: free · license: Apache-2.0 · platforms: WEB · deployment: browser
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: agents-last-exam.org · github.com. These links do not verify the individual claims.
Agents' Last Exam sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It provides a standardized way to evaluate and compare AI agents on complex, real-world tasks. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers. Basis unknown · not verified: Agents' Last Exam costs nothing to use. Basis unknown · not verified: It ships for the web.
Berkeley RDI and contributors builds and maintains Agents' Last Exam, and it first shipped in 2026. Development happens publicly on GitHub with 759 stars and 87 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, leaderboard, and task repository.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Agent benchmarking
- Leaderboard
- Task repository
- Cross-domain evaluation
- Open contributions
Topics: Inferred · not functionally tested
Built with & integrations
- Next.js
- x-nextjs-prerender header · /_next/static/ in the HTML · __next_f in the HTML
- Vercel
- x-vercel-id header · x-vercel-cache header
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about Agents' Last Exam
- What is Agents' Last Exam?
- Inferred · not functionally tested: Agents' Last Exam provides a standardized way to evaluate and compare AI agents on complex, real-world tasks. It is catalogued under Agent evaluation & testing on PulseGate.
- Who should use Agents' Last Exam?
- Inferred · not functionally tested: Agents' Last Exam is an open-source project built for AI researchers and developers.
- Does Agents' Last Exam have a free plan?
- Basis unknown · not verified: Yes — Agents' Last Exam is free to use.
- What platforms does Agents' Last Exam run on?
- Basis unknown · not verified: Agents' Last Exam runs on the web.
- Is Agents' Last Exam still maintained?
- PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 87 commits in the last 90 days.
- What projects are similar to Agents' Last Exam?
- Similar projects tracked by PulseGate include AIAgentBenchmark, agent-exam, and Agents School.AIAgentBenchmarkagent-examAgents School
- Who develops Agents' Last Exam?
- Agents' Last Exam is developed by Berkeley RDI and contributors, based in the United States.
- How long has Agents' Last Exam been around?
- Agents' Last Exam first shipped in 2026.
Similar projects
Closest matches by what these projects do