agent-exam is an Apache-2.0 evaluation framework for testing agent skills across Claude Code, Codex CLI, Copilot CLI, and OpenCode. It is intended for developers building and benchmarking agent workflows.
In the LLM eval & observability space, agent-exam takes a focused approach. It focuses on evaluating and comparing agent skills consistently across multiple coding-agent environments. It is built as an open-source project for developers building and evaluating AI agent workflows. The project is open source (Apache-2.0). agent-exam is available on the web and the command line, and it can be self-hosted.
It is developed by Zyte Data, and it first shipped in 2026. Development happens publicly on GitHub with 10 commits in the last 90 days. Among its 5 catalogued features are agent skill evaluation, cross-agent testing, and CLI evaluation.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do