agent-skill-eval is an open-source CLI framework for evaluating the skills of code-generating agents across models like OpenCode, Claude Code, and Codex. It enables researchers and developers to benchmark agent performance using standardized tests.
In the LLM eval & observability space, agent-skill-eval takes a focused approach. It focuses on providing a standardized way to evaluate and benchmark agent skills across different LLM code models. agent-skill-eval is an open-source project aimed at AI researchers and developers. agent-skill-eval is open source under the MIT license. agent-skill-eval is available on the command line.
Behind agent-skill-eval is tardigrde, and it first shipped in 2026. Development happens publicly on GitHub with 53 commits in the last 90 days. Key capabilities include agent benchmarking, skill evaluation, and LLM support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do