caliper-eval is an open-source CLI tool designed for evaluating Claude Code skills and AI agents. It provides developers and researchers with tools to assess and benchmark AI agent performance from the command line.
caliper-eval is a LLM eval & observability project. Enabling developers to evaluate and benchmark Claude Code skills and AI agents via the command line. caliper-eval is an open-source project aimed at AI developers and researchers. The project is open source (MIT). caliper-eval is available on the command line.
edonadei builds and maintains caliper-eval, and it first shipped in 2026. The project is developed in the open on GitHub with 23 stars and 108 commits in the last 90 days. Among its 5 catalogued features are AI evaluation, Claude Code skills, and agent testing.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do