rubric-eval is an open-source CLI tool for testing and evaluating the behavior of agents in LLM applications. It auto-captures agent runs, integrates with evaluation tools, and helps developers catch regressions in CI pipelines. The tool is designed for developers working with AI agents and LLMs.
In the Agent evaluation & testing space, rubric-eval takes a focused approach. It focuses on testing and evaluating agent behavior in LLM-powered applications. It is built as an open-source project for developers building and testing LLM agent apps. The project is open source (MIT). rubric-eval is available on the command line.
It is developed by Kareem-Rashed, and it first shipped in 2026. The project is developed in the open on GitHub with 13 stars and 5 commits in the last 90 days. Key capabilities include agent run capture, eval tool integration, and regression testing. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do