Contract-based testing framework for LLM agents — hybrid metrics (objective + LLM-as-judge + human), multi-agent and A2A protocol support, deterministic record/replay envelope.
agentanvil is a LLM eval & observability project. It focuses on testing and evaluating LLM agents with objective, LLM-based, and human metrics in multi-agent environments. agentanvil is an open-source project aimed at AI developers and researchers. The project is open source (MIT). agentanvil is available on the web and the command line, and it can be self-hosted.
Behind agentanvil is cchinchilla-dev, and it first shipped in 2026. Development happens publicly on GitHub with 33 commits in the last 90 days. Among its 5 catalogued features are contract-based testing, hybrid metrics, and multi-agent support.
What PulseGate has recorded for this listing
Closest matches by what these projects do