Skip to content
Back to the index

agentanvil

github.comInfrastructure

PulseGate's liveness check found it on 9 Oct 2026; it is registered on GitHub and PyPI and has been in the index since 16 Jun 2026. How this is checked

Contract-based testing framework for LLM agents — hybrid metrics (objective + LLM-as-judge + human), multi-agent and A2A protocol support, deterministic record/replay envelope.

Inferred · not functionally tested

Open SourceMITWebCLISelf-hosted
agentanvil preview
Visit github.com

Overview

5 features

Purpose: Testing and evaluating LLM agents with objective, LLM-based, and human metrics in multi-agent environments.

Inferred · not functionally tested

Audience: AI developers and researchers

Inferred · not functionally tested

Functions: monitoring, agents

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: browser, cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: github.com. These links do not verify the individual claims.

agentanvil is an Agent evaluation & testing project. Inferred · not functionally tested: It focuses on testing and evaluating LLM agents with objective, LLM-based, and human metrics in multi-agent environments. Inferred · not functionally tested: agentanvil is an open-source project aimed at AI developers and researchers. Basis unknown · not verified: The project is open source (MIT). Basis unknown · not verified: agentanvil is available on the web and the command line, and it can be self-hosted.

Behind agentanvil is cchinchilla-dev, and it first shipped in 2026. Development happens publicly on GitHub with 33 commits in the last 90 days. Inferred · not functionally tested: Among its 5 catalogued features are contract-based testing, hybrid metrics, and multi-agent support.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Contract-based testing
  • Hybrid metrics
  • Multi-agent support
  • A2A protocol support
  • Deterministic record/replay

Topics: Inferred · not functionally tested

Tags
llm-testingmulti-agenta2a-protocolevaluation-frameworkopen-source

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
BrowserCLISelf-hosted

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tierGitHubActive maintenance

Indexing history

What PulseGate has recorded for this listing

Nothing recorded for this listing in this window.

Frequently asked questions about agentanvil

What is agentanvil?
Inferred · not functionally tested: Agentanvil focuses on testing and evaluating LLM agents with objective, LLM-based, and human metrics in multi-agent environments. It is catalogued under Agent evaluation & testing on PulseGate.
Who should use agentanvil?
Inferred · not functionally tested: agentanvil is an open-source project built for AI developers and researchers.
Does agentanvil have a free plan?
Basis unknown · not verified: Yes — agentanvil is open source under the MIT license and free to use.
What platforms does agentanvil run on?
Basis unknown · not verified: agentanvil runs on the web and the command line. It can also be self-hosted.
Is agentanvil still active?
PulseGate's liveness check found it on 9 Oct 2026. Its GitHub repository shows 33 commits in the last 90 days.
What projects are similar to agentanvil?
Similar projects tracked by PulseGate include agentlint-dev, agent-belt, and agent-eval.agentlint-devagent-beltagent-eval
Who develops agentanvil?
agentanvil is developed by cchinchilla-dev.
How long has agentanvil been around?
agentanvil first shipped in 2026.

Similar projects

Closest matches by what these projects do