AgentRisk is a platform that aggregates and analyzes verifiable behavior records from millions of AI agents across platforms like Hugging Face, GitHub, and others. It provides trust proofs, categorizes failures, measures deception rates, and offers per-platform health dashboards. Users can query specific agents or browse by category to understand risks such as code deletion, runaway costs, or misleading outputs. All records are anchored with hash chains for independent verification.
In the LLM eval & observability space, AgentRisk takes a focused approach. It focuses on evaluating and verifying the reliability and past behavior of AI agents before trusting them in production workflows. AgentRisk is a B2B product aimed at AI developers and agent operators. A free plan is available. It ships for the web and API.
AgentRisk builds and maintains AgentRisk, and it first shipped in 2025. Key capabilities include Agent Trust Scoring, Failure Archive, and Platform Comparison. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do