ciagent is an open-source Python library that analyzes the reliability of AI agent evaluations. It generates stability reports with flip attribution, performs LLM judge audits, and runs deterministic checks derived from a knowledge base to identify when evals may be misleading. Primarily used by AI developers and researchers building autonomous agents.
In the LLM eval & observability space, ciagent takes a focused approach. It focuses on detecting when AI agent evaluation metrics are unreliable or misleading due to instability or judge errors. It is built as an open-source project for developers. ciagent is open source under the Apache-2.0 license. ciagent is available on the command line and API, and it can be self-hosted.
Suniel builds and maintains ciagent, and the product first shipped in 2026. Development happens publicly on GitHub with 98 commits in the last 90 days. Key capabilities include Stability Reports, Flip Attribution, and LLM Judge Audits. It exposes integrations via a public API.
Latest indexed changes and source events
Other apps tracked under the same category.