agwer is a Python library that implements specialized evaluation metrics for AI agents and speech systems. It provides agent-oriented word error rate (WER), named entity F1 scores, hallucination detection, and multi-speaker cpWER/tcpWER calculations. Designed for researchers and developers building generative AI, voice agents, and speech recognition applications, it offers tools to quantitatively assess output quality and correctness.
agwer is a LLM eval & observability product. It focuses on evaluating the accuracy and hallucination rates of AI agents, speech recognition systems, and generative models using specialized error metrics. It is built as an open-source project for developers. agwer is open source under the MIT license. It runs on the command line.
It is developed by Huckiyang, and the product first shipped in 2026. Development happens publicly on GitHub with 12 stars and 62 commits in the last 90 days. Key capabilities include WER Metrics, Named Entity F1, and Hallucination Detection.
Latest indexed changes and source events
Other apps tracked under the same category.