One-Eval is an open-source, agent-based evaluation framework for large language models, enabling users to automate evaluation workflows and generate reports from natural language requirements. It supports traceable, interruptible, and scalable evaluation loops, and integrates with platforms like Claude Code. Designed for AI researchers and developers, it streamlines LLM evaluation processes.
One-Eval sits in PulseGate's LLM eval & observability category. It focuses on automating and simplifying the evaluation of large language models using agent-based workflows. One-Eval is an open-source project aimed at AI researchers and developers. The project is open source (Apache-2.0). One-Eval is available on the web and the command line, and it can be self-hosted.
Behind One-Eval is OpenDCAI, based in China, and the product first shipped in 2025. The project is developed in the open on GitHub with 150 stars and 21 commits in the last 90 days. Among its 8 catalogued features are automated evaluation, agent orchestration, and natural language input.
Latest indexed changes and source events
OpenDCAI/One-Eval verified by the PulseGate indexer
Other apps tracked under the same category.