proofrag is an open-source command-line tool for evaluating RAG applications and agents against documentation. It generates golden test sets and produces LLM-as-judge and retrieval scorecards for developers.
proofrag is an Agent evaluation & testing project. It focuses on evaluating RAG retrieval quality and agent answers without manually creating test sets or scorecards. It is built as an open-source project for AI and RAG developers. The project is open source (MIT). It runs on the command line, and it can be self-hosted.
unshDee builds and maintains proofrag, and it first shipped in 2026. The project is developed in the open on GitHub with 32 commits in the last 90 days. Key capabilities include golden test sets, LLM-as-judge, and retrieval scoring.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do