SlopCodeBench is a community benchmark that evaluates coding agents through repeated requirement changes and extensions. Each problem consists of a sequence of checkpoints where the agent implements an initial version and then extends its own solution as new requirements arrive. It uses black-box evaluation based only on a CLI or API contract, with no prescribed architecture, allowing assessment of real-world software development practices with AI agents.
In the LLM eval & observability space, SlopCodeBench takes a focused approach. It focuses on evaluating how well AI coding agents maintain code quality when iteratively extending their own solutions under changing requirements. It is built as an open-source project for AI researchers and developers. SlopCodeBench is free to use. SlopCodeBench is available on the web, the command line, and API.
Behind SlopCodeBench is SlopCodeBench Community, and the product first shipped in 2025. Development happens publicly on GitHub with 99 stars and 5 commits in the last 90 days. Key capabilities include leaderboard, Problem Checkpoints, and Erosion Metrics. It exposes integrations via a public API.
Latest indexed changes and source events
SlopCodeBench verified by the PulseGate indexer
Other apps tracked under the same category.