SlopCodeBench
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and has been in the index since 28 Jul 2026. How this is checked
SlopCodeBench is a community benchmark that evaluates coding agents through repeated requirement changes and extensions. Each problem consists of a sequence of checkpoints where the agent implements an initial version and then extends its own solution as new requirements arrive. It uses black-box evaluation based only on a CLI or API contract, with no prescribed architecture, allowing assessment of real-world software development practices with AI agents.
Inferred · not functionally tested
Overview
5 featuresPurpose: Evaluating how well AI coding agents maintain code quality when iteratively extending their own solutions under changing requirements.
Inferred · not functionally tested
Audience: AI researchers and developers
Inferred · not functionally tested
Functions: code_generation
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: free · license: MIT · platforms: API, CLI, WEB · deployment: browser, cli, api_only
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: scbench.ai · github.com. These links do not verify the individual claims.
SlopCodeBench sits in PulseGate's LLM evaluation & benchmarks category. Inferred · not functionally tested: It focuses on evaluating how well AI coding agents maintain code quality when iteratively extending their own solutions under changing requirements. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers. Basis unknown · not verified: SlopCodeBench costs nothing to use. Basis unknown · not verified: It ships for the web, the command line, and API.
SlopCodeBench Community builds and maintains SlopCodeBench, and it first shipped in 2025. The project is developed in the open on GitHub with 99 stars and 5 commits in the last 90 days. Inferred · not functionally tested: Among its 5 catalogued features are leaderboard, Problem Checkpoints, and Erosion Metrics. Inferred · not functionally tested: Catalogued interfaces include a public API.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Leaderboard
- Problem Checkpoints
- Erosion Metrics
- Model Comparison
- Black-box Evaluation
Topics: Inferred · not functionally tested
Built with & integrations
- Next.js
- /_next/static/ in the HTML · __next_f in the HTML
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed28 Jul · 01:36 UTCSlopCodeBench seen via Hacker News firehose (Algolia)Source: Hacker News firehose (Algolia) · Open
Frequently asked questions about SlopCodeBench
- What is SlopCodeBench?
- Inferred · not functionally tested: SlopCodeBench focuses on evaluating how well AI coding agents maintain code quality when iteratively extending their own solutions under changing requirements. It is catalogued under LLM evaluation & benchmarks on PulseGate.
- Who should use SlopCodeBench?
- Inferred · not functionally tested: SlopCodeBench is an open-source project built for AI researchers and developers.
- Does SlopCodeBench have a free plan?
- Basis unknown · not verified: Yes — SlopCodeBench is free to use.
- What platforms does SlopCodeBench run on?
- Basis unknown · not verified: SlopCodeBench runs on the web, the command line, and API.
- Is SlopCodeBench still active?
- PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 5 commits in the last 90 days.
- What projects are similar to SlopCodeBench?
- Similar projects tracked by PulseGate include slopscore, slopcount, and slop-o-meter.slopscoreslopcountslop-o-meter
- Who develops SlopCodeBench?
- SlopCodeBench is developed by SlopCodeBench Community.
- When did SlopCodeBench launch?
- SlopCodeBench first shipped in 2025.
Similar projects
Closest matches by what these projects do