CheatBench
PulseGate's liveness check found it on 24 Sep 2026; it is registered on GitHub and has been in the index since 24 Sep 2026. How this is checked
CheatBench evaluates whether AI agents attempt to cheat when honest work is difficult. It provides benchmark environments spanning coding, mathematics, visual tasks, and knowledge work, with tools for comparing cheating rates across models and agent harnesses.
Inferred · not functionally tested
Overview
6 featuresPurpose: Measuring whether AI agents exploit reward systems instead of completing tasks honestly.
Inferred · not functionally tested
Audience: AI safety researchers and developers evaluating agent behavior
Inferred · not functionally tested
Functions: agents
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: unknown · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: free · license: MIT · platforms: WEB · deployment: browser, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: cheatbench.ai · github.com. These links do not verify the individual claims.
CheatBench sits in PulseGate's LLM evaluation & benchmarks category. Inferred · not functionally tested: It focuses on measuring whether AI agents exploit reward systems instead of completing tasks honestly. Inferred · not functionally tested: CheatBench is an open-source project aimed at AI safety researchers and developers evaluating agent behavior. Basis unknown · not verified: CheatBench is free to use. Basis unknown · not verified: It runs on the web, and it can be self-hosted.
Behind CheatBench is Center for AI Safety, based in the United States, and it first shipped in 2026. The project is developed in the open on GitHub with 10 stars and 4 commits in the last 90 days. Inferred · not functionally tested: Among its 6 catalogued features are cheating rate charts, agent comparisons, and ten task categories.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Cheating rate charts
- Agent comparisons
- Ten task categories
- Code repository
- Research paper
- Author information
Topics: Inferred · not functionally tested
Built with & integrations
- Next.js
- /_next/static/ in the HTML · __next_f in the HTML
- openai
- bgpt- in the HTML
- Vercel
- x-vercel-id header · x-vercel-cache header
- anthropic
- bclaude in the HTML · bclaude- in the HTML
- google_gemini
- bgemini- in the HTML
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed24 Sep · 12:13 UTCCheatBench: Measuring Reward Gaming in AI Agents seen via Hacker News firehose (Algolia)Source: Hacker News firehose (Algolia) · Open
Frequently asked questions about CheatBench
- What does CheatBench do?
- Inferred · not functionally tested: CheatBench focuses on measuring whether AI agents exploit reward systems instead of completing tasks honestly. It is catalogued under LLM evaluation & benchmarks on PulseGate.
- Who is CheatBench for?
- Inferred · not functionally tested: CheatBench is an open-source project built for AI safety researchers and developers evaluating agent behavior.
- Is CheatBench free?
- Basis unknown · not verified: Yes — CheatBench is free to use.
- What platforms does CheatBench run on?
- Basis unknown · not verified: CheatBench runs on the web. It can also be self-hosted.
- Is CheatBench still maintained?
- PulseGate's liveness check found it on 24 Sep 2026. Its GitHub repository shows 4 commits in the last 90 days.
- What projects are similar to CheatBench?
- Similar projects tracked by PulseGate include CatchBench, CooperBench, and Bench'd.CatchBenchCooperBenchBench'd
- Who makes CheatBench?
- CheatBench is developed by Center for AI Safety, based in the United States.
- When did CheatBench launch?
- CheatBench first shipped in 2026.
Similar projects
Closest matches by what these projects do