dig.bench is a benchmark for evaluating human and AI-agent scientific discovery through 70 text-based games with undisclosed rules. It provides a shared game interface, difficulty tiers, leaderboards, public API access, and research documentation for comparing discovery performance.
Bench sits in PulseGate's Agent evaluation & testing category. It focuses on measuring whether humans and AI agents can discover unknown rules through controlled interactive experiments. Bench is an open-source project aimed at AI researchers and developers evaluating language-model agents. Bench is open source under the MIT license. It runs on the web and API.
Bench first shipped in 2026. The project is developed in the open on GitHub with 15 stars and 3 commits in the last 90 days. Key capabilities include interactive games, game leaderboard, and difficulty tiers. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match