MazeBench
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 7 Jul 2026. How this is checked
MazeBench is an open-source tool for benchmarking coding agents through interactive ice-maze puzzles. It includes a web-based puzzle site, a world editor for creating custom mazes, and a CLI for running benchmarks. It is designed for AI researchers and developers working on agent evaluation.
Inferred · not functionally tested
Overview
5 featuresPurpose: Benchmarking and evaluating coding agents using interactive maze puzzles and custom environments.
Inferred · not functionally tested
Audience: ai researchers
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI, WEB · deployment: browser, cli
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: mazebench.com · github.com. These links do not verify the individual claims.
In the Agent evaluation & testing space, MazeBench takes a focused approach. Inferred · not functionally tested: It focuses on benchmarking and evaluating coding agents using interactive maze puzzles and custom environments. Inferred · not functionally tested: It is built as an open-source project for ai researchers. Basis unknown · not verified: The project is open source (MIT). Basis unknown · not verified: It runs on the web and the command line.
MazeBench first shipped in 2026. Development happens publicly on GitHub with 348 commits in the last 90 days. Inferred · not functionally tested: Among its 5 catalogued features are puzzle site, world editor, and benchmark runner.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Puzzle site
- World editor
- Benchmark runner
- Coding agent evaluation
- Local web server
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about MazeBench
- What does MazeBench do?
- Inferred · not functionally tested: MazeBench focuses on benchmarking and evaluating coding agents using interactive maze puzzles and custom environments. It is catalogued under Agent evaluation & testing on PulseGate.
- Who should use MazeBench?
- Inferred · not functionally tested: MazeBench is an open-source project built for ai researchers.
- Does MazeBench have a free plan?
- Basis unknown · not verified: Yes — MazeBench is open source under the MIT license and free to use.
- What platforms does MazeBench run on?
- Basis unknown · not verified: MazeBench runs on the web and the command line.
- Is MazeBench still maintained?
- PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 348 commits in the last 90 days.
- When did MazeBench launch?
- MazeBench first shipped in 2026.
- Is MazeBench open source?
- Basis unknown · not verified: Yes — MazeBench is open source under the MIT license, developed on GitHub.
Similar projects
Closest matches by what these projects do