coder-eval
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 9 Jul 2026. How this is checked
coder-eval is an open-source command-line tool for evaluating, benchmarking, and A/B testing AI coding agents. It uses sandboxed, reproducible YAML task suites, making it suitable for AI researchers and developers assessing agent performance.
Inferred · not functionally tested
Overview
5 featuresPurpose: Providing reproducible evaluation and benchmarking of AI coding agents using standardized task suites.
Inferred · not functionally tested
Audience: AI researchers
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
coder-eval is an Agent evaluation & testing project. Inferred · not functionally tested: It focuses on providing reproducible evaluation and benchmarking of AI coding agents using standardized task suites. Inferred · not functionally tested: coder-eval is an open-source project aimed at AI researchers. Basis unknown · not verified: The project is open source (Apache-2.0). Basis unknown · not verified: It ships for the command line, and it can be self-hosted.
Behind coder-eval is UiPath, and it first shipped in 2026. Development happens publicly on GitHub with 5 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent evaluation, benchmarking, and A/B testing.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Agent evaluation
- Benchmarking
- A/B testing
- YAML task suites
- Sandboxed execution
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
4What PulseGate has recorded for this listing
Frequently asked questions about coder-eval
- What does coder-eval do?
- Inferred · not functionally tested: Coder-eval focuses on providing reproducible evaluation and benchmarking of AI coding agents using standardized task suites. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is coder-eval for?
- Inferred · not functionally tested: coder-eval is an open-source project built for AI researchers.
- Is coder-eval free?
- Basis unknown · not verified: Yes — coder-eval is open source under the Apache-2.0 license and free to use.
- What platforms does coder-eval run on?
- Basis unknown · not verified: coder-eval runs on the command line. It can also be self-hosted.
- Is coder-eval still maintained?
- PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 5 commits in the last 90 days.
- What are alternatives to coder-eval?
- Similar projects tracked by PulseGate include Coder Eval, caliper-eval, and agent-evaluation-lab.Coder Evalcaliper-evalagent-evaluation-lab
- Who makes coder-eval?
- coder-eval is developed by UiPath.
- How long has coder-eval been around?
- coder-eval first shipped in 2026.
Similar projects
Closest matches by what these projects do