aws-bench
PulseGate's liveness check found it on 5 Oct 2026; it has been in the index since 5 Sep 2026. How this is checked
aws-bench is an open-source benchmark for measuring AI coding agents and model combinations on real AWS tasks. It evaluates capabilities such as diagnosing cloud misconfigurations, provisioning infrastructure, and operating live environments for researchers and engineering teams.
Inferred · not functionally tested
Overview
6 featuresPurpose: Evaluating whether AI coding agents can reliably diagnose, provision, and operate real AWS environments.
Inferred · not functionally tested
Audience: AI researchers and developers evaluating coding agents
Inferred · not functionally tested
Functions: agents, analytics
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Open Source · platforms: WEB · deployment: browser, self_hosted, cli
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: github.com. These links do not verify the individual claims.
In the Agent evaluation & testing space, aws-bench takes a focused approach. Inferred · not functionally tested: It focuses on evaluating whether AI coding agents can reliably diagnose, provision, and operate real AWS environments. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers evaluating coding agents. Basis unknown · not verified: aws-bench is open source under the Open Source license. Basis unknown · not verified: It ships for the web and the command line, and it can be self-hosted.
aws-bench contributors builds and maintains aws-bench. Inferred · not functionally tested: Among its 6 catalogued features are AWS task evaluation, agent benchmarking, and model comparison. Inferred · not functionally tested: Catalogued interfaces include a public API.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- AWS task evaluation
- Agent benchmarking
- Model comparison
- Misconfiguration diagnosis
- Infrastructure provisioning
- Live environment operations
Topics: Inferred · not functionally tested
Built with & integrations
- anthropic
- bclaude in the HTML · bclaude- in the HTML
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed5 Sep · 00:56 UTCAWS-bench: Benchmark for evaluating AI coding agents on real-world AWS tasks seen via Hacker News firehose (Algolia)Source: Hacker News firehose (Algolia) · Open
Frequently asked questions about aws-bench
- What is aws-bench?
- Inferred · not functionally tested: Aws-bench focuses on evaluating whether AI coding agents can reliably diagnose, provision, and operate real AWS environments. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is aws-bench for?
- Inferred · not functionally tested: aws-bench is an open-source project built for AI researchers and developers evaluating coding agents.
- Is aws-bench free?
- Basis unknown · not verified: Yes — aws-bench is open source under the Open Source license and free to use.
- What platforms does aws-bench run on?
- Basis unknown · not verified: aws-bench runs on the web and the command line. It can also be self-hosted.
- Is aws-bench still active?
- PulseGate's liveness check found it on 5 Oct 2026.
- What projects are similar to aws-bench?
- Similar projects tracked by PulseGate include CooperBench, CatchBench, and AIAgentBenchmark.CooperBenchCatchBenchAIAgentBenchmark
- Who develops aws-bench?
- aws-bench is developed by aws-bench contributors.
- Is aws-bench open source?
- Basis unknown · not verified: Yes — aws-bench is open source under the Open Source license.
Similar projects
Closest matches by what these projects do