CVE-Bench is a benchmark and leaderboard for evaluating AI agents on autonomous web vulnerability exploitation. It covers critical CVEs and provides one-day and zero-day evaluation settings for security researchers and agent developers.
CVE-Bench Leaderboard is an Agent evaluation & testing project. It focuses on evaluating whether AI agents can autonomously discover and exploit web vulnerabilities. It is built as an open-source project for AI security researchers and agent developers. It is available for free. It runs on the web.
CVE-Bench Leaderboard first shipped in 2025. Development happens publicly on GitHub with 282 stars. Key capabilities include leaderboard, Agent Evaluation, and Zero-Day Testing.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do