swe-duel is a Python package and command-line arena for adversarial evaluation of LLM coding agents. It provides contamination-resistant benchmark tasks for researchers and developers testing software-engineering agents.
swe-duel sits in PulseGate's Agent evaluation & testing category. It focuses on evaluating coding agents against adversarial tasks without benchmark contamination. It is built as an open-source project for AI researchers and developers evaluating coding agents. The project is open source (Open Source). It runs on the command line, and it can be self-hosted.
Weimin N. builds and maintains swe-duel, and it first shipped in 2026. Among its 6 catalogued features are adversarial tasks, coding agent arena, and contamination resistance.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match