debatebench
PulseGate's liveness check found it on 17 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 17 Sep 2026. How this is checked
debatebench is an MIT-licensed Python package and CLI for running structured, multi-turn adversarial debates between language models. It scores debate outcomes against a fixed rubric for researchers and developers building or evaluating LLM systems.
Inferred · not functionally tested
Overview
5 featuresPurpose: Evaluating the quality and performance of adversarial multi-turn LLM debates against a consistent rubric.
Inferred · not functionally tested
Audience: AI researchers and developers evaluating language models
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
debatebench is a LLM evaluation & benchmarks project. Inferred · not functionally tested: It focuses on evaluating the quality and performance of adversarial multi-turn LLM debates against a consistent rubric. Inferred · not functionally tested: debatebench is an open-source project aimed at AI researchers and developers evaluating language models. Basis unknown · not verified: debatebench is open source under the MIT license. Basis unknown · not verified: debatebench is available on the command line, and it can be self-hosted.
Behind debatebench is Robert E. Roy, and it first shipped in 2026. The project is developed in the open on GitHub with 100 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include multi-turn debates, adversarial evaluation, and fixed rubric scoring.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Multi-turn debates
- Adversarial evaluation
- Fixed rubric scoring
- Benchmarking
- CLI interface
Topics: Inferred · not functionally tested
Built with & integrations
- Claude Code
- commit addd77519977 · since Sep 2026
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
Frequently asked questions about debatebench
- What does debatebench do?
- Inferred · not functionally tested: Debatebench focuses on evaluating the quality and performance of adversarial multi-turn LLM debates against a consistent rubric. It is catalogued under LLM evaluation & benchmarks on PulseGate.
- Who should use debatebench?
- Inferred · not functionally tested: debatebench is an open-source project built for AI researchers and developers evaluating language models.
- Is debatebench free?
- Basis unknown · not verified: Yes — debatebench is open source under the MIT license and free to use.
- What platforms does debatebench run on?
- Basis unknown · not verified: debatebench runs on the command line. It can also be self-hosted.
- Is debatebench still active?
- PulseGate's liveness check found it on 17 Sep 2026. Its GitHub repository shows 100 commits in the last 90 days.
- What projects are similar to debatebench?
- Similar projects tracked by PulseGate include porchbench, bench-my-llm, and LitigationBench.porchbenchbench-my-llmLitigationBench
- Who makes debatebench?
- debatebench is developed by Robert E. Roy.
- How long has debatebench been around?
- debatebench first shipped in 2026.
Similar projects
Closest matches by what these projects do