Skip to content
Back to the index

modelduel

PyPIInfrastructure

PulseGate's liveness check found it on 1 Oct 2026; it is registered on GitHub and PyPI and has been in the index since 1 Oct 2026. How this is checked

modelduel is an open-source command-line tool for comparing AI models on coding tasks. It runs the same pytest-based tests against multiple models to support repeatable evaluation.

Inferred · not functionally tested

Open SourceMITCLISelf-hosted
Visit PyPI

Overview

5 features

Purpose: Comparing AI models consistently on coding tasks with the same test suite.

Inferred · not functionally tested

Audience: developers and AI evaluators

Inferred · not functionally tested

Functions: code_generation, analytics

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli, self_hosted

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org · github.com. These links do not verify the individual claims.

modelduel sits in PulseGate's LLM evaluation & benchmarks category. Inferred · not functionally tested: It focuses on comparing AI models consistently on coding tasks with the same test suite. Inferred · not functionally tested: It is built as an open-source project for developers and AI evaluators. Basis unknown · not verified: The project is open source (MIT). Basis unknown · not verified: It ships for the command line, and it can be self-hosted.

It is developed by BertMarti, and it first shipped in 2026. The project is developed in the open on GitHub with 104 commits in the last 90 days. Inferred · not functionally tested: Among its 5 catalogued features are model comparison, coding benchmarks, and pytest tests.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Model comparison
  • Coding benchmarks
  • Pytest tests
  • CLI workflow
  • Repeatable evaluation

Topics: Inferred · not functionally tested

Tags
llm-benchmarkingcoding-evaluationpytest-testingmodel-comparison
AI capabilities
Code
Inference: Cloud API

JSON profile · Text profile · Access guide

Built with & integrations

Written with
Claude Code
Runs on
CLISelf-hosted
Written with — evidence
Claude Code
commit 286bbc449e96 · since Sep 2026

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceGitHubActive maintenance

Indexing history

2

What PulseGate has recorded for this listing

  1. Indexed1 Oct · 20:41 UTC
    modelduel seen via PyPI Fresh Feed
    Source: PyPI Fresh Feed · Open
  2. Indexed1 Oct · 17:29 UTC
    modelduel seen via PyPI Bulk Enumerator
    Source: PyPI Bulk Enumerator · Open

Frequently asked questions about modelduel

What does modelduel do?
Inferred · not functionally tested: Modelduel focuses on comparing AI models consistently on coding tasks with the same test suite. It is catalogued under LLM evaluation & benchmarks on PulseGate.
Who should use modelduel?
Inferred · not functionally tested: modelduel is an open-source project built for developers and AI evaluators.
Does modelduel have a free plan?
Basis unknown · not verified: Yes — modelduel is open source under the MIT license and free to use.
What platforms does modelduel run on?
Basis unknown · not verified: modelduel runs on the command line. It can also be self-hosted.
Is modelduel still maintained?
PulseGate's liveness check found it on 1 Oct 2026. Its GitHub repository shows 104 commits in the last 90 days.
What projects are similar to modelduel?
Similar projects tracked by PulseGate include contamcheck, quantdiff, and rerun-bench.contamcheckquantdiffrerun-bench
Who develops modelduel?
modelduel is developed by BertMarti.
How long has modelduel been around?
modelduel first shipped in 2026.

Also in LLM evaluation & benchmarks

Same category — not a similarity match