Skip to content
Back to the index

agent-skill-eval

PyPIInfrastructure

PulseGate's liveness check found it on 13 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 30 Jun 2026. How this is checked

agent-skill-eval is an open-source CLI framework for evaluating the skills of code-generating agents across models like OpenCode, Claude Code, and Codex. It enables researchers and developers to benchmark agent performance using standardized tests.

Inferred · not functionally tested

Open SourceMITCLI
Visit PyPI

Overview

4 features

Purpose: Providing a standardized way to evaluate and benchmark agent skills across different LLM code models.

Inferred · not functionally tested

Audience: AI researchers and developers

Inferred · not functionally tested

Functions: code_generation

Inferred · not functionally tested

Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown

Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: pypi.org · github.com. These links do not verify the individual claims.

In the Agent evaluation & testing space, agent-skill-eval takes a focused approach. Inferred · not functionally tested: It focuses on providing a standardized way to evaluate and benchmark agent skills across different LLM code models. Inferred · not functionally tested: agent-skill-eval is an open-source project aimed at AI researchers and developers. Basis unknown · not verified: agent-skill-eval is open source under the MIT license. Basis unknown · not verified: agent-skill-eval is available on the command line.

Behind agent-skill-eval is tardigrde, and it first shipped in 2026. Development happens publicly on GitHub with 53 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, skill evaluation, and LLM support.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Agent benchmarking
  • Skill evaluation
  • LLM support
  • Command-line interface

Topics: Inferred · not functionally tested

Tags
agent-evaluationllm-benchmarkingcode-eval
AI capabilities
Code

JSON profile · Text profile · Access guide

Built with & integrations

AI providers
openai
Runs on
CLI

Trust & compliance

License
MIT
Public signals
HTTPSOpen SourceFree tierGitHubActive maintenance

Indexing history

What PulseGate has recorded for this listing

Nothing recorded for this listing in this window.

Frequently asked questions about agent-skill-eval

What does agent-skill-eval do?
Inferred · not functionally tested: Agent-skill-eval focuses on providing a standardized way to evaluate and benchmark agent skills across different LLM code models. It is catalogued under Agent evaluation & testing on PulseGate.
Who is agent-skill-eval for?
Inferred · not functionally tested: agent-skill-eval is an open-source project built for AI researchers and developers.
Is agent-skill-eval free?
Basis unknown · not verified: Yes — agent-skill-eval is open source under the MIT license and free to use.
What platforms does agent-skill-eval run on?
Basis unknown · not verified: agent-skill-eval runs on the command line.
Is agent-skill-eval still active?
PulseGate's liveness check found it on 13 Sep 2026. Its GitHub repository shows 53 commits in the last 90 days.
What are alternatives to agent-skill-eval?
Similar projects tracked by PulseGate include agent-exam, agent-skill-description-optimizer, and agent-evaluation-lab.agent-examagent-skill-description-optimizeragent-evaluation-lab
Who makes agent-skill-eval?
agent-skill-eval is developed by tardigrde.
How long has agent-skill-eval been around?
agent-skill-eval first shipped in 2026.

Similar projects

Closest matches by what these projects do