agent-skill-eval
PulseGate's liveness check found it on 13 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 30 Jun 2026. How this is checked
agent-skill-eval is an open-source CLI framework for evaluating the skills of code-generating agents across models like OpenCode, Claude Code, and Codex. It enables researchers and developers to benchmark agent performance using standardized tests.
Inferred · not functionally tested
Overview
4 featuresPurpose: Providing a standardized way to evaluate and benchmark agent skills across different LLM code models.
Inferred · not functionally tested
Audience: AI researchers and developers
Inferred · not functionally tested
Functions: code_generation
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
In the Agent evaluation & testing space, agent-skill-eval takes a focused approach. Inferred · not functionally tested: It focuses on providing a standardized way to evaluate and benchmark agent skills across different LLM code models. Inferred · not functionally tested: agent-skill-eval is an open-source project aimed at AI researchers and developers. Basis unknown · not verified: agent-skill-eval is open source under the MIT license. Basis unknown · not verified: agent-skill-eval is available on the command line.
Behind agent-skill-eval is tardigrde, and it first shipped in 2026. Development happens publicly on GitHub with 53 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent benchmarking, skill evaluation, and LLM support.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Agent benchmarking
- Skill evaluation
- LLM support
- Command-line interface
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about agent-skill-eval
- What does agent-skill-eval do?
- Inferred · not functionally tested: Agent-skill-eval focuses on providing a standardized way to evaluate and benchmark agent skills across different LLM code models. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is agent-skill-eval for?
- Inferred · not functionally tested: agent-skill-eval is an open-source project built for AI researchers and developers.
- Is agent-skill-eval free?
- Basis unknown · not verified: Yes — agent-skill-eval is open source under the MIT license and free to use.
- What platforms does agent-skill-eval run on?
- Basis unknown · not verified: agent-skill-eval runs on the command line.
- Is agent-skill-eval still active?
- PulseGate's liveness check found it on 13 Sep 2026. Its GitHub repository shows 53 commits in the last 90 days.
- What are alternatives to agent-skill-eval?
- Similar projects tracked by PulseGate include agent-exam, agent-skill-description-optimizer, and agent-evaluation-lab.agent-examagent-skill-description-optimizeragent-evaluation-lab
- Who makes agent-skill-eval?
- agent-skill-eval is developed by tardigrde.
- How long has agent-skill-eval been around?
- agent-skill-eval first shipped in 2026.
Similar projects
Closest matches by what these projects do