skill-eval-harness
PulseGate's liveness check found it on 14 Sep 2026; it is registered on GitHub and PyPI and has been in the index since 8 Jul 2026. How this is checked
skill-eval-harness is an open-source CLI tool for evaluating agent skills using benchmarks, paired variants, and holdout splits. It supports repeated-run statistics, script assertions, Anthropic-compatible exports, and Jetty adapter integration, targeting AI researchers and developers.
Inferred · not functionally tested
Overview
6 featuresPurpose: Standardizing and automating the evaluation of agent skills across benchmarks and variants.
Inferred · not functionally tested
Audience: AI researchers and developers evaluating agent performance
Inferred · not functionally tested
Functions: analytics
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: cli
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: pypi.org · github.com. These links do not verify the individual claims.
In the Agent evaluation & testing space, skill-eval-harness takes a focused approach. Inferred · not functionally tested: It focuses on standardizing and automating the evaluation of agent skills across benchmarks and variants. Inferred · not functionally tested: It is built as an open-source project for AI researchers and developers evaluating agent performance. Basis unknown · not verified: skill-eval-harness is open source under the MIT license. Basis unknown · not verified: It runs on the command line.
Behind skill-eval-harness is adewale, and it first shipped in 2026. The project is developed in the open on GitHub with 54 stars and 76 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include skill evaluation, paired variants, and holdout splits. Inferred · not functionally tested: Catalogued interfaces include a public API.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Skill evaluation
- Paired variants
- Holdout splits
- Repeated-run stats
- Script assertions
- Anthropic export
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about skill-eval-harness
- What is skill-eval-harness?
- Inferred · not functionally tested: Skill-eval-harness focuses on standardizing and automating the evaluation of agent skills across benchmarks and variants. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is skill-eval-harness for?
- Inferred · not functionally tested: skill-eval-harness is an open-source project built for AI researchers and developers evaluating agent performance.
- Is skill-eval-harness free?
- Basis unknown · not verified: Yes — skill-eval-harness is open source under the MIT license and free to use.
- What platforms does skill-eval-harness run on?
- Basis unknown · not verified: skill-eval-harness runs on the command line.
- Is skill-eval-harness still maintained?
- PulseGate's liveness check found it on 14 Sep 2026. Its GitHub repository shows 76 commits in the last 90 days.
- What projects are similar to skill-eval-harness?
- Similar projects tracked by PulseGate include skill-harness, skill-inspect, and skillfed.skill-harnessskill-inspectskillfed
- Who develops skill-eval-harness?
- skill-eval-harness is developed by adewale.
- When did skill-eval-harness launch?
- skill-eval-harness first shipped in 2026.
Similar projects
Closest matches by what these projects do