skilldiff is an open-source CLI for A/B testing AI agent skills, prompts, and tool definitions. It uses LLM-as-a-Judge evaluation to compare agent configurations, with support for Gemini and Ollama-based workflows.
skilldiff is a LLM eval & observability project. It focuses on comparing AI agent skills, prompts, and tool definitions consistently without manually judging outputs. It is built as an open-source project for AI developers building and evaluating agent workflows. The project is open source (MIT). It ships for the command line, and it can be self-hosted.
Behind skilldiff is shunvel, and it first shipped in 2026. The project is developed in the open on GitHub with 11 commits in the last 90 days. Among its 6 catalogued features are A/B testing, prompt evaluation, and tool definition testing.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do