mcp-llm-eval is an open-source CLI tool and MCP server that packages LLM evaluation gates as reusable primitives for CI/CD workflows. It enables developers and ML engineers to automate and standardize the evaluation of large language models within their development pipelines. The tool supports benchmarking and integration with MCP for scalable model assessment.
mcp-llm-eval sits in PulseGate's LLM eval & observability category. It focuses on automating and standardizing LLM evaluation in CI/CD pipelines for developers and ML engineers. It is built as an open-source project for machine learning engineers. The project is open source (MIT). mcp-llm-eval is available on the command line and API, and it can be self-hosted.
It is developed by berkayildi, and it first shipped in 2026. The project is developed in the open on GitHub with 57 commits in the last 90 days. Key capabilities include LLM evaluation, CI/CD integration, and reusable gates. It exposes integrations via an MCP server and a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do