Mendmark-evals is a Python package for performing mutation testing on evaluation suites used for LLM-based agents. It helps developers assess how well their test frameworks detect failures and measure the quality of autonomous agent implementations.
mendmark-evals sits in PulseGate's LLM eval & observability category. It focuses on evaluating the robustness and reliability of AI agent testing frameworks. It is built as an open-source project for AI developers. The project is open source (MIT). It ships for the command line.
It is developed by Daniel Gaskins, and it first shipped in 2026. Development happens publicly on GitHub with 14 commits in the last 90 days. Key capabilities include mutation testing, agent evaluation, and test suite analysis.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match