ai-eval-forge
PulseGate's liveness check found it on 9 Oct 2026; it is registered on GitHub and PyPI and has been in the index since 16 Jun 2026. How this is checked
Zero-dependency eval harness for LLM and agent regression testing. Scores outputs with exact, contains, regex, JSON, citation, and token-F1 checks. Compares two runs to flag regressions.
Inferred · not functionally tested
Overview
5 featuresPurpose: Automating regression testing and evaluation of LLM and agent outputs for developers and researchers.
Inferred · not functionally tested
Audience: AI developers and researchers
Inferred · not functionally tested
Functions: analytics, monitoring
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: unknown
Recorded constraints: pricing: open_source · license: MIT · platforms: CLI · deployment: browser, cli
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: github.com. These links do not verify the individual claims.
ai-eval-forge sits in PulseGate's LLM evaluation & benchmarks category. Inferred · not functionally tested: It focuses on automating regression testing and evaluation of LLM and agent outputs for developers and researchers. Inferred · not functionally tested: ai-eval-forge is an open-source project aimed at AI developers and researchers. Basis unknown · not verified: The project is open source (MIT). Basis unknown · not verified: It ships for the web and the command line.
It is developed by Mukunda Katta, and it first shipped in 2026. Development happens publicly on GitHub with 4 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include regression testing, output scoring, and multiple check types.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Regression testing
- Output scoring
- Multiple check types
- Run comparison
- Zero dependencies
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
What PulseGate has recorded for this listing
Frequently asked questions about ai-eval-forge
- What does ai-eval-forge do?
- Inferred · not functionally tested: Ai-eval-forge focuses on automating regression testing and evaluation of LLM and agent outputs for developers and researchers. It is catalogued under LLM evaluation & benchmarks on PulseGate.
- Who should use ai-eval-forge?
- Inferred · not functionally tested: ai-eval-forge is an open-source project built for AI developers and researchers.
- Is ai-eval-forge free?
- Basis unknown · not verified: Yes — ai-eval-forge is open source under the MIT license and free to use.
- What platforms does ai-eval-forge run on?
- Basis unknown · not verified: ai-eval-forge runs on the web and the command line.
- Is ai-eval-forge still maintained?
- PulseGate's liveness check found it on 9 Oct 2026. Its GitHub repository shows 4 commits in the last 90 days.
- What projects are similar to ai-eval-forge?
- Similar projects tracked by PulseGate include agent-eval, evalforge, and tool-eval.agent-evalevalforgetool-eval
- Who makes ai-eval-forge?
- ai-eval-forge is developed by Mukunda Katta.
- How long has ai-eval-forge been around?
- ai-eval-forge first shipped in 2026.
Similar projects
Closest matches by what these projects do