EvalView
PulseGate's liveness check found it on 3 Oct 2026; it is registered on GitHub and has been in the index since 26 Aug 2026. How this is checked
EvalView is an open-source pytest-style testing framework for AI agents. Its CLI records tool calls, parameters, sequences, outputs, costs, and latency, then compares runs with golden baselines using rule-based, semantic, and LLM-based scoring for CI/CD regression checks.
Inferred · not functionally tested
Overview
6 featuresPurpose: Detecting silent behavior, tool-calling, and output regressions in AI agents before production.
Inferred · not functionally tested
Audience: AI agent developers and engineering teams
Inferred · not functionally tested
Functions: analytics, monitoring
Inferred · not functionally tested
Interfaces: API: unknown · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI · deployment: cli, self_hosted
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: evalview.com · github.com. These links do not verify the individual claims.
EvalView sits in PulseGate's Agent evaluation & testing category. Inferred · not functionally tested: It focuses on detecting silent behavior, tool-calling, and output regressions in AI agents before production. Inferred · not functionally tested: EvalView is an open-source project aimed at AI agent developers and engineering teams. Basis unknown · not verified: EvalView is open source under the Apache-2.0 license. Basis unknown · not verified: It ships for the command line, and it can be self-hosted.
hidai25 builds and maintains EvalView, and it first shipped in 2025. Development happens publicly on GitHub with 130 stars and 15 commits in the last 90 days. Inferred · not functionally tested: Key capabilities include agent test generation, behavior snapshots, and golden baselines.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Agent test generation
- Behavior snapshots
- Golden baselines
- Tool-call diffing
- Output regression checks
- JSON schema checks
Topics: Inferred · not functionally tested
Built with & integrations
- local_oss
- bollama in the HTML
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed26 Aug · 17:57 UTChidai25/eval-view seen via GitHub Search (Smart)Source: GitHub Search (Smart) · Open
Frequently asked questions about EvalView
- What is EvalView?
- Inferred · not functionally tested: EvalView focuses on detecting silent behavior, tool-calling, and output regressions in AI agents before production. It is catalogued under Agent evaluation & testing on PulseGate.
- Who is EvalView for?
- Inferred · not functionally tested: EvalView is an open-source project built for AI agent developers and engineering teams.
- Is EvalView free?
- Basis unknown · not verified: Yes — EvalView is open source under the Apache-2.0 license and free to use.
- What platforms does EvalView run on?
- Basis unknown · not verified: EvalView runs on the command line. It can also be self-hosted.
- Is EvalView still active?
- PulseGate's liveness check found it on 3 Oct 2026. Its GitHub repository shows 15 commits in the last 90 days.
- What are alternatives to EvalView?
- Similar projects tracked by PulseGate include evalite, Aeval, and evalgrid.evaliteAevalevalgrid
- Who makes EvalView?
- EvalView is developed by hidai25.
- When did EvalView launch?
- EvalView first shipped in 2025.
Similar projects
Closest matches by what these projects do