LLM eval & observability Tools & Software
LLM eval & observability is part of AI on PulseGate. PulseGate tracks 1,118 LLM eval & observability products — 45 indexed in the past week, most recently sentinel-ai-auditor.
sentinel-ai-auditor is a deterministic, offline-first security auditor for AI agent skills and instruction bundles.
reprosieve reduces failed agent traces into redacted offline predicate reproductions.
CrewScore provides an offline structural production-readiness scorecard for AI agent system prompts.
critic-orchestrator is an MCP server that spawns parallel adversarial AI reviewers to verify claims and reduce confabulation in AI coding agents.
SlopCodeBench is a community-driven benchmark measuring code erosion in AI coding agents as they iteratively extend solutions across checkpoints.
flightdeck-verify is a standalone offline verifier for FlightDeck evidence packs generated by AI agent fleets.
Evarness is a Python package for proving AI agent harnesses with traces, invariants, and offline verification.
Stepproof is a Python package providing step-level verification and tamper-evident audit trails for AI agents.
agent-detective is a blame analysis tool for multi-agent runs that processes OTEL traces.
detective-ci is a deterministic blame-level golden replay and CI gate for Agent Detective.
blame-engine is a pure, I/O-free blame analysis engine for multi-agent execution graphs.
model-integrity-cli is a command-line tool that performs pre/post system state checksums to verify AI model runtime integrity.
ctxlens-cli is a context-window profiler for AI agents that identifies token waste and optimizes context usage.
answerproof provides verifiable, tamper-evident receipts for RAG and agent answers.
Pluto LLM Diff provides behavioral regression testing for language models.
iFixAi is an open-source CLI diagnostic tool that screens AI agents and deployments for misalignment risks.
kibsu reads coding-agent instructions and reports which of them can actually be verified.
Agentdiscover scans infrastructure to detect and inventory all AI agents using static analysis, runtime monitoring, and cloud audit logs.
Agentverity provides decision stability and coverage checks for AI agent tests.
semantic-entropy-gate is a Python package to detect LLM confabulation using semantic entropy and gate agent actions.
mcplint-cli lints and benchmarks MCP tool contracts to prevent ambiguous descriptions from breaking LLM tool selection.
steward-agent-governance provides citation-verified effective-access analysis for AI agent fleets.
quickstarted is a tool to test whether an AI agent can complete your quickstart by following your documentation.
agent-cost-attribution provides per-stage token and cost tracking for multi-agent AI workflows.
PandaProbe is an open-source platform for tracing, evaluating, and monitoring AI agents in production.
whatbroke-recorder is a Python package that records AI agent traces as whatbroke JSONL for later diffing and regression analysis.
hermes-token-consumption-tracker is a Python package that tracks per-request LLM token consumption for the Hermes Agent.
hermes-llm-api-call-logger is a Python package that logs every LLM API call with full request and response data for the Hermes Agent.
agentgauge-harness is a statistical regression harness for MCP tool descriptions that measures impact on agent task success, with a deterministic defect linter.
Budgie-firewall is a spend firewall for AI agents that prices cloud commands before execution and blocks over-budget ones.
Intentos is an open-source flight recorder for AI agents that tracks actions, failures, and costs.
ToolBench is a platform and CLI for building benchmarks for agentic tools and LLM harnesses.
attested is a Python package for independent verification of AI agent work and claims.
mcp-gauntlet is an agentic evaluation harness for MCP servers that tests whether AI agents can accomplish real tasks using the server's tools.
instar-harness is an open-source LLM measurement and evaluation harness.
webrtrace provides causality tracing for multi-agent AI systems to identify the node that broke the chain.
ActionRail is an open-source runtime enforcement layer that verifies AI agent actions against policies and live systems of record before execution.
ModelBias.ai is a platform that compares hidden defaults, associations, and biases across 100 AI models by running standardized prompts.
harness-bench-fast is a self-contained 298-task benchmark for evaluating AI agents on file operations, code edits, data pipelines, and memory discipline.
eval-integrity provides dependency-free statistical checks for AI evaluation claims including multiple-comparisons, judge-bias, resolution, and fragility as a CLI and MCP server.
ciagent is a Python package providing stability reports, flip attribution, LLM judge audits, and deterministic checks for AI agent evaluations.
Emisar is a zero-trust MCP server that safely connects AI agents like Claude and ChatGPT to production infrastructure with policy enforcement and audit trails.
journeygraph is a local-first graph analytics library for AI agent traces and event data.

ChirpPal is a monitoring tool that wraps AI API calls to detect model deprecations, output changes, and performance issues in real time.
trajectory-causal-attribution records AI agent trajectories to identify failure-causing steps via counterfactual replay.
context-tracker provides context window forensics, token tracking, and optimization for Claude Code with hooks, dashboard, and MCP server.
FreshBench provides local-first benchmarks for OpenAI-compatible language models.
agwer is a Python package providing agent-oriented word error rate, named entity F1, hallucination, and multi-speaker error metrics for evaluating generative AI and speech systems.
Path-index is a diagnostic framework for assessing the execution integrity and health of AI agent trajectories.
Agentlinter estimates costs and lints agent workflow YAML files.