LLM eval & observability Tools & Software

Live LLM eval & observability listings on PulseGate.

LLM eval & observability is part of AI on PulseGate. PulseGate tracks 1,118 LLM eval & observability products — 45 indexed in the past week, most recently sentinel-ai-auditor.

sentinel-ai-auditor
pypi.org
LLM eval & observability · 8h ago

sentinel-ai-auditor is a deterministic, offline-first security auditor for AI agent skills and instruction bundles.

LLM eval & observability
8h ago
reprosieve
pypi.org
LLM eval & observability · 8h ago

reprosieve reduces failed agent traces into redacted offline predicate reproductions.

LLM eval & observability
8h ago
crewscore
crewscore.ai
LLM eval & observability · 12h ago

CrewScore provides an offline structural production-readiness scorecard for AI agent system prompts.

LLM eval & observability
12h ago
critic-orchestrator
pypi.org
LLM eval & observability · 16h ago

critic-orchestrator is an MCP server that spawns parallel adversarial AI reviewers to verify claims and reduce confabulation in AI coding agents.

LLM eval & observability
16h ago
SlopCodeBench logo
SlopCodeBench
scbench.ai
LLM eval & observability · 23h ago

SlopCodeBench is a community-driven benchmark measuring code erosion in AI coding agents as they iteratively extend solutions across checkpoints.

LLM eval & observability
23h ago
flightdeck-verify
pypi.org
LLM eval & observability · 1d ago

flightdeck-verify is a standalone offline verifier for FlightDeck evidence packs generated by AI agent fleets.

LLM eval & observability
1d ago
evarness
pypi.org
LLM eval & observability · 1d ago

Evarness is a Python package for proving AI agent harnesses with traces, invariants, and offline verification.

LLM eval & observability
1d ago
stepproof
pypi.org
LLM eval & observability · 1d ago

Stepproof is a Python package providing step-level verification and tamper-evident audit trails for AI agents.

LLM eval & observability
1d ago
agent-detective
pypi.org
LLM eval & observability · 1d ago

agent-detective is a blame analysis tool for multi-agent runs that processes OTEL traces.

LLM eval & observability
1d ago
detective-ci
pypi.org
LLM eval & observability · 1d ago

detective-ci is a deterministic blame-level golden replay and CI gate for Agent Detective.

LLM eval & observability
1d ago
blame-engine
pypi.org
LLM eval & observability · 1d ago

blame-engine is a pure, I/O-free blame analysis engine for multi-agent execution graphs.

LLM eval & observability
1d ago
model-integrity-cli
pypi.org
LLM eval & observability · 1d ago

model-integrity-cli is a command-line tool that performs pre/post system state checksums to verify AI model runtime integrity.

LLM eval & observability
1d ago
ctxlens-cli
pypi.org
LLM eval & observability · 1d ago

ctxlens-cli is a context-window profiler for AI agents that identifies token waste and optimizes context usage.

LLM eval & observability
1d ago
answerproof
pypi.org
LLM eval & observability · 1d ago

answerproof provides verifiable, tamper-evident receipts for RAG and agent answers.

LLM eval & observability
1d ago
pluto-llm-diff
pypi.org
LLM eval & observability · 1d ago

Pluto LLM Diff provides behavioral regression testing for language models.

LLM eval & observability
1d ago
iFixAi logo
iFixAi
ifixai.ai
LLM eval & observability · 1d ago

iFixAi is an open-source CLI diagnostic tool that screens AI agents and deployments for misalignment risks.

LLM eval & observability
1d ago
kibsu
pypi.org
LLM eval & observability · 2d ago

kibsu reads coding-agent instructions and reports which of them can actually be verified.

LLM eval & observability
2d ago
agentdiscover
pypi.org
LLM eval & observability · 2d ago

Agentdiscover scans infrastructure to detect and inventory all AI agents using static analysis, runtime monitoring, and cloud audit logs.

LLM eval & observability
2d ago
agentverity
pypi.org
LLM eval & observability · 2d ago

Agentverity provides decision stability and coverage checks for AI agent tests.

LLM eval & observability
2d ago
semantic-entropy-gate
pypi.org
LLM eval & observability · 2d ago

semantic-entropy-gate is a Python package to detect LLM confabulation using semantic entropy and gate agent actions.

LLM eval & observability
2d ago
mcplint-cli
pypi.org
LLM eval & observability · 2d ago

mcplint-cli lints and benchmarks MCP tool contracts to prevent ambiguous descriptions from breaking LLM tool selection.

LLM eval & observability
2d ago
steward-agent-governance
pypi.org
LLM eval & observability · 2d ago

steward-agent-governance provides citation-verified effective-access analysis for AI agent fleets.

LLM eval & observability
2d ago
quickstarted
pypi.org
LLM eval & observability · 2d ago

quickstarted is a tool to test whether an AI agent can complete your quickstart by following your documentation.

LLM eval & observability
2d ago
agent-cost-attribution
pypi.org
LLM eval & observability · 3d ago

agent-cost-attribution provides per-stage token and cost tracking for multi-agent AI workflows.

LLM eval & observability
3d ago
PandaProbe logo
PandaProbe
chirpz.ai
LLM eval & observability · 3d ago

PandaProbe is an open-source platform for tracing, evaluating, and monitoring AI agents in production.

LLM eval & observability
3d ago
whatbroke-recorder
pypi.org
LLM eval & observability · 3d ago

whatbroke-recorder is a Python package that records AI agent traces as whatbroke JSONL for later diffing and regression analysis.

LLM eval & observability
3d ago
hermes-token-consumption-tracker
pypi.org
LLM eval & observability · 3d ago

hermes-token-consumption-tracker is a Python package that tracks per-request LLM token consumption for the Hermes Agent.

LLM eval & observability
3d ago
hermes-llm-api-call-logger
pypi.org
LLM eval & observability · 3d ago

hermes-llm-api-call-logger is a Python package that logs every LLM API call with full request and response data for the Hermes Agent.

LLM eval & observability
3d ago
agentgauge-harness
pypi.org
LLM eval & observability · 3d ago

agentgauge-harness is a statistical regression harness for MCP tool descriptions that measures impact on agent task success, with a deterministic defect linter.

LLM eval & observability
3d ago
budgie-firewall
pypi.org
LLM eval & observability · 3d ago

Budgie-firewall is a spend firewall for AI agents that prices cloud commands before execution and blocks over-budget ones.

LLM eval & observability
3d ago
intentos
pypi.org
LLM eval & observability · 4d ago

Intentos is an open-source flight recorder for AI agents that tracks actions, failures, and costs.

LLM eval & observability
4d ago
ToolBench
pypi.org
LLM eval & observability · 4d ago

ToolBench is a platform and CLI for building benchmarks for agentic tools and LLM harnesses.

LLM eval & observability
4d ago
attested
pypi.org
LLM eval & observability · 4d ago

attested is a Python package for independent verification of AI agent work and claims.

LLM eval & observability
4d ago
mcp-gauntlet
pypi.org
LLM eval & observability · 5d ago

mcp-gauntlet is an agentic evaluation harness for MCP servers that tests whether AI agents can accomplish real tasks using the server's tools.

LLM eval & observability
5d ago
instar-harness
pypi.org
LLM eval & observability · 5d ago

instar-harness is an open-source LLM measurement and evaluation harness.

LLM eval & observability
5d ago
webrtrace
pypi.org
LLM eval & observability · 5d ago

webrtrace provides causality tracing for multi-agent AI systems to identify the node that broke the chain.

LLM eval & observability
5d ago
ActionRail logo
ActionRail
actionrail.ai
LLM eval & observability · 5d ago

ActionRail is an open-source runtime enforcement layer that verifies AI agent actions against policies and live systems of record before execution.

LLM eval & observability
5d ago
ModelBias.ai logo
ModelBias.ai
modelbias.ai
LLM eval & observability · 5d ago

ModelBias.ai is a platform that compares hidden defaults, associations, and biases across 100 AI models by running standardized prompts.

LLM eval & observability
5d ago
harness-bench-fast
pypi.org
LLM eval & observability · 5d ago

harness-bench-fast is a self-contained 298-task benchmark for evaluating AI agents on file operations, code edits, data pipelines, and memory discipline.

LLM eval & observability
5d ago
eval-integrity
pypi.org
LLM eval & observability · 5d ago

eval-integrity provides dependency-free statistical checks for AI evaluation claims including multiple-comparisons, judge-bias, resolution, and fragility as a CLI and MCP server.

LLM eval & observability
5d ago
ciagent
pypi.org
LLM eval & observability · 5d ago

ciagent is a Python package providing stability reports, flip attribution, LLM judge audits, and deterministic checks for AI agent evaluations.

LLM eval & observability
5d ago
emisar logo
emisar
emisar.dev
LLM eval & observability · 5d ago

Emisar is a zero-trust MCP server that safely connects AI agents like Claude and ChatGPT to production infrastructure with policy enforcement and audit trails.

LLM eval & observability
5d ago
journeygraph
pypi.org
LLM eval & observability · 6d ago

journeygraph is a local-first graph analytics library for AI agent traces and event data.

LLM eval & observability
6d ago
ChirpPal logo
ChirpPal
chirppal.com
LLM eval & observability · 6d ago

ChirpPal is a monitoring tool that wraps AI API calls to detect model deprecations, output changes, and performance issues in real time.

LLM eval & observability
6d ago
trajectory-causal-attribution
pypi.org
LLM eval & observability · 6d ago

trajectory-causal-attribution records AI agent trajectories to identify failure-causing steps via counterfactual replay.

LLM eval & observability
6d ago
context-tracker
pypi.org
LLM eval & observability · 7d ago

context-tracker provides context window forensics, token tracking, and optimization for Claude Code with hooks, dashboard, and MCP server.

LLM eval & observability
7d ago
FreshBench
pypi.org
LLM eval & observability · 7d ago

FreshBench provides local-first benchmarks for OpenAI-compatible language models.

LLM eval & observability
7d ago
agwer
pypi.org
LLM eval & observability · 7d ago

agwer is a Python package providing agent-oriented word error rate, named entity F1, hallucination, and multi-speaker error metrics for evaluating generative AI and speech systems.

LLM eval & observability
7d ago
path-index
pypi.org
LLM eval & observability · 8d ago

Path-index is a diagnostic framework for assessing the execution integrity and health of AI agent trajectories.

LLM eval & observability
8d ago
agentlinter
pypi.org
LLM eval & observability · 8d ago

Agentlinter estimates costs and lints agent workflow YAML files.

LLM eval & observability
8d ago
Showing 150 of 1,118
Explore results. Clear filters to return to Top 500.