Skip to content
AI
LLM eval & observability
LLM eval & observability inside AI.
Niches in LLM eval & observability
Newest through the gate
- ami-surveypypi.orgami-survey measures cost, token usage, and duration for completed agent workflows from session logs.
- VSArenavercel.appVSArena is a browser-based evaluation arena for testing embodied robot policies on stacking tasks.
- chainbreakpypi.orgchainbreak is an empirical benchmark for authorization behavior in delegated and agentic cloud systems.
- Open SLM Leaderboardhuggingface.coOpen SLM Leaderboard lets users browse and compare open-source language models by size, architecture, datasets, and performance.
- AgentCrashpypi.orgAgentCrash is an open-source harness for security-testing AI agents in synthetic workplaces.
- rubricapypi.orgrubrica generates skill-based test suites for evaluating agentic systems.
- Container Host AIopsgithub.comContainer Host AIops provides governed diagnosis and guarded remediation for Docker and Portainer container hosts.
- Endpoint AIopsgithub.comEndpoint AIops provides governed AI operations for managed endpoint fleets, including login-storm and configuration-drift analysis.
- Postgres AIopsgithub.comPostgres AIops provides governed AI operations for diagnosing and managing PostgreSQL databases.
- Observability AIopsgithub.comObservability AIops provides governed AI operations for self-hosted Prometheus and Grafana monitoring environments.
- MySQL AIopsgithub.comMySQL AIops provides governed AI operations for diagnosing MySQL and MariaDB database issues.
- Inference AIopsgithub.comInference AIops provides governed GPU inference operations for diagnosing latency, scaling services, and draining workloads.
- Queue AIopsgithub.comQueue AIops provides governed root-cause analysis for Redis and RabbitMQ queue and connection issues.
- tracesweeppypi.orgtracesweep is a planned developer tool for diagnosing problems in agent traces.
- HungryGPUhungrygpu.comHungryGPU tracks AI models, patches, recipes, and papers by hardware compatibility.
- ContextBurngithub.comContextBurn measures how efficiently coding agents turn paid tokens into generated output.
- Harness Optimizationhuggingface.coHarness Optimization runs an automated loop that improves an agent harness around a fixed model and displays benchmark score changes.
- KnowMeNotknowmenot.comKnowMeNot compares how well your AI assistant knows you against a new model using a private quiz.
- Try ObserverBenchgithub.ioObserverBench benchmarks AI monitoring methods by testing the decisions and harms they miss.
- AI Tracevisualstudio.comA flight recorder for AI coding.
- MissionBellvisualstudio.comMission control + native notifications for Claude Code (Windows & macOS): see every session (working / done / needs answer) and get alerted from any window.
- Pacify-X Control Planevisualstudio.comGoverned operational dashboard and native VS Code control surface for Pacify-X.
- TurnStagevisualstudio.comTest LLM chat and agent APIs with streaming diagnostics, regression suites, red-team cases, and evidence
- callwitnesspypi.orgcallwitness records every tool call made by an AI agent for observability and security.
This is the newest 24 of 1,566. Open LLM eval & observability in the live index →
Subscribe to LLM eval & observability by RSS — new listings in this category, in your reader, no account.
Elsewhere in AI
Foundation models & chatCoding AI & assistantsImage generationVideo generationVoice, TTS & speechAutonomous agents & workflowsRAG, search & retrievalData science & ML workbenchFine-tuning & trainingInference & model servingOther AIComputer vision, OCR & document AIWriting & editingAI security & guardrails3D generation