LitigationBench is a benchmark developed by Litco that measures language models on litigation tasks. It runs each model on identical tasks twice, once without safeguards and once with Litco’s safeguards active inside its production agent. The benchmark publishes both scores, the performance gap, fabricated authorities, false premises, and costs to support evaluation of model behavior in legal contexts.
A composite quality score appears for each setting, shown alongside the score before penalties so the effect of penalties is visible. Columns for fabricated authorities and false premises use color coding: green for zero instances, amber for an adopted false premise, and red for a fabrication. Full flag details are available in expanded row panels and a candor matrix. Cost reflects the metered provider bill for the tasks, while the self-hosted row runs on Litco’s own hardware and lists price in its cost column.
The leaderboard allows sorting by column, with arrows indicating the preferred direction. Clicking a row displays both settings side by side along with serving pins. Models missing a setting appear below scored rows with an explanatory note. The benchmark focuses on frontier and self-hosted models and includes every failure observed, even those occurring inside the product.
LitigationBench is a LLM eval & observability project. It focuses on measuring and comparing how well AI models perform on complex litigation tasks while identifying safety and accuracy gaps. It is built as a B2B product for AI researchers and legal tech developers. It ships for the web and the command line.
Behind LitigationBench is Litco, and it first shipped in 2025. Among its 6 catalogued features are leaderboard, Model Comparison, and Safeguard Testing. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do