DeepResearch Bench is an interactive leaderboard for evaluating deep research and LLM-with-search models. Users can filter models by name or category and compare rankings, scores, and related benchmark details.
DeepResearch Bench sits in PulseGate's LLM eval & observability category. It focuses on comparing the performance of deep research and search-enabled language models using standardized leaderboard scores. DeepResearch Bench is a B2B product aimed at AI researchers and developers evaluating research-capable language models. DeepResearch Bench costs nothing to use. It ships for the web, and it can be self-hosted.
muset-ai builds and maintains DeepResearch Bench. Key capabilities include model search, category filters, and rankings table.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do