Munibench provides a flexible, callback-driven framework for running benchmarks against memory-augmented LLM systems. It is designed specifically for long-context and memory evaluation tasks, allowing researchers to easily instrument and measure how well different memory architectures perform on complex retrieval and reasoning benchmarks. The package is distributed via PyPI and is intended for use in research and development workflows.
Munibench is a Data science & ML workbench project. It focuses on evaluating and benchmarking memory systems for large language models. Munibench is an open-source project aimed at AI researchers and developers. The project is open source (Open Source). Munibench is available on the command line.
Munibench first shipped in 2026. Among its 3 catalogued features are callback-based runner, memory system evaluation, and LLM benchmarking.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do