Munibench provides a flexible, callback-driven framework for running benchmarks against memory-augmented LLM systems. It is designed specifically for long-context and memory evaluation tasks, allowing researchers to easily instrument and measure how well different memory architectures perform on complex retrieval and reasoning benchmarks. The package is distributed via PyPI and is intended for use in research and development workflows.
Munibench is a Data science & ML workbench product. It focuses on evaluating and benchmarking memory systems for large language models. It is built as an open-source project for AI researchers and developers. Munibench is open source under the Open Source license. It runs on the command line.
Munibench first shipped in 2026. Key capabilities include callback-based runner, memory system evaluation, and LLM benchmarking.
Latest indexed changes and source events
Other apps tracked under the same category.