An LLM benchmark that measures ability to do long calculations without chain-of-thought
In the LLM evaluation & benchmarks space, Latentmathbench takes a focused approach. It focuses on measuring latent mathematical reasoning and long calculations without relying on visible chain-of-thought. It is built as an open-source project for AI researchers and machine learning engineers. Latentmathbench is open source under the MIT license. It ships for the web and the command line, and it can be self-hosted.
It is developed by Maarten Baert, and it first shipped in 2026. Development happens publicly on GitHub with 3 commits in the last 90 days. Among its 5 catalogued features are math benchmarks, latent reasoning tests, and long calculation evaluation.
What PulseGate has recorded for this listing
Same category — not a similarity match