Llamathon is a Python benchmarking tool for LLM inference servers. It orchestrates tests using llama-benchy with automatic hardware and model detection, realistic workload generation, and produces charts and detailed reports. It supports popular backends including Ollama, vLLM, and llama.cpp.
In the AI & ML space, llamathon takes a focused approach. Accurately measuring and reporting the real-world LLM inference performance and capacity of a server. It is built as an open-source project for ML engineers and infrastructure teams. llamathon is open source under the MIT license. It runs on the command line, and it can be self-hosted.
llamathon first shipped in 2026. The project is developed in the open on GitHub with 15 commits in the last 90 days. Among its 5 catalogued features are LLM Benchmarking, Inference Testing, and auto-detection.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do