Open Source AI Inference Benchmark, also referred to as InferenceX and InferenceMAX, is an open-source benchmarking suite focused on evaluating AI inference performance. The platform provides transparent and reproducible benchmarks that measure how AI models perform during inference across a range of hardware, including NVIDIA and AMD GPUs, and software stacks. It is designed to address the need for vendor-neutral, continuously updated benchmarks as models and inference systems evolve, offering the AI community real-world performance data that supports informed decision-making.
The tool emphasizes transparency and reproducibility, delivering data on metrics such as token throughput, performance per dollar, and tokens per megawatt. It benchmarks the latest software optimizations and inference engines, including those using features like FP4, MTP, speculative decode, and wide-EP, to reflect actual deployment scenarios. The platform runs benchmarks nightly, ensuring that the data remains current as new models, hardware, and optimizations are introduced. It supports a variety of inference stacks and frameworks, such as vLLM, SGLang, TensorRT-LLM, and PyTorch, and tracks the performance of models like MiniMax M3, Qwen, and Kimi K2.
InferenceX is intended for a broad audience, including researchers, engineers, practitioners, platform teams, and operators of large-scale datacenters. Its openly available results help these users compare inference efficiency, throughput, and cost across different hardware and software configurations. The platform is community-driven, with contributions and collaboration from various organizations in the AI ecosystem, and is trusted by stakeholders from academia, industry, and cloud infrastructure providers.
The benchmarking suite is delivered as an open-source project, allowing the community to access, review, and build upon its methodology and results. By providing public, reproducible benchmarks under realistic workloads, it aims to strengthen the AI inference ecosystem and support ongoing innovation in both hardware and software. The platform does not mention specific pricing or licensing terms beyond its open-source nature.
Open Source AI Inference Benchmark is a LLM eval & observability project. It focuses on providing transparent, reproducible benchmarks for AI inference performance across hardware and frameworks. It is built as an open-source project for AI researchers and ML engineers. The project is open source (Apache-2.0). It runs on the web, and it can be self-hosted.
SemiAnalysis builds and maintains Open Source AI Inference Benchmark, and it first shipped in 2025. The project is developed in the open on GitHub with 1.2k stars and 594 commits in the last 90 days. Key capabilities include AI benchmarking, GPU comparison, and framework analysis.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do