VeloxQuant-MLX is an MIT-licensed Python package for quantizing LLM key-value caches on Apple Silicon using MLX. It includes 43 research-adapted compression methods and Metal kernels for developers optimizing local inference workloads.
In the Inference & model serving space, VeloxQuant-MLX takes a focused approach. It focuses on reducing LLM KV-cache memory usage and improving local inference performance on Apple Silicon. VeloxQuant-MLX is an open-source project aimed at ML engineers and developers running LLM inference on Apple Silicon. VeloxQuant-MLX is open source under the MIT license. It runs on the command line, and it can be self-hosted.
It is developed by rajveer43, and it first shipped in 2026. The project is developed in the open on GitHub with 14 stars and 502 commits in the last 90 days. Key capabilities include KV-cache quantization, 43 compression methods, and metal kernels.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do