vMLX is a free, open-source macOS application and inference engine for running LLMs locally on Apple Silicon using MLX. It provides prefix and paged KV caching, continuous batching, MCP tool support, and an OpenAI-compatible API for developers and local AI users.
vMLX sits in PulseGate's Inference & model serving category. It focuses on running and serving language models locally on Apple Silicon without sending data to the cloud. vMLX is an open-source project aimed at mac developers and users running local LLMs. vMLX is open source under the Apache-2.0 license. It ships for the web, the command line, macOS, and API, and it can be self-hosted.
It is developed by jjang-ai, and it first shipped in 2026. Development happens publicly on GitHub with 839 stars and 2.7k commits in the last 90 days. Among its 10 catalogued features are prefix caching, Paged KV cache, and continuous batching. It exposes integrations via an MCP server and a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do