mlx-serve is a free and open source native application for macOS that enables users to run large language models (LLMs) and generative AI models entirely on Apple Silicon Macs. Designed for privacy and performance, the tool allows users to interact with AI for tasks such as chat, coding agents, image editing, video creation, music composition, 3D model generation, and voice cloning, all processed locally without any data leaving the device.
The platform supports a wide range of AI capabilities, including editing photos with text prompts, animating images into videos with synchronized sound, generating original music, converting photos into textured 3D models, and cloning voices from short audio samples. Coding features such as running Claude Code offline without an API key are also available. Users can summon AI through a launcher, voice commands, or scheduled actions, and sandbox AI agents in a Linux virtual machine to ensure system security. The tool is compatible with OpenAI and Anthropic APIs, allowing existing apps built for these APIs to connect seamlessly to mlx-serve. It also natively supports Ollama's protocol, enabling integration with tools designed for Ollama without modification.
Performance is a key focus, with benchmarks indicating that mlx-serve operates faster than comparable tools like LM Studio on identical hardware and models. The application leverages speculative decoding and prompt lookup decoding to increase token generation speed, with adaptive mechanisms to maintain output quality. It supports running advanced models such as DeepSeek V4 Flash with up to 284 billion parameters on Macs equipped with sufficient unified memory. The software is distributed as a single self-contained binary, requiring no external dependencies or setup scripts, and includes features such as eager warmup for quick initial responses and fast conversational turnarounds.
mlx-serve is available for Apple Silicon Macs running macOS 26 or later (M1 through M4). It is open source, downloadable at no cost, and does not require accounts, subscriptions, or cloud calls, making it suitable for users seeking private, high-performance local AI capabilities.
In the AI space, mlx-serve takes a focused approach. It focuses on running large language models and generative AI locally on Apple Silicon Macs without Python dependencies. mlx-serve is an open-source project aimed at developers and AI researchers using Apple Silicon Macs. mlx-serve is open source under the MIT license. mlx-serve is available on the web, the command line, macOS, and API.
It is developed by ddalcu, and it first shipped in 2026. The project is developed in the open on GitHub with 194 stars and 147 commits in the last 90 days. Among its 11 catalogued features are menu bar app, Local LLM inference, and OpenAI API compatible. It exposes integrations via a public API and an MCP server.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do