This is a 6-bit quantized MLX version of the Qwen3-VL-4B vision-language model, optimized for Apple Silicon devices. It allows local multimodal inference combining vision and language understanding. The model is targeted at developers who want to run capable vision-language models on Macs or other local hardware without cloud services.
Qwen3 VL 4B Instruct is a Multimodal & vision project. It focuses on running vision-language models efficiently on local Apple hardware with limited memory. It is built as an open-source project for developers. Qwen3 VL 4B Instruct is open source under the MIT license. Qwen3 VL 4B Instruct is available on the web and the command line, and it can be self-hosted.
lmstudio-community builds and maintains Qwen3 VL 4B Instruct, and it first shipped in 2024. The project is developed in the open on GitHub with 5.2k stars and 419 commits in the last 90 days. It competes in a saturated segment with 25 similar projects in PulseGate's index. Among its 3 catalogued features are vision-language, Quantized MLX, and multimodal understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do