This is an Unsloth-optimized 4-bit quantized version of Alibaba's Qwen3-VL-2B-Instruct vision-language model. It supports understanding both images and text, making it suitable for multimodal tasks. The bnb-4bit format allows efficient local inference on modest hardware while preserving vision-language capabilities.
In the Multimodal & vision space, Qwen3 VL 2B Instruct Unsloth Bnb takes a focused approach. It focuses on running a small, efficient multimodal vision-language model locally with minimal memory usage. It is built as an open-source project for developers. Qwen3 VL 2B Instruct Unsloth Bnb is open source under the Apache-2.0 license. It ships for the web and the command line, and it can be self-hosted.
Behind Qwen3 VL 2B Instruct Unsloth Bnb is Alibaba Cloud, based in China, and it first shipped in 2023. Development happens publicly on GitHub with 68.7k stars and 1.2k commits in the last 90 days. The category is crowded — PulseGate's index counts 25 comparable projects. Key capabilities include vision-language, 4-bit quantized, and multimodal.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do