Qwen3-VL-4B-Instruct-NVFP4 is a 4-billion parameter vision-language model based on the Qwen3 series, optimized with NVFP4 quantization for efficient local inference. It supports multimodal inputs including text, images, and video, along with advanced features such as tool calling and custom chat templates. Hosted on Hugging Face, it is designed for developers building local multimodal AI applications that require vision understanding and language generation without cloud dependencies.
In the Foundation models & chat space, Qwen3 VL 4B Instruct takes a focused approach. It focuses on running high-performance multimodal vision-language inference locally without relying on proprietary cloud APIs. Qwen3 VL 4B Instruct is an open-source project aimed at developers. Qwen3 VL 4B Instruct is open source under the Open Source license. It runs on the web and API.
nm-testing builds and maintains Qwen3 VL 4B Instruct, and it first shipped in 2025. It competes in a saturated segment with 25 similar projects in PulseGate's index. Among its 6 catalogued features are Vision Language Model, Multimodal Understanding, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do