A quantized version of the Qwen2.5-VL-3B-Instruct model optimized with AWQ. It accepts both image and video inputs along with text and follows natural language instructions for vision-language tasks. The model is hosted on Hugging Face and can be loaded via the transformers library or run with inference engines supporting GGUF/AWQ formats.
Qwen2.5 VL 3B Instruct sits in PulseGate's Multimodal & vision category. It focuses on running efficient multimodal (vision + language) models locally with reduced memory requirements. Qwen2.5 VL 3B Instruct is an open-source project aimed at developers and researchers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
It is developed by Qwen, and it first shipped in 2024. The project is developed in the open on GitHub with 19.6k stars. It competes in a saturated segment with 25 similar projects in PulseGate's index. Among its 4 catalogued features are vision-Language, Instruction Following, and AWQ Quantized.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do