This is an AWQ-quantized release of Alibaba's Qwen2-VL-72B-Instruct, a 72-billion-parameter vision-language model capable of understanding both images and video. The quantization enables more efficient local or self-hosted deployment while preserving strong performance on multimodal tasks. It is intended for developers building applications that require high-capability image and video reasoning using open-source models.
In the Foundation models & chat space, Qwen2 VL 72B Instruct takes a focused approach. It focuses on running a large-scale multimodal vision-language model locally with reduced VRAM usage. It is built as an open-source project for developers. The project is open source (Apache-2.0). It ships for the web and the command line, and it can be self-hosted.
It is developed by Qwen, and it first shipped in 2024. The project is developed in the open on GitHub with 19.7k stars. The category is crowded — PulseGate's index counts 25 comparable apps. Among its 5 catalogued features are vision-Language, AWQ Quantization, and Image Understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do