Qwen2.5-VL-72B-Instruct is a 72 billion parameter vision-language model developed by the Qwen team at Alibaba. It processes images, videos, and text together, supporting advanced multimodal reasoning and instruction following. The model is openly available on Hugging Face for local or hosted inference.
Qwen2.5 VL 72B Instruct is a Multimodal & vision project. It focuses on understanding and reasoning over images, videos, and text using a single large multimodal model. Qwen2.5 VL 72B Instruct is an open-source project aimed at developers. The project is open source (Apache-2.0). Qwen2.5 VL 72B Instruct is available on the web, the command line, and API.
It is developed by Qwen, and it first shipped in 2024. Development happens publicly on GitHub with 19.6k stars. It competes in a saturated segment with 25 similar projects in PulseGate's index. Key capabilities include vision-language understanding, image and video processing, and instruction following.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do