Qwen2.5-VL-72B-Instruct is a 72 billion parameter vision-language model developed by the Qwen team at Alibaba. It processes images, videos, and text together, supporting advanced multimodal reasoning and instruction following. The model is openly available on Hugging Face for local or hosted inference.
Qwen2.5 VL 72B Instruct sits in PulseGate's Foundation models & chat category. It focuses on understanding and reasoning over images, videos, and text using a single large multimodal model. It is built as an open-source project for developers. Qwen2.5 VL 72B Instruct is open source under the Apache-2.0 license. It runs on the web, the command line, and API.
It is developed by Qwen, and the product first shipped in 2024. Development happens publicly on GitHub with 19.6k stars. The category is crowded — PulseGate's index counts 25 comparable apps. Key capabilities include vision-language understanding, image and video processing, and instruction following.
Latest indexed changes and source events
Qwen/Qwen2.5-VL-72B-Instruct verified by the PulseGate indexer
Other apps tracked under the same category.