This is a GGUF-quantized version of Alibaba's Qwen2.5-VL-7B-Instruct model, optimized for local inference. It supports understanding both images and video content alongside text, making it suitable for multimodal applications. The model is popular in the LM Studio community for local AI use.
Qwen2.5 VL 7B Instruct is a Multimodal & vision project. It focuses on running multimodal vision-language models locally with support for images and video. It is built as an open-source project for developers. Qwen2.5 VL 7B Instruct is open source under the MIT license. It ships for the web, the command line, and API.
Behind Qwen2.5 VL 7B Instruct is lmstudio-community, and it first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. The category is crowded — PulseGate's index counts 23 comparable projects. Among its 4 catalogued features are Vision Language Model, Image Understanding, and Video Understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do