Qwen2.5-VL-7B-Instruct-GGUF is a repository on Hugging Face that supplies GGUF quantized files of the Qwen2.5-VL 7B vision-language model. It is intended for local inference of a multimodal model capable of processing both images and video alongside text. The repository is provided by unsloth.
The files include variants such as Qwen2.5-VL-7B-Instruct-BF16.gguf and Qwen2.5-VL-7B-Instruct-IQ4_NL.gguf. A chat template is defined that handles messages containing text, image, or video content by inserting specific vision and image or video pad tokens. The template supports counting multiple images or videos in a conversation and adds optional identifiers such as "Picture 1:" or "Video 1:" before the visual tokens. It uses special tokens including vision_start, vision_end, image_pad, video_pad, im_start, and im_end.
The total file size across the GGUF files is 15237851776 bytes. These quantized models are distributed for use with compatible inference engines that accept the GGUF format. The repository forms part of the broader collection of foundation models available on the platform.
No pricing, licensing terms, or additional deployment details beyond the GGUF files and chat template are stated.
Qwen2.5 VL 7B Instruct sits in PulseGate's Multimodal & vision category. It focuses on running a capable 7B vision-language model efficiently on consumer hardware using quantized GGUF format. Qwen2.5 VL 7B Instruct is an open-source project aimed at Local AI users and developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
It is developed by Unsloth, and it first shipped in 2024. Development happens publicly on GitHub with 19.6k stars. The category is crowded — PulseGate's index counts 25 comparable projects. Key capabilities include Vision Language, GGUF Quantization, and Local Inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do