Qwen2.5-VL-7B-Instruct-GGUF is a repository on Hugging Face that supplies GGUF quantized files of the Qwen2.5-VL 7B vision-language model. It is intended for local inference of a multimodal model capable of processing both images and video alongside text. The repository is provided by unsloth.
The files include variants such as Qwen2.5-VL-7B-Instruct-BF16.gguf and Qwen2.5-VL-7B-Instruct-IQ4_NL.gguf. A chat template is defined that handles messages containing text, image, or video content by inserting specific vision and image or video pad tokens. The template supports counting multiple images or videos in a conversation and adds optional identifiers such as "Picture 1:" or "Video 1:" before the visual tokens. It uses special tokens including vision_start, vision_end, image_pad, video_pad, im_start, and im_end.
The total file size across the GGUF files is 15237851776 bytes. These quantized models are distributed for use with compatible inference engines that accept the GGUF format. The repository forms part of the broader collection of foundation models available on the platform.
No pricing, licensing terms, or additional deployment details beyond the GGUF files and chat template are stated.
Qwen2.5 VL 7B Instruct is a Foundation models & chat product. It focuses on running a capable 7B vision-language model efficiently on consumer hardware using quantized GGUF format. Qwen2.5 VL 7B Instruct is an open-source project aimed at Local AI users and developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
It is developed by Unsloth, and the product first shipped in 2024. The project is developed in the open on GitHub with 19.6k stars. It competes in a saturated segment with 25 similar apps in PulseGate's index. Among its 4 catalogued features are Vision Language, GGUF Quantization, and Local Inference.
Latest indexed changes and source events
unsloth/Qwen2.5-VL-7B-Instruct-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.