This is a 2B parameter multimodal model from Alibaba NLP based on Qwen2-VL. It is instruction-tuned to handle both image and text inputs for various vision-language tasks. The model supports chat-style interactions with images and video and is available for download and local inference through the Hugging Face ecosystem.
Gme Qwen2 VL 2B Instruct is a Foundation models & chat project. It focuses on performing vision-language understanding and generation tasks with a compact multimodal model. It is built as an open-source project for machine learning developers. Gme Qwen2 VL 2B Instruct is open source under the Open Source license. It ships for the web and API.
It is developed by Alibaba-NLP (China), and it first shipped in 2024. It operates in a well-populated space: PulseGate tracks 12 similar projects. Among its 3 catalogued features are Vision Language, Multimodal Input, and Instruction Tuning.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do