InternVL2_5-4B is a multimodal large language model hosted on Hugging Face under the OpenGVLab organization. The model accepts both image and video inputs together with text and follows a chat-style conversation format that includes special tokens for different content types.
Its tokenizer configuration defines specific system prompts that identify it as 书生·万象 with the English name InternVL. The prompt template supports roles for system, user, and assistant messages and inserts placeholder tokens such as <image or <video when the content type requires them. The end-of-sentence token is defined as <|im_end|, the beginning-of-sentence marker as <|im_start|, and padding uses <|endoftext|.
The model page records over 80000 recent downloads and more than 723000 downloads in total. It was created on 2024-11-20 and remains open for community discussions. The repository provides the model weights, configuration files, and a chat template that enables direct inference through the Hugging Face ecosystem.
As a foundation model, InternVL2_5-4B is intended for developers and researchers who integrate multimodal understanding capabilities into applications. The page does not specify licensing details, training data, or parameter count.
In the Foundation models & chat space, InternVL2 5 4B takes a focused approach. It focuses on understanding and reasoning over both images and video content using a compact multimodal model. InternVL2 5 4B is an open-source project aimed at developers. The project is open source (MIT). It runs on the web, the command line, and API.
Behind InternVL2 5 4B is OpenGVLab, based in China, and the product first shipped in 2023. The project is developed in the open on GitHub with 10.1k stars. Among its 3 catalogued features are image understanding, video understanding, and multilingual support.
Latest indexed changes and source events
OpenGVLab/InternVL2_5-4B verified by the PulseGate indexer
Other apps tracked under the same category.