InternVL3-8B-hf is an 8B parameter multimodal model hosted on Hugging Face. It belongs to the class of foundation models and accepts both image and video inputs alongside text.
The model defines a chat template that processes messages containing text, image, or video content. When an image appears the template inserts an image context token, while a video triggers a dedicated video token. The template also handles role-based formatting with special start and end markers for each turn and supplies an assistant prompt when generation is required. A processor configuration specifies an end-of-text token and leaves the unknown token undefined.
It is delivered as a repository on the Hugging Face platform under the identifier OpenGVLab/InternVL3-8B-hf. The repository was created on 18 April 2025 and last modified on 23 April 2025. It has recorded more than 432000 downloads since upload and remains publicly accessible for download and inference.
No pricing, licensing terms, or usage restrictions are stated in the repository metadata.
InternVL3 8B Hf is a Foundation models & chat project. It focuses on providing a strong open-source multimodal model for image and video understanding tasks. It is built as an open-source project for AI researchers and application developers. InternVL3 8B Hf is open source under the Open Source license. It runs on the web, the command line, and API.
Behind InternVL3 8B Hf is OpenGVLab, and it first shipped in 2025. It operates in a well-populated space: PulseGate tracks 11 similar projects. Key capabilities include vision-Language, Video Support, and Hugging Face Format. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do