InternVL3_5-30B-A3B is a multimodal foundation model hosted on Hugging Face by OpenGVLab. It processes text along with image and video inputs through a unified chat template that inserts special tokens for different content types.
The model uses a specific tokenizer configuration with defined pad, bos, and eos tokens. Its chat template supports messages that interleave text with image and video markers, followed by an assistant generation prompt. This structure enables the model to handle mixed-modality conversations in a consistent format.
The repository provides the model weights and configuration files for download and local use. It records over 110,000 recent downloads and nearly 760,000 downloads in total. The model was created on August 25 2025 and last modified on August 29 2025.
It belongs to the class of foundation models and is delivered as a downloadable asset on the Hugging Face platform. No pricing, licensing details, or specific task benchmarks are stated in the repository metadata.
InternVL3 5 30B A3B sits in PulseGate's Multimodal & vision category. It focuses on processing and understanding mixed image, video, and text inputs at scale with a single model. It is built as an open-source project for computer vision and multimodal AI researchers. The project is open source (MIT). It runs on the web, the command line, and API.
Behind InternVL3 5 30B A3B is OpenGVLab, and it first shipped in 2023. The project is developed in the open on GitHub with 10.1k stars. Key capabilities include Multimodal Understanding, Image Processing, and Video Processing. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do