InternVL3_5-30B-A3B is a multimodal foundation model hosted on Hugging Face by OpenGVLab. It processes text along with image and video inputs through a unified chat template that inserts special tokens for different content types.
The model uses a specific tokenizer configuration with defined pad, bos, and eos tokens. Its chat template supports messages that interleave text with image and video markers, followed by an assistant generation prompt. This structure enables the model to handle mixed-modality conversations in a consistent format.
The repository provides the model weights and configuration files for download and local use. It records over 110,000 recent downloads and nearly 760,000 downloads in total. The model was created on August 25 2025 and last modified on August 29 2025.
It belongs to the class of foundation models and is delivered as a downloadable asset on the Hugging Face platform. No pricing, licensing details, or specific task benchmarks are stated in the repository metadata.
InternVL3 5 30B A3B sits in PulseGate's Foundation models & chat category. It focuses on processing and understanding mixed image, video, and text inputs at scale with a single model. InternVL3 5 30B A3B is an open-source project aimed at computer vision and multimodal AI researchers. The project is open source (MIT). The product ships for the web, the command line, and API.
Behind InternVL3 5 30B A3B is OpenGVLab, and the product first shipped in 2023. The project is developed in the open on GitHub with 10.1k stars. Among its 4 catalogued features are Multimodal Understanding, Image Processing, and Video Processing. It exposes integrations via a public API.
Latest indexed changes and source events
OpenGVLab/InternVL3_5-30B-A3B verified by the PulseGate indexer
Other apps tracked under the same category.