InternVL3-1B-hf is a 1-billion-parameter vision-language model developed by OpenGVLab. It supports multimodal inputs including images, video, and text, and uses a custom chat template for conversational interactions. The model is available in Hugging Face format and can be used for various vision-language tasks.
InternVL3 1B Hf is a Multimodal & vision project. It focuses on understanding and reasoning over both images/videos and text in a single unified model. It is built as an open-source project for AI researchers and developers. The project is open source (Open Source). It ships for the web, the command line, and API.
OpenGVLab builds and maintains InternVL3 1B Hf. It operates in a well-populated space: PulseGate tracks 5 similar projects. Key capabilities include vision-Language, multimodal, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do