LLaVA-OneVision-2-8B-Instruct is an open multimodal model that understands images, video, and text. It follows instructions and can be used for a variety of vision-language tasks. The model is distributed on Hugging Face and is intended for researchers and developers building multimodal AI applications.
LLaVA OneVision 2 8B Instruct is a Foundation models & chat project. It focuses on providing open-source multimodal understanding capabilities for images, video, and text in a single model. It is built as an open-source project for developers. LLaVA OneVision 2 8B Instruct is open source under the Open Source license. LLaVA OneVision 2 8B Instruct is available on the web, the command line, and API.
It is developed by lmms-lab, and it first shipped in 2025. Key capabilities include Vision-Language Understanding, Image Processing, and Video Understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do