LLaVA v1.6 Vicuna 13B is an open-source multimodal model combining a vision encoder with a Vicuna language model for image-to-text generation and visual reasoning tasks. It supports visual question answering, image captioning, and complex multimodal instructions. Hosted on Hugging Face, it is used by developers and researchers via the Transformers library for local or self-hosted inference.
Llava V1.6 Vicuna 13b sits in PulseGate's Multimodal & vision category. It focuses on running state-of-the-art open multimodal vision-language models locally without proprietary APIs. It is built as an open-source project for developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
It is developed by liuhaotian, and it first shipped in 2023. The project is developed in the open on GitHub with 24.9k stars. Key capabilities include Image-Text Understanding, Visual Question Answering, and Multimodal Reasoning.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do