LLaVA v1.6 Vicuna 13B is an open-source multimodal model combining a vision encoder with a Vicuna language model for image-to-text generation and visual reasoning tasks. It supports visual question answering, image captioning, and complex multimodal instructions. Hosted on Hugging Face, it is used by developers and researchers via the Transformers library for local or self-hosted inference.
Llava V1.6 Vicuna 13b is a Foundation models & chat product. It focuses on running state-of-the-art open multimodal vision-language models locally without proprietary APIs. It is built as an open-source project for developers. Llava V1.6 Vicuna 13b is open source under the Apache-2.0 license. The product ships for the web, the command line, and API.
It is developed by liuhaotian, and the product first shipped in 2023. Development happens publicly on GitHub with 24.9k stars. Key capabilities include Image-Text Understanding, Visual Question Answering, and Multimodal Reasoning.
Latest indexed changes and source events
liuhaotian/llava-v1.6-vicuna-13b verified by the PulseGate indexer
Other apps tracked under the same category.