LLaVA v1.5-7B is a vision-language model that connects a CLIP vision encoder with a Vicuna language model. It can answer questions about images, describe visual content, and follow complex multimodal instructions. The model is widely used for research and local multimodal applications.
In the Foundation models & chat space, Llava V1.5 7b takes a focused approach. It focuses on enabling language models to reason about the content of images alongside text instructions. It is built as an open-source project for developers. Llava V1.5 7b is open source under the Apache-2.0 license. The product ships for the web, the command line, and API.
Behind Llava V1.5 7b is liuhaotian, and the product first shipped in 2023. Development happens publicly on GitHub with 24.9k stars.
Latest indexed changes and source events
liuhaotian/llava-v1.5-7b verified by the PulseGate indexer
Other apps tracked under the same category.