LLaVA v1.5-7B is a vision-language model that connects a CLIP vision encoder with a Vicuna language model. It can answer questions about images, describe visual content, and follow complex multimodal instructions. The model is widely used for research and local multimodal applications.
Llava V1.5 7b sits in PulseGate's Multimodal & vision category. It focuses on enabling language models to reason about the content of images alongside text instructions. Llava V1.5 7b is an open-source project aimed at developers. Llava V1.5 7b is open source under the Apache-2.0 license. It ships for the web, the command line, and API.
liuhaotian builds and maintains Llava V1.5 7b, and it first shipped in 2023. The project is developed in the open on GitHub with 24.9k stars.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do