Tarsier-7B is a 7 billion parameter vision-language model from omni-research. Built on the LLaVA architecture, it supports text generation conditioned on image or video inputs. The model is available on Hugging Face and can be used with the Transformers library for multimodal research and applications.
In the Foundation models & chat space, Tarsier 7b takes a focused approach. It focuses on enabling large language models to understand and reason about visual content in images and videos. Tarsier 7b is an open-source project aimed at AI researchers and developers. The project is open source (Apache-2.0). The product ships for the web, the command line, and API.
Omni Research builds and maintains Tarsier 7b, and the product first shipped in 2024. The project is developed in the open on GitHub with 548 stars.
Latest indexed changes and source events
omni-research/Tarsier-7b verified by the PulseGate indexer
Other apps tracked under the same category.