LLaVA-NeXT-Video-7B-hf is an open-source multimodal model that extends LLaVA to video understanding. It accepts video input along with text prompts and can answer questions, describe scenes, or follow instructions related to video content. The model is provided with full weights on Hugging Face and supports standard inference pipelines for research and application development in video-language tasks.
LLaVA NeXT Video 7B Hf sits in PulseGate's Foundation models & chat category. It focuses on enabling large language models to comprehend and reason about video content alongside text instructions. LLaVA NeXT Video 7B Hf is an open-source project aimed at developers. The project is open source (Apache-2.0). LLaVA NeXT Video 7B Hf is available on the web, the command line, and API.
LLaVA builds and maintains LLaVA NeXT Video 7B Hf, and the product first shipped in 2024. The project is developed in the open on GitHub with 4.7k stars and 2 commits in the last 90 days. Among its 3 catalogued features are video understanding, multimodal input, and visual question answering.
Latest indexed changes and source events
llava-hf/LLaVA-NeXT-Video-7B-hf verified by the PulseGate indexer
Other apps tracked under the same category.