LLaVA-NeXT-Video-7B-hf is an open-source multimodal model that extends LLaVA to video understanding. It accepts video input along with text prompts and can answer questions, describe scenes, or follow instructions related to video content. The model is provided with full weights on Hugging Face and supports standard inference pipelines for research and application development in video-language tasks.
LLaVA NeXT Video 7B Hf is a Multimodal & vision project. It focuses on enabling large language models to comprehend and reason about video content alongside text instructions. LLaVA NeXT Video 7B Hf is an open-source project aimed at developers. LLaVA NeXT Video 7B Hf is open source under the Apache-2.0 license. LLaVA NeXT Video 7B Hf is available on the web, the command line, and API.
LLaVA builds and maintains LLaVA NeXT Video 7B Hf, and it first shipped in 2024. The project is developed in the open on GitHub with 4.7k stars and 2 commits in the last 90 days. Among its 3 catalogued features are video understanding, multimodal input, and visual question answering.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do