This is a specialized version of the Qwen2.5-VL-7B vision-language model adapted for video classification. It includes a custom chat template that supports both image and video inputs with appropriate vision and video padding tokens. The model can process video content and classify it according to trained categories.
In the Foundation models & chat space, Qwen2.5 VL 7B For VideoCls takes a focused approach. It focuses on adapting general vision-language models for specialized video understanding and classification. It is built as an open-source project for computer vision researchers and developers. The project is open source (Open Source). It runs on the web and API.
muziyongshixin builds and maintains Qwen2.5 VL 7B For VideoCls, and it first shipped in 2025. Among its 4 catalogued features are Video Classification, vision-Language, and multimodal. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do