Florence-2 for Videos is a Hugging Face Space that processes uploaded videos by first creating a caption for the entire clip and then detecting and tracking referenced objects. It overlays bounding boxes and labels on the video output. The tool is useful for developers and researchers working on multimodal AI, video analysis, and vision-language models.
Florence-2 for Videos sits in PulseGate's AI & ML category. Automatically generating captions and tracking mentioned objects across video frames. Florence-2 for Videos is an open-source project aimed at developers. It is available for free. It runs on the web, and it can be self-hosted.
Behind Florence-2 for Videos is SkalskiP, and it first shipped in 2024. Among its 4 catalogued features are Video Upload, Auto Captioning, and Object Tracking.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do