Florence-2-large is a large-scale vision-language foundation model hosted on Hugging Face. It supports a wide range of multimodal tasks including image captioning, visual question answering, object detection, and more through a unified image-text-to-text interface. The model is designed for developers and researchers building computer vision and multimodal applications using the Transformers library.
Florence 2 Large sits in PulseGate's Foundation models & chat category. It focuses on accessing and using a powerful open-source vision-language model for multimodal AI tasks. Florence 2 Large is an open-source project aimed at developers. The project is open source (Open Source). It runs on the web, the command line, and API.
florence-community builds and maintains Florence 2 Large, and the product first shipped in 2024. Among its 5 catalogued features are vision-language understanding, image captioning, and visual question answering.
Latest indexed changes and source events
florence-community/Florence-2-large verified by the PulseGate indexer
Other apps tracked under the same category.