Florence-2-base-ft is a vision-language model hosted on Hugging Face for image-text-to-text tasks. Developed by Microsoft and released under the MIT license, it forms part of the Florence-2 family and is provided as an openly accessible model card with associated weights.
The model can be loaded and run through the Transformers library from Hugging Face. Code examples show instantiation via a high-level pipeline for image-text-to-text or through direct loading of an AutoProcessor and AutoModelForMultimodalLM with trust_remote_code enabled. It supports deployment using vLLM after installation from pip. Additional integration paths include Google Colab and Kaggle notebooks as well as local applications.
The underlying implementation relies on PyTorch and Safetensors formats. Its model card references the arXiv preprint 2311.06242. The page indicates 20.8k likes and positions the model within collections related to vision and custom code capabilities.
It is intended for developers and researchers working with multimodal inputs who require a base fine-tuned checkpoint that can be adapted for specific image-to-text workflows.
Florence 2 Base Ft sits in PulseGate's Foundation models & chat category. It focuses on accessing a fine-tuned open-source vision-language model for advanced image understanding tasks. Florence 2 Base Ft is an open-source project aimed at developers. The project is open source (Open Source). The product ships for the web, the command line, and API.
It is developed by Microsoft, and the product first shipped in 2024. Among its 4 catalogued features are vision-language processing, fine-tuned capabilities, and image understanding.
Latest indexed changes and source events
microsoft/Florence-2-base-ft verified by the PulseGate indexer
Other apps tracked under the same category.