PaliGemma-3B is a 3-billion parameter vision-language model from Google, fine-tuned on the COCO captions dataset at 448px resolution. It excels at image-to-text tasks including captioning and visual question answering. The model is distributed openly on Hugging Face for use with the Transformers library.
Paligemma 3b Ft Cococap 448 sits in PulseGate's Foundation models & chat category. It focuses on generating accurate textual descriptions and answers from images using a vision-language model. Paligemma 3b Ft Cococap 448 is an open-source project aimed at developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
Google builds and maintains Paligemma 3b Ft Cococap 448, and the product first shipped in 2022. The project is developed in the open on GitHub with 3.5k stars. Among its 3 catalogued features are image captioning, visual question answering, and image-text-to-text.
Latest indexed changes and source events
google/paligemma-3b-ft-cococap-448 verified by the PulseGate indexer
Other apps tracked under the same category.