PaliGemma-3B-PT-224 is an open-source 3 billion parameter vision-language model developed by Google. It is designed for a wide range of image-to-text and multimodal tasks including visual question answering, captioning, and visual reasoning. The model is available on Hugging Face for local or cloud inference and is intended for developers and researchers working on computer vision and multimodal applications.
Paligemma 3b Pt 224 is a Foundation models & chat product. It focuses on enabling open-source multimodal AI that can reason jointly over images and text without relying on closed APIs. Paligemma 3b Pt 224 is an open-source project aimed at developers and researchers. The project is open source (Apache-2.0). The product ships for the web, the command line, and API, and it can be self-hosted.
It is developed by Google (United States), and the product first shipped in 2018. The project is developed in the open on GitHub with 36k stars and 1.9k commits in the last 90 days. Among its 4 catalogued features are Image Understanding, Visual Question Answering, and Multimodal Reasoning.
Latest indexed changes and source events
google/paligemma-3b-pt-224 verified by the PulseGate indexer
Other apps tracked under the same category.