A 3B parameter fine-tuned version of PaliGemma 2 specialized for document-centric image-to-text tasks (DOCCI). It processes images and text prompts to generate relevant textual outputs and is distributed via the Transformers library for easy integration into vision-language applications.
Paligemma2 3b Ft Docci 448 is a Foundation models & chat product. It focuses on understanding and extracting information from document images using vision-language models. Paligemma2 3b Ft Docci 448 is an open-source project aimed at developers. The project is open source (Apache-2.0). Paligemma2 3b Ft Docci 448 is available on the web and the command line, and it can be self-hosted.
google builds and maintains Paligemma2 3b Ft Docci 448, and the product first shipped in 2022. The project is developed in the open on GitHub with 3.5k stars. Among its 3 catalogued features are Image-Text Understanding, Document AI, and vision-Language.
Latest indexed changes and source events
google/paligemma2-3b-ft-docci-448 verified by the PulseGate indexer
Other apps tracked under the same category.