ColPali v1.2 is a multimodal model specialized in visual document retrieval. It processes document images to create rich embeddings that capture both textual and visual information, enabling more accurate retrieval than text-only approaches. It is part of the ViDoRe project and available through the Hugging Face ecosystem.
Colpali is a Foundation models & chat product. It focuses on retrieving relevant document pages using both text and visual information. It is built as an open-source project for machine learning developers. Colpali is open source under the MIT license. The product ships for the web, the command line, and API.
vidore builds and maintains Colpali, and the product first shipped in 2024. Development happens publicly on GitHub with 1 commits in the last 90 days. Key capabilities include Visual Document Retrieval, Multimodal Embeddings, and ColPali Architecture.
Latest indexed changes and source events
vidore/colpali-v1.2 verified by the PulseGate indexer
Other apps tracked under the same category.