SigLIP2 is Google's open-weight vision-language model that maps images and text into a shared embedding space. The so400m-patch16-384 variant processes 384x384 images using a 16x16 patch size. It supports tasks such as image-text similarity, zero-shot classification, and retrieval. The model is distributed on Hugging Face and can be used locally or via inference providers.
In the Other AI space, Siglip2 So400m Patch16 384 takes a focused approach. It focuses on aligning images and text in a shared embedding space for zero-shot classification and retrieval tasks. It is built as an open-source project for developers. Siglip2 So400m Patch16 384 is open source under the Open Source license. It runs on the web and API.
Google builds and maintains Siglip2 So400m Patch16 384, and the product first shipped in 2025. Key capabilities include vision-Language, Image-Text Alignment, and Patch Embedding.
Latest indexed changes and source events
google/siglip2-so400m-patch16-384 verified by the PulseGate indexer