This model is a SigLIP2-trained Vision Transformer (ViT) with 400 million parameters and 384px resolution, hosted in the timm library. It excels at zero-shot image classification and producing high-quality image embeddings for downstream computer vision tasks. The openly available weights allow researchers and engineers to leverage state-of-the-art contrastive vision capabilities.
In the Other AI space, ViT SO400M 16 SigLIP2 384 takes a focused approach. It focuses on performing zero-shot image classification and retrieval without task-specific training data. ViT SO400M 16 SigLIP2 384 is an open-source project aimed at developers. The project is open source (Apache-2.0). The product ships for the web and API.
It is developed by timm, and the product first shipped in 2022. The project is developed in the open on GitHub with 3.5k stars. Among its 3 catalogued features are zero-shot classification, image embeddings, and sigLIP2 training.
Latest indexed changes and source events
timm/ViT-SO400M-16-SigLIP2-384 verified by the PulseGate indexer