This model is a SigLIP variant of a Vision Transformer (ViT-B/16) from the timm library. It supports zero-shot image classification and is trained using contrastive learning on web-scale data. It can be used through open_clip or the transformers ecosystem.
In the Other AI space, ViT B 16 SigLIP takes a focused approach. It focuses on enabling zero-shot image classification and retrieval without task-specific training. It is built as an open-source project for developers. The project is open source (Apache-2.0). It ships for the web and API.
Behind ViT B 16 SigLIP is timm, and it first shipped in 2022. Development happens publicly on GitHub with 3.5k stars. Among its 3 catalogued features are Zero-shot Classification, Vision Transformer, and sigLIP.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do