This model is a Vision Transformer (ViT) base with patch size 16 trained using the CLIP objective on the LAION-400M dataset (epoch 31). It is part of the timm library and supports zero-shot image classification and image embedding generation. The model is distributed on Hugging Face for use with the OpenCLIP library.
In the Other AI space, Vit Base Patch16 Clip 224.laion400m E31 takes a focused approach. It focuses on providing pre-trained CLIP vision encoders for zero-shot classification and image-text similarity tasks. It is built as an open-source project for computer vision researchers and developers. Vit Base Patch16 Clip 224.laion400m E31 is open source under the Open Source license. Vit Base Patch16 Clip 224.laion400m E31 is available on the web and API.
timm builds and maintains Vit Base Patch16 Clip 224.laion400m E31. Key capabilities include Zero-shot Image Classification, CLIP Embeddings, and Vision Transformer.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do