This is a Vision Transformer (ViT-B/16 at 224px) model from the timm library, loaded with OpenAI CLIP weights for image feature extraction. It is compatible with both timm and Transformers libraries and has been widely downloaded. The model produces embeddings suitable for similarity search, classification, or other computer vision downstream tasks.
Vit Base Patch16 Clip 224.openai is an Other AI project. It focuses on obtaining high-quality image embeddings or features from a CLIP-pretrained Vision Transformer without training it yourself. It is built as an open-source project for developers. The project is open source (Apache-2.0). It ships for the web and API.
timm builds and maintains Vit Base Patch16 Clip 224.openai, and it first shipped in 2019. Development happens publicly on GitHub with 37k stars and 49 commits in the last 90 days. Key capabilities include Image Feature Extraction, CLIP Vision Encoder, and pyTorch.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do