This is a Vision Transformer (ViT-B/16 at 224px) model from the timm library, loaded with OpenAI CLIP weights for image feature extraction. It is compatible with both timm and Transformers libraries and has been widely downloaded. The model produces embeddings suitable for similarity search, classification, or other computer vision downstream tasks.
In the Other AI space, Vit Base Patch16 Clip 224.openai takes a focused approach. It focuses on obtaining high-quality image embeddings or features from a CLIP-pretrained Vision Transformer without training it yourself. It is built as an open-source project for developers. Vit Base Patch16 Clip 224.openai is open source under the Apache-2.0 license. The product ships for the web and API.
It is developed by timm, and the product first shipped in 2019. Development happens publicly on GitHub with 37k stars and 49 commits in the last 90 days. Key capabilities include Image Feature Extraction, CLIP Vision Encoder, and pyTorch.
Latest indexed changes and source events
timm/vit_base_patch16_clip_224.openai verified by the PulseGate indexer
Other apps tracked under the same category.