This is an Apple-developed variant of the CLIP model using a ViT-B/16 vision transformer backbone and a Data Filtering Network (DFN). It produces aligned embeddings for images and text, enabling zero-shot classification, retrieval, and other vision-language tasks. The model is available on Hugging Face for use with the Transformers library.
In the Other AI space, DFN2B CLIP ViT B 16 takes a focused approach. It focuses on aligning visual and textual representations for zero-shot image classification and retrieval. DFN2B CLIP ViT B 16 is an open-source project aimed at developers. The project is open source (Open Source). The product ships for the web and API.
It is developed by Apple, and the product first shipped in 2023.
Latest indexed changes and source events
apple/DFN2B-CLIP-ViT-B-16 verified by the PulseGate indexer
Other apps tracked under the same category.