DFN5B-CLIP-ViT-H-14 is a large-scale contrastive language-image pretraining (CLIP) model released by Apple on Hugging Face. It features a Vision Transformer (ViT-H/14) backbone trained on the DFN5B dataset. The model produces aligned embeddings for images and text, enabling zero-shot classification, retrieval, and other multimodal tasks. It is distributed as open weights and can be used with standard CLIP inference pipelines. The model is targeted at researchers and engineers building computer vision and multimodal applications.
DFN5B CLIP ViT H 14 is an Other AI product. It focuses on providing high-quality open-weight vision-language embeddings for image-text retrieval and multimodal AI applications. It is built as an open-source project for AI researchers and developers. DFN5B CLIP ViT H 14 is open source under the Apache-2.0 license. It runs on the web, and it can be self-hosted.
It is developed by Apple (United States), and the product first shipped in 2023. Development happens publicly on GitHub with 2.4k stars and 155 commits in the last 90 days. Key capabilities include Image-Text Alignment, Vision Encoding, and Contrastive Learning.
Latest indexed changes and source events
apple/DFN5B-CLIP-ViT-H-14 verified by the PulseGate indexer
Other apps tracked under the same category.