CLIP ViT H 14 laion2B s32B b79K Alternatives
CLIP-ViT-H-14-laion2B-s32B-b79K is a large vision transformer model trained using the CLIP objective on the LAION-2B dataset. Below are 20 other ai apps with similar functionality to CLIP ViT H 14 laion2B s32B b79K, matched by what each product actually does — not ranked or scored. Explore each to find the closest fit for your use case.
- CLIP ViT L 14 laion2B s32B b82Khuggingface.co
This is an open-weight CLIP model variant trained by LAION on the LAION-2B dataset. It maps images and text into a shared embedding space for tasks such as zero-shot image classification, image-text retrieval, and multimodal similarity scoring. The model is distributed on Hugging Face and can be used with the OpenCLIP or transformers libraries.
- CLIP ViT B 32 laion2B s34B b79Khuggingface.co
laion/CLIP-ViT-B-32-laion2B-s34B-b79K is an open-source vision-language model that generates embeddings for both images and text. It is widely used for zero-shot classification, search, and retrieval tasks, providing a robust foundation for multimodal AI applications. The model is suitable for researchers and developers building advanced AI systems.
- CLIP ViT B 16 laion2B s34B b88Khuggingface.co
CLIP-ViT-B-16-laion2B-s34B-b88K is a vision transformer model trained with the CLIP objective on the large-scale LAION-2B dataset. It produces aligned embeddings for images and text that support zero-shot classification, retrieval, and similarity tasks. The model weights are publicly available on Hugging Face for research and commercial applications.
- CLIP ViT bigG 14 laion2B 39B B160khuggingface.co
This is a large ViT-bigG/14 CLIP model trained by LAION on the LAION-2B dataset. It produces powerful joint embeddings for images and text, enabling zero-shot image classification, retrieval, and other multimodal tasks. The model is widely used in the open-source AI community and available on Hugging Face.
- CLIP ViT B 32 Xlm Roberta Base laion5B s13B B90khuggingface.co
This LAION model combines a ViT-B/32 image encoder with an XLM-RoBERTa-base text encoder, trained on the LAION-5B dataset. It is part of the OpenCLIP ecosystem and supports cross-lingual image-text similarity, retrieval, and zero-shot classification. The model is distributed as open weights and integrates directly with the OpenCLIP library.
- CLIP Convnext Base laion400M s13B b51Khuggingface.co
This is an OpenCLIP model based on ConvNeXt architecture trained on the LAION-400M dataset with 13B samples seen. It produces embeddings for images and text that can be used for zero-shot classification, image-text retrieval, and other multimodal tasks. The model is provided as open weights on Hugging Face for use with the OpenCLIP library and is intended for researchers and developers building vision-language applications.
- CLIP Convnext Base W laion2B s13B b82Khuggingface.co
This is a CLIP model that uses a ConvNeXt base visual backbone trained on the LAION-2B dataset. It produces aligned image and text embeddings suitable for zero-shot image classification, retrieval, and multimodal tasks. The model is provided by LAION on Hugging Face.
- Vit Base Patch32 Clip 224.laion400m E32huggingface.co
This is a Vision Transformer (ViT) model trained with the CLIP objective on the LAION-400M dataset. It enables zero-shot image classification and image-text similarity computation. The model is provided through the timm library and OpenCLIP on Hugging Face.
- CLIP Convnext Large D 320.laion2B s29B b131K Ft Souphuggingface.co
This model is a large ConvNeXt-based CLIP variant trained on the LAION-2B dataset with 29 billion samples and further fine-tuned. It maps images and text into a shared embedding space for tasks such as zero-shot classification, retrieval, and similarity computation. It is distributed on Hugging Face for use by machine learning developers building multimodal AI applications.
- CLIP ViT B 16 DataComp.XL s13B b90Khuggingface.co
This is a CLIP ViT-B/16 model trained on the DataComp XL dataset (s13B-b90K). It is designed for vision-language tasks including zero-shot image classification and retrieval. The model is available on Hugging Face and uses a standard CLIP architecture for aligning image and text embeddings in a shared space.
- Vit Base Patch32 Clip 224.laion400m E31huggingface.co
This model is a Vision Transformer (ViT) trained with the CLIP objective on the LAION-400M dataset using the OpenCLIP library. It is hosted in the timm model hub and supports zero-shot image classification as well as image and text embedding generation for downstream computer vision tasks.
- Vit Base Patch16 Clip 224.laion400m E31huggingface.co
This model is a Vision Transformer (ViT) base with patch size 16 trained using the CLIP objective on the LAION-400M dataset (epoch 31). It is part of the timm library and supports zero-shot image classification and image embedding generation. The model is distributed on Hugging Face for use with the OpenCLIP library.
- DFN2B CLIP ViT B 16huggingface.co
This is an Apple-developed variant of the CLIP model using a ViT-B/16 vision transformer backbone and a Data Filtering Network (DFN). It produces aligned embeddings for images and text, enabling zero-shot classification, retrieval, and other vision-language tasks. The model is available on Hugging Face for use with the Transformers library.
- DFN5B CLIP ViT H 14huggingface.co
DFN5B-CLIP-ViT-H-14 is a large-scale contrastive language-image pretraining (CLIP) model released by Apple on Hugging Face. It features a Vision Transformer (ViT-H/14) backbone trained on the DFN5B dataset. The model produces aligned embeddings for images and text, enabling zero-shot classification, retrieval, and other multimodal tasks. It is distributed as open weights and can be used with standard CLIP inference pipelines. The model is targeted at researchers and engineers building computer vision and multimodal applications.
- Clip Vit Large Patch14 336huggingface.co
clip-vit-large-patch14-336 is an open-source vision-language model by OpenAI that enables joint image and text embedding for tasks such as zero-shot classification and retrieval. It is widely used by AI researchers and developers for multimodal applications.
- Convnext Base.clip Laion2b Augreg Ft In12k In1khuggingface.co
This ConvNeXt Base model has been pre-trained using CLIP on LAION-2B data and fine-tuned on ImageNet-12k and ImageNet-1k. It is part of the PyTorch Image Models (timm) library and can be used for image classification and as a feature extractor. The model offers strong performance on standard vision benchmarks.
- Convnext Base.clip Laion2bhuggingface.co
This model is a ConvNeXt Base image encoder trained with CLIP on the LAION-2B dataset and provided through the timm library. It generates rich feature embeddings suitable for image-text similarity, zero-shot classification, and retrieval. It integrates with both timm and Hugging Face Transformers for easy use in computer vision pipelines.
- ViT B 32 Openaihuggingface.co
This repository provides ONNX exports of the OpenCLIP ViT-B-32 model originally from OpenAI. It is specifically intended for use within Immich, a self-hosted photo and video library management application. The model enables image feature extraction and semantic search capabilities.
- Vit Base Patch16 Clip 224.openaihuggingface.co
This is a Vision Transformer (ViT-B/16 at 224px) model from the timm library, loaded with OpenAI CLIP weights for image feature extraction. It is compatible with both timm and Transformers libraries and has been widely downloaded. The model produces embeddings suitable for similarity search, classification, or other computer vision downstream tasks.
- Clip Vit Base Patch32huggingface.co
A port of OpenAI's CLIP ViT-B/32 model by Xenova, optimized for execution in JavaScript environments via ONNX Runtime or WebNN. It enables zero-shot image classification, image-text similarity, and multimodal embeddings directly in the browser or on the server with Node.js. The model is widely used for client-side computer vision tasks.