CLIP-ViT-H-14-laion2B-s32B-b79K is a large vision transformer model trained using the CLIP objective on the LAION-2B dataset. It maps images and text into a shared embedding space, enabling zero-shot image classification, image-text retrieval, and other multimodal tasks. The model is fully open-source with publicly available weights on Hugging Face, allowing researchers and developers to use, fine-tune, or run inference locally or via cloud providers.
In the Other AI space, CLIP ViT H 14 laion2B s32B b79K takes a focused approach. It focuses on training large-scale multimodal models without proprietary datasets or closed-source weights. CLIP ViT H 14 laion2B s32B b79K is an open-source project aimed at AI researchers and developers. The project is open source (Open Source). It runs on the web and API.
It is developed by LAION (Germany), and the product first shipped in 2021. The project is developed in the open on GitHub with 14k stars and 115 commits in the last 90 days. Among its 4 catalogued features are Vision-Language Model, Contrastive Pretraining, and Image-Text Retrieval.
Latest indexed changes and source events
laion/CLIP-ViT-H-14-laion2B-s32B-b79K verified by the PulseGate indexer
Other apps tracked under the same category.