CLIP-ViT-H-14-laion2B-s32B-b79K is a large vision transformer model trained using the CLIP objective on the LAION-2B dataset. It maps images and text into a shared embedding space, enabling zero-shot image classification, image-text retrieval, and other multimodal tasks. The model is fully open-source with publicly available weights on Hugging Face, allowing researchers and developers to use, fine-tune, or run inference locally or via cloud providers.
CLIP ViT H 14 laion2B s32B b79K is a Multimodal & vision project. It focuses on training large-scale multimodal models without proprietary datasets or closed-source weights. It is built as an open-source project for AI researchers and developers. The project is open source (Open Source). It ships for the web and API.
Behind CLIP ViT H 14 laion2B s32B b79K is LAION, based in Germany, and it first shipped in 2021. The project is developed in the open on GitHub with 14k stars and 115 commits in the last 90 days. Among its 4 catalogued features are Vision-Language Model, Contrastive Pretraining, and Image-Text Retrieval.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do