CLIP-ViT-B-16-laion2B-s34B-b88K is a vision transformer model trained with the CLIP objective on the large-scale LAION-2B dataset. It produces aligned embeddings for images and text that support zero-shot classification, retrieval, and similarity tasks. The model weights are publicly available on Hugging Face for research and commercial applications.
CLIP ViT B 16 laion2B s34B b88K is an Embeddings & retrieval project. It focuses on creating robust open-source multimodal embeddings for image and text without relying on proprietary training data. It is built as an open-source project for AI researchers and developers. The project is open source (Open Source). It ships for the web and API.
LAION builds and maintains CLIP ViT B 16 laion2B s34B b88K, and it first shipped in 2021. The project is developed in the open on GitHub with 14k stars and 115 commits in the last 90 days. Key capabilities include Image-Text Alignment, Zero-Shot Classification, and Contrastive Learning.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do