This LAION model combines a ViT-B/32 image encoder with an XLM-RoBERTa-base text encoder, trained on the LAION-5B dataset. It is part of the OpenCLIP ecosystem and supports cross-lingual image-text similarity, retrieval, and zero-shot classification. The model is distributed as open weights and integrates directly with the OpenCLIP library.
In the Multimodal & vision space, CLIP ViT B 32 Xlm Roberta Base laion5B s13B B90k takes a focused approach. It focuses on enabling zero-shot image classification and retrieval across many languages using contrastive vision-language embeddings. It is built as an open-source project for developers. CLIP ViT B 32 Xlm Roberta Base laion5B s13B B90k is open source under the Open Source license. It runs on the web and API.
Behind CLIP ViT B 32 Xlm Roberta Base laion5B s13B B90k is LAION, based in Germany, and it first shipped in 2021. The project is developed in the open on GitHub with 14k stars and 110 commits in the last 90 days. Among its 3 catalogued features are contrastive image-text model, multilingual text encoder, and openCLIP compatible.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do