M-CLIP/XLM-Roberta-Large-Vit-B-16Plus extends OpenAI's CLIP to support 48+ languages by pairing a multilingual XLM-RoBERTa text encoder with a Vision Transformer image encoder. It enables cross-lingual image-text retrieval and zero-shot classification. The model is available through Hugging Face Transformers for research and production use.
In the Multimodal & vision space, XLM Roberta Large Vit B 16Plus takes a focused approach. It focuses on creating multilingual vision-language representations beyond English-only CLIP models. XLM Roberta Large Vit B 16Plus is an open-source project aimed at developers. The project is open source (Open Source). It runs on the web and the command line, and it can be self-hosted.
Behind XLM Roberta Large Vit B 16Plus is Multilingual-CLIP, and it first shipped in 2021. The project is developed in the open on GitHub with 14k stars and 110 commits in the last 90 days. Key capabilities include Multilingual Text, Image-Text Alignment, and Zero-Shot Retrieval.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do