M-CLIP/XLM-Roberta-Large-Vit-B-16Plus extends OpenAI's CLIP to support 48+ languages by pairing a multilingual XLM-RoBERTa text encoder with a Vision Transformer image encoder. It enables cross-lingual image-text retrieval and zero-shot classification. The model is available through Hugging Face Transformers for research and production use.
In the Other AI space, XLM Roberta Large Vit B 16Plus takes a focused approach. It focuses on creating multilingual vision-language representations beyond English-only CLIP models. XLM Roberta Large Vit B 16Plus is an open-source project aimed at developers. The project is open source (Open Source). XLM Roberta Large Vit B 16Plus is available on the web and the command line, and it can be self-hosted.
Behind XLM Roberta Large Vit B 16Plus is Multilingual-CLIP, and the product first shipped in 2021. The project is developed in the open on GitHub with 14k stars and 110 commits in the last 90 days. Among its 3 catalogued features are Multilingual Text, Image-Text Alignment, and Zero-Shot Retrieval.
Latest indexed changes and source events
M-CLIP/XLM-Roberta-Large-Vit-B-16Plus verified by the PulseGate indexer
Other apps tracked under the same category.