Minilm Bo En Sim is a sentence similarity model published on Hugging Face for Tibetan-English cross-lingual tasks. It generates vector representations of sentences that support computation of semantic similarity scores between text in Tibetan and English.
The model is based on the MiniLM architecture and is tagged for use with the sentence-transformers library. It accepts lists of sentences as input, encodes them into embeddings, and can compute a similarity matrix directly from those embeddings. Example code demonstrates loading the model, encoding English sentences, and producing a 3-by-3 similarity tensor.
It is listed under the sentence-transformers framework and carries a BSD license. The repository includes model files in Safetensors format and references the arXiv paper 1908.10084. The page provides instructions for integration with inference providers, Google Colab, and Kaggle notebooks.
Khyentse Vision Project maintains the model card, which outlines intended use, model details, evaluation metrics, and training information without further elaboration on those sections in the visible excerpt.
In the Foundation models & chat space, Minilm Bo En Sim takes a focused approach. It focuses on computing semantic similarity between Tibetan and English sentences for cross-lingual applications. It is built as an open-source project for developers. Minilm Bo En Sim is open source under the Apache-2.0 license. It runs on the web and API.
It is developed by Khyentse Vision Project, and the product first shipped in 2019. Development happens publicly on GitHub with 18.9k stars and 96 commits in the last 90 days. Key capabilities include Sentence Embeddings, Cross-lingual Similarity, and Tibetan Support.
Latest indexed changes and source events
khyentsevision/minilm-bo-en-sim verified by the PulseGate indexer
Other apps tracked under the same category.