BAAI/bge-large-zh-v1.5 is a model hosted on Hugging Face for generating embeddings from Chinese text. Developed by the Beijing Academy of Artificial Intelligence, it belongs to the class of foundation models focused on feature extraction and sentence similarity.
The model is provided as a PyTorch implementation based on the BERT architecture and is tagged for Chinese language processing, text embeddings, and sentence-transformers compatibility. It can produce vector representations of sentences that support similarity calculations. Example code demonstrates loading the model to encode lists of sentences and compute a similarity matrix between them.
Integration occurs through standard libraries. The sentence-transformers package allows direct instantiation and encoding, while the Transformers library supports it via a feature-extraction pipeline. These options enable use in local applications, notebooks, or inference providers. The model card lists five associated arXiv papers and indicates an MIT license.
It is followed by over 4,000 users on the platform and has received 643 likes. The repository includes model files, versions, and community contributions.
In the Embeddings & retrieval space, Bge Large Zh takes a focused approach. It focuses on providing high-quality Chinese text embeddings for semantic similarity, search, and NLP applications. Bge Large Zh is an open-source project aimed at NLP researchers and developers. Bge Large Zh is open source under the Open Source license. Bge Large Zh is available on the web, the command line, and API, and it can be self-hosted.
Behind Bge Large Zh is Beijing Academy of Artificial Intelligence, based in China, and it first shipped in 2023. Among its 5 catalogued features are text embeddings, sentence similarity, and feature extraction.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do