Qdrant/all_miniLM_L6_v2_with_attentions is an ONNX export of the sentence-transformers/all-MiniLM-L6-v2 model that has been adjusted to return attention weights. It belongs to the class of sentence similarity and text embedding models hosted on Hugging Face. The model is intended for BM42 searches and is supposed to be used with Qdrant.
It supports loading via the Transformers library with AutoTokenizer and AutoModel, including an example that runs the model on automatic device mapping. An inference example with FastEmbed is also provided. The model card lists it under the tasks of sentence similarity, feature extraction, and text embeddings inference, with English as the language and the bert architecture.
The model carries an Apache 2.0 license. It is published by the organization Qdrant on the Hugging Face platform, where it can be deployed or copied to a bucket. No pricing information is stated because the model is openly available for download and local or hosted inference.
All miniLM L6 V2 With Attentions is a Foundation models & chat product. It focuses on providing attention-aware embeddings for hybrid search and retrieval in vector databases like Qdrant. All miniLM L6 V2 With Attentions is an open-source project aimed at developers. The project is open source (Apache-2.0). All miniLM L6 V2 With Attentions is available on the web and API.
Qdrant builds and maintains All miniLM L6 V2 With Attentions, and the product first shipped in 2023. The project is developed in the open on GitHub with 3.1k stars and 4 commits in the last 90 days. Among its 3 catalogued features are Sentence Embeddings, Attention Weights, and BM42 Search.
Latest indexed changes and source events
Qdrant/all_miniLM_L6_v2_with_attentions verified by the PulseGate indexer
Other apps tracked under the same category.