Qdrant/all_miniLM_L6_v2_with_attentions is an ONNX export of the sentence-transformers/all-MiniLM-L6-v2 model that has been adjusted to return attention weights. It belongs to the class of sentence similarity and text embedding models hosted on Hugging Face. The model is intended for BM42 searches and is supposed to be used with Qdrant.
It supports loading via the Transformers library with AutoTokenizer and AutoModel, including an example that runs the model on automatic device mapping. An inference example with FastEmbed is also provided. The model card lists it under the tasks of sentence similarity, feature extraction, and text embeddings inference, with English as the language and the bert architecture.
The model carries an Apache 2.0 license. It is published by the organization Qdrant on the Hugging Face platform, where it can be deployed or copied to a bucket. No pricing information is stated because the model is openly available for download and local or hosted inference.
In the Embeddings & retrieval space, All miniLM L6 V2 With Attentions takes a focused approach. It focuses on providing attention-aware embeddings for hybrid search and retrieval in vector databases like Qdrant. All miniLM L6 V2 With Attentions is an open-source project aimed at developers. All miniLM L6 V2 With Attentions is open source under the Apache-2.0 license. It runs on the web and API.
Qdrant builds and maintains All miniLM L6 V2 With Attentions, and it first shipped in 2023. The project is developed in the open on GitHub with 3.1k stars and 4 commits in the last 90 days. Among its 3 catalogued features are Sentence Embeddings, Attention Weights, and BM42 Search.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do