Blip2 Opt 2.7b is a vision-language model hosted on Hugging Face under the Salesforce organization. It carries the pipeline tag image-text-to-text and belongs to the blip-2 model family. The model combines visual understanding with language generation to support tasks such as visual question answering and image-to-text conversion.
It is distributed as an open model repository that includes PyTorch and safetensors weights. The library_name field identifies it as compatible with the Transformers library, allowing integration into standard machine-learning workflows that rely on this framework. Model files are publicly accessible without gating.
The repository records more than 17 million all-time downloads and approximately 500 thousand recent downloads. It has accumulated 446 likes from the Hugging Face community. The model card was first created in February 2023 and received its most recent update in February 2025.
No pricing, licensing terms, or specific training details appear in the available metadata. The entry contains no further description of architecture components, supported input formats, or example use cases beyond the pipeline and tag information.
In the Multimodal & vision space, Blip2 Opt 2.7b takes a focused approach. It focuses on bridging vision and language understanding to enable image-based question answering and captioning. Blip2 Opt 2.7b is an open-source project aimed at developers. Blip2 Opt 2.7b is open source under the BSD-3-Clause license. It ships for the web, the command line, and API.
Salesforce builds and maintains Blip2 Opt 2.7b, and it first shipped in 2022. Development happens publicly on GitHub with 11.3k stars and 1 commit in the last 90 days. Among its 3 catalogued features are Visual Question Answering, Image Captioning, and vision-Language. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do