BLIP-2 is a vision-language model from Salesforce that leverages a frozen image encoder and a large language model (OPT-6.7B). It excels at visual question answering, image captioning, and other image-to-text tasks. The model is available through the Hugging Face Transformers library and can be used for research and application development.
Blip2 Opt 6.7b sits in PulseGate's Multimodal & vision category. It focuses on connecting vision and language models to enable image understanding and visual question answering. It is built as an open-source project for AI developers. The project is open source (BSD-3-Clause). It runs on the web, the command line, and API.
Salesforce builds and maintains Blip2 Opt 6.7b, and it first shipped in 2022. Development happens publicly on GitHub with 11.3k stars and 1 commit in the last 90 days. Among its 3 catalogued features are visual question answering, image captioning, and image-to-text.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do