BLIP-2 is a vision-language model from Salesforce that leverages a frozen image encoder and a large language model (OPT-6.7B). It excels at visual question answering, image captioning, and other image-to-text tasks. The model is available through the Hugging Face Transformers library and can be used for research and application development.
Blip2 Opt 6.7b is a Foundation models & chat product. It focuses on connecting vision and language models to enable image understanding and visual question answering. Blip2 Opt 6.7b is an open-source project aimed at AI developers. The project is open source (BSD-3-Clause). The product ships for the web, the command line, and API.
Salesforce builds and maintains Blip2 Opt 6.7b, and the product first shipped in 2022. The project is developed in the open on GitHub with 11.3k stars and 1 commits in the last 90 days. Among its 3 catalogued features are visual question answering, image captioning, and image-to-text.
Latest indexed changes and source events
Salesforce/blip2-opt-6.7b verified by the PulseGate indexer
Other apps tracked under the same category.