nlpconnect/vit-gpt2-image-captioning is an open-source model that combines a Vision Transformer (ViT) image encoder with a GPT-2 text decoder to generate natural language captions from input images. It is designed for image-to-text tasks and is available through the Hugging Face Transformers library. Researchers and developers use it for building applications that require automatic image description.
Vit Gpt2 Image Captioning sits in PulseGate's Foundation models & chat category. Automatically generating descriptive captions for images using open-source models. It is built as an open-source project for developers. Vit Gpt2 Image Captioning is open source under the Apache-2.0 license. It runs on the web and API.
nlpconnect builds and maintains Vit Gpt2 Image Captioning, and the product first shipped in 2018. Development happens publicly on GitHub with 162.8k stars and 745 commits in the last 90 days. Key capabilities include Image Captioning, Vision Transformer, and GPT-2 Decoder.
Latest indexed changes and source events
nlpconnect/vit-gpt2-image-captioning verified by the PulseGate indexer
Other apps tracked under the same category.