SmolVLM is a Hugging Face web application for asking questions about one or more uploaded images. It combines visual inputs with text queries to generate written responses and is based on an openly available vision-language model.
SmolVLM sits in PulseGate's Multimodal & vision category. It focuses on understanding and querying image content without manually inspecting or describing it. It is built as an open-source project for developers, researchers, and users exploring multimodal AI. SmolVLM is free to use. It ships for the web, and it can be self-hosted.
HuggingFaceTB builds and maintains SmolVLM. Key capabilities include Image Upload, Image Question Answering, and Multi-image Input.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do