Keye-VL-1_5-8B is an 8 billion parameter vision-language foundation model hosted on Hugging Face. It processes images and videos together with text through a specialized chat template that inserts vision tokens for multimodal input.
The model employs a custom chat template written in Jinja2 that tracks separate counts for images and videos. When a message contains an image or image_url it inserts vision_start, image_pad, and vision_end tokens, optionally prefixing with "Picture N:". Video content receives analogous video_pad tokens and an optional "Video N:" label. Plain text segments are inserted directly. The template also handles system prompts, user and assistant roles, and generation prompts using im_start and im_end delimiters.
It is distributed as open weights on the Hugging Face model repository under the repository name Kwai-Keye/Keye-VL-1_5-8B. The page indicates compatibility with standard Transformers pipelines that consume the provided tokenizer_config and chat_template. No pricing, licensing terms, or specific end-user audience are stated beyond the general Hugging Face open-science context.
Keye VL 1 5 8B sits in PulseGate's Foundation models & chat category. It focuses on understanding and reasoning over both images and videos using a single unified language model. It is built as an open-source project for Multimodal AI developers and researchers. Keye VL 1 5 8B is open source under the Open Source license. The product ships for the web, the command line, and API.
It is developed by Kwai (China), and the product first shipped in 2025. Development happens publicly on GitHub with 806 stars and 17 commits in the last 90 days. Key capabilities include Vision Language, Image Understanding, and Video Understanding.
Latest indexed changes and source events
Kwai-Keye/Keye-VL-1_5-8B verified by the PulseGate indexer
Other apps tracked under the same category.