GLM-4V-9B is a multimodal variant of the GLM-4 large language model series. It combines a vision encoder with a 9-billion-parameter language model, enabling it to process images alongside text. The model supports both Chinese and English and is suitable for vision-language tasks such as image captioning, visual question answering, and document understanding.
In the Multimodal & vision space, Glm 4v 9b takes a focused approach. It focuses on providing an open multimodal model capable of understanding both images and text, with strong Chinese language performance. Glm 4v 9b is an open-source project aimed at AI developers. Glm 4v 9b is open source under the Apache-2.0 license. It ships for the web and API.
Behind Glm 4v 9b is ZAI-ORG, and it first shipped in 2024. Development happens publicly on GitHub with 7.1k stars. Key capabilities include Vision Language Model, multimodal, and Chinese English Support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do