GLM-4.6V-Flash-MLX-4bit is a 4-bit quantized version of the GLM-4.6V-Flash multimodal model hosted on Hugging Face by the lmstudio-community. It is distributed as part of the collection of models available for local AI runtimes.
The model includes a chat template that supports tools through XML-formatted function calls. It handles multimodal inputs by processing both text strings and image content marked with specific tokens such as <|begin_of_image| and <|image|. The configuration specifies a pad token of <|endoftext|.
It is delivered as a repository on the Hugging Face platform, where models can be downloaded for integration into compatible local inference frameworks. The page appears under the lmstudio-community organization, indicating suitability for users running models through LM Studio or similar environments.
No pricing or licensing details are stated on the model page. The surrounding Hugging Face site promotes open source and open science but does not attribute a specific license to this quantized variant.
GLM 4.6V Flash is a Foundation models & chat product. It focuses on enabling efficient on-device or local-server inference of large vision-language models. It is built as an open-source project for developers. GLM 4.6V Flash is open source under the MIT license. GLM 4.6V Flash is available on the web, the command line, and API.
It is developed by lmstudio-community, and the product first shipped in 2024. Development happens publicly on GitHub with 5.2k stars and 419 commits in the last 90 days. Key capabilities include Multimodal Inference, Quantized Model, and Vision Support.
Latest indexed changes and source events
lmstudio-community/GLM-4.6V-Flash-MLX-4bit verified by the PulseGate indexer
Other apps tracked under the same category.