GLM-4.6V-Flash-MLX-4bit is a 4-bit quantized version of the GLM-4.6V-Flash multimodal model hosted on Hugging Face by the lmstudio-community. It is distributed as part of the collection of models available for local AI runtimes.
The model includes a chat template that supports tools through XML-formatted function calls. It handles multimodal inputs by processing both text strings and image content marked with specific tokens such as <|begin_of_image| and <|image|. The configuration specifies a pad token of <|endoftext|.
It is delivered as a repository on the Hugging Face platform, where models can be downloaded for integration into compatible local inference frameworks. The page appears under the lmstudio-community organization, indicating suitability for users running models through LM Studio or similar environments.
No pricing or licensing details are stated on the model page. The surrounding Hugging Face site promotes open source and open science but does not attribute a specific license to this quantized variant.
GLM 4.6V Flash is a Multimodal & vision project. It focuses on enabling efficient on-device or local-server inference of large vision-language models. GLM 4.6V Flash is an open-source project aimed at developers. The project is open source (MIT). It ships for the web, the command line, and API.
Behind GLM 4.6V Flash is lmstudio-community, and it first shipped in 2024. Development happens publicly on GitHub with 5.2k stars and 419 commits in the last 90 days. Among its 3 catalogued features are Multimodal Inference, Quantized Model, and Vision Support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do