GLM-4.6V-Flash-MLX-6bit is a quantized multimodal model hosted on Hugging Face by the lmstudio-community. It belongs to the class of foundation models and is distributed as an MLX-optimized 6-bit version intended for local inference on compatible hardware.
The model includes a chat template that supports tool calling. When tools are provided, the template instructs the model to output function calls in a specific XML format enclosed in tool_call tags, with individual arguments expressed through arg_key and arg_value elements. The template also defines handling for multimodal content that mixes text and images, using special tokens such as begin_of_image and image.
It is delivered as a repository on the Hugging Face platform, allowing download of the model files for use with the MLX framework. The page references a pad_token set to the end-of-text token and supplies the full chat template as part of the configuration.
No pricing, licensing terms, or explicit target audience beyond the general Hugging Face modeling community are stated.
In the Foundation models & chat space, GLM 4.6V Flash takes a focused approach. It focuses on running advanced multimodal AI models locally with reduced memory requirements. It is built as an open-source project for developers. GLM 4.6V Flash is open source under the MIT license. It ships for the web, the command line, and API.
Behind GLM 4.6V Flash is lmstudio-community, and it first shipped in 2024. Development happens publicly on GitHub with 5.2k stars and 419 commits in the last 90 days. Among its 4 catalogued features are Multimodal Input, Tool Calling, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do