GLM-4.6V-Flash-MLX-6bit is a quantized multimodal model hosted on Hugging Face by the lmstudio-community. It belongs to the class of foundation models and is distributed as an MLX-optimized 6-bit version intended for local inference on compatible hardware.
The model includes a chat template that supports tool calling. When tools are provided, the template instructs the model to output function calls in a specific XML format enclosed in tool_call tags, with individual arguments expressed through arg_key and arg_value elements. The template also defines handling for multimodal content that mixes text and images, using special tokens such as begin_of_image and image.
It is delivered as a repository on the Hugging Face platform, allowing download of the model files for use with the MLX framework. The page references a pad_token set to the end-of-text token and supplies the full chat template as part of the configuration.
No pricing, licensing terms, or explicit target audience beyond the general Hugging Face modeling community are stated.
GLM 4.6V Flash sits in PulseGate's Foundation models & chat category. It focuses on running advanced multimodal AI models locally with reduced memory requirements. GLM 4.6V Flash is an open-source project aimed at developers. The project is open source (MIT). The product ships for the web, the command line, and API.
lmstudio-community builds and maintains GLM 4.6V Flash, and the product first shipped in 2024. The project is developed in the open on GitHub with 5.2k stars and 419 commits in the last 90 days. Among its 4 catalogued features are Multimodal Input, Tool Calling, and Chat Template.
Latest indexed changes and source events
lmstudio-community/GLM-4.6V-Flash-MLX-6bit verified by the PulseGate indexer
Other apps tracked under the same category.