This is a 6-bit quantized version of the GLM-4.7-Flash model, prepared by the LM Studio community for efficient inference using the MLX framework on Apple devices. It includes support for tool calling and follows a standard chat template. The model is hosted on Hugging Face for easy local deployment.
In the Foundation models & chat space, GLM 4.7 Flash takes a focused approach. It focuses on running a capable GLM-4 model efficiently on local Apple hardware through quantization and MLX optimization. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (MIT). It runs on the web, the command line, and API.
lmstudio-community builds and maintains GLM 4.7 Flash, and it first shipped in 2023. Development happens publicly on GitHub with 6.4k stars and 21 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 5 similar projects. Among its 4 catalogued features are Large Language Model, Tool Calling, and Quantized Inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do