This repository hosts an AWQ 4-bit quantized version of the GLM-4.7-Flash model, optimized for reduced memory usage and faster inference. It retains tool-calling capabilities and a full chat template, making it suitable for local or hosted deployment where full-precision models would be too large. The model is provided for developers seeking efficient open-weight alternatives.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on running the GLM-4.7-Flash model efficiently on consumer hardware through 4-bit quantization. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
cyankiwi builds and maintains GLM 4.7 Flash, and it first shipped in 2025. The project is developed in the open on GitHub with 4.4k stars. Among its 3 catalogued features are quantized weights, tool calling, and chat template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do