GLM-4.7-Flash-AWQ is an AWQ-quantized version of the GLM-4.7-Flash model optimized for fast and memory-efficient inference. It supports advanced features such as function calling and custom chat templates. The model is distributed on Hugging Face and can be run locally using Transformers or Docker.
In the Foundation models & chat space, GLM 4.7 Flash takes a focused approach. It focuses on deploying large language models with significantly reduced memory footprint while preserving tool-use and conversational capabilities. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (Apache-2.0). GLM 4.7 Flash is available on the web, the command line, and API.
QuantTrio builds and maintains GLM 4.7 Flash, and it first shipped in 2025. The project is developed in the open on GitHub with 4.4k stars. Key capabilities include AWQ Quantization, Function Calling, and Chat Templates.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do