This repository contains GGUF quantized weights for the GLM-4.7-Flash model. It is compatible with llama.cpp and other GGUF-compatible runtimes, allowing users to run the model locally with support for tool calling and standard chat templates.
In the Foundation models & chat space, GLM 4.7 Flash takes a focused approach. It focuses on enabling efficient CPU and GPU inference of the GLM-4.7 model on consumer hardware using the GGUF format. It is built as an open-source project for developers. GLM 4.7 Flash is open source under the Open Source license. GLM 4.7 Flash is available on the command line.
Behind GLM 4.7 Flash is ggml-org. It operates in a well-populated space: PulseGate tracks 7 similar tools.
Latest indexed changes and source events
ggml-org/GLM-4.7-Flash-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.