This repository contains GGUF quantized weights for the GLM-4.7-Flash model. It is compatible with llama.cpp and other GGUF-compatible runtimes, allowing users to run the model locally with support for tool calling and standard chat templates.
GLM 4.7 Flash sits in PulseGate's Quantised & converted weights category. It focuses on enabling efficient CPU and GPU inference of the GLM-4.7 model on consumer hardware using the GGUF format. It is built as an open-source project for developers. The project is open source (Open Source). GLM 4.7 Flash is available on the command line.
ggml-org builds and maintains GLM 4.7 Flash. It operates in a well-populated space: PulseGate tracks 7 similar projects.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do