This repository provides GGUF quantized weights of the GLM-4.7-Flash model, optimized for use with LM Studio, llama.cpp, and other GGUF-compatible engines. It includes support for tool calling and is intended for local CPU/GPU inference of the GLM series of large language models.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on running the GLM-4.7-Flash model efficiently on consumer hardware using popular local LLM tools. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (MIT). GLM 4.7 Flash is available on the web, the command line, and API.
It is developed by lmstudio-community, and the product first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. PulseGate's similarity index places it among 8 comparable tools. Among its 3 catalogued features are GGUF Format, quantized, and Tool Calling.
Latest indexed changes and source events
lmstudio-community/GLM-4.7-Flash-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.