This repository provides GGUF quantized weights of the GLM-4.7-Flash model, optimized for use with LM Studio, llama.cpp, and other GGUF-compatible engines. It includes support for tool calling and is intended for local CPU/GPU inference of the GLM series of large language models.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on running the GLM-4.7-Flash model efficiently on consumer hardware using popular local LLM tools. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (MIT). It ships for the web, the command line, and API.
Behind GLM 4.7 Flash is lmstudio-community, and it first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 8 similar projects. Key capabilities include GGUF Format, quantized, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do