GLM-4.7-Flash-NVFP4 is a quantized variant of the GLM-4.7 language model hosted on Hugging Face. It is provided by user GadflyII under the repository name GLM-4.7-Flash-NVFP4 and uses NVFP4 precision for reduced memory and compute requirements during inference.
The model includes a chat template that supports tool calling. When tools are supplied, the template instructs the model to output function calls in a specific XML format enclosed in tool_call tags, with individual arguments expressed as arg_key and arg_value pairs. A visible_text macro handles content rendering for strings, iterables, or mappings containing text items. The tokenizer configuration specifies an end-of-text token as the pad token.
It is distributed as a model repository on the Hugging Face platform. Users obtain the files through the standard Hugging Face ecosystem for loading with compatible inference libraries.
The page provides no information on licensing, pricing, or intended audience beyond the general Hugging Face context of open-source and open-science AI advancement.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on deploying large language models efficiently with reduced memory footprint and faster inference speeds. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
Behind GLM 4.7 Flash is GadflyII, and the product first shipped in 2023. The project is developed in the open on GitHub with 2.3k commits in the last 90 days. PulseGate's similarity index places it among 7 comparable tools. Among its 4 catalogued features are text generation, tool calling, and quantized inference.
Latest indexed changes and source events
GadflyII/GLM-4.7-Flash-NVFP4 verified by the PulseGate indexer
Other apps tracked under the same category.