GLM-4.7-Flash-NVFP4 is a quantized variant of the GLM-4.7 language model hosted on Hugging Face. It is provided by user GadflyII under the repository name GLM-4.7-Flash-NVFP4 and uses NVFP4 precision for reduced memory and compute requirements during inference.
The model includes a chat template that supports tool calling. When tools are supplied, the template instructs the model to output function calls in a specific XML format enclosed in tool_call tags, with individual arguments expressed as arg_key and arg_value pairs. A visible_text macro handles content rendering for strings, iterables, or mappings containing text items. The tokenizer configuration specifies an end-of-text token as the pad token.
It is distributed as a model repository on the Hugging Face platform. Users obtain the files through the standard Hugging Face ecosystem for loading with compatible inference libraries.
The page provides no information on licensing, pricing, or intended audience beyond the general Hugging Face context of open-source and open-science AI advancement.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on deploying large language models efficiently with reduced memory footprint and faster inference speeds. GLM 4.7 Flash is an open-source project aimed at developers. GLM 4.7 Flash is open source under the Apache-2.0 license. It ships for the web, the command line, and API.
GadflyII builds and maintains GLM 4.7 Flash, and it first shipped in 2023. The project is developed in the open on GitHub with 2.3k commits in the last 90 days. PulseGate's similarity index places it among 7 comparable projects. Key capabilities include text generation, tool calling, and quantized inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do