unsloth/GLM-4.7-Flash-FP8-Dynamic is an FP8 dynamically quantized version of the GLM-4.7-Flash model. Created by Unsloth, it enables efficient local inference with significantly reduced VRAM requirements while maintaining model quality. The model includes a custom chat template and is compatible with the Transformers library and other inference engines.
In the Foundation models & chat space, GLM 4.7 Flash FP8 Dynamic takes a focused approach. It focuses on running large language models with reduced memory and faster inference using FP8 quantization. GLM 4.7 Flash FP8 Dynamic is an open-source project aimed at developers and researchers. The project is open source (Apache-2.0). GLM 4.7 Flash FP8 Dynamic is available on the web, the command line, and API.
Behind GLM 4.7 Flash FP8 Dynamic is Unsloth, and it first shipped in 2023. Development happens publicly on GitHub with 68.4k stars and 1.2k commits in the last 90 days.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do