This repository hosts GGUF quantized files for the QwQ-32B model. It provides multiple quantization levels so the 32-billion-parameter model can run efficiently on consumer hardware. The model includes support for tool calling and follows standard chat templates, making it suitable for local deployment of a capable reasoning and conversational AI.
In the Foundation models & chat space, QwQ 32B takes a focused approach. It focuses on running a high-parameter 32B language model locally with manageable resource usage. It is built as an open-source project for AI developers and enthusiasts. QwQ 32B is open source under the MIT license. QwQ 32B is available on the web, the command line, and API.
Maziyar Panahi builds and maintains QwQ 32B, and it first shipped in 2023. The project is developed in the open on GitHub with 122.5k stars and 1.2k commits in the last 90 days. Among its 4 catalogued features are Large Language Model, Quantized GGUF, and Local Inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do