Qwen3-32B-NVFP4 is an FP4 quantized version of the Qwen3-32B model created by NVIDIA. It includes support for tool calling and is optimized for high-performance inference on NVIDIA GPUs. The model is distributed on Hugging Face for developers seeking state-of-the-art performance with reduced memory and compute requirements.
Qwen3 32B is a Foundation models & chat project. It focuses on enabling efficient inference of large 32B parameter models on NVIDIA hardware using advanced low-precision formats. It is built as an open-source project for developers. Qwen3 32B is open source under the Open Source license. Qwen3 32B is available on the web and API.
It is developed by NVIDIA (United States), and it first shipped in 2025. PulseGate's similarity index places it among 18 comparable projects. Among its 4 catalogued features are FP4 quantization, tool calling, and high performance.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do