Qwen3-32B-FP8 is an FP8-quantized version of Alibaba's Qwen3 32B large language model. It supports advanced capabilities including tool calling, long-context understanding, and multilingual performance while significantly reducing VRAM usage compared to the original. The model is designed for local and self-hosted inference by developers building AI applications.
In the Foundation models & chat space, Qwen3 32B takes a focused approach. It focuses on running a high-performance 32-billion-parameter language model with reduced memory requirements via FP8 quantization. Qwen3 32B is an open-source project aimed at developers. Qwen3 32B is open source under the Open Source license. It runs on the web, the command line, and API, and it can be self-hosted.
Behind Qwen3 32B is Qwen, and it first shipped in 2024. The project is developed in the open on GitHub with 27.4k stars. It competes in a saturated segment with 23 similar projects in PulseGate's index. Among its 3 catalogued features are Large Language Model, FP8 Quantization, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do