Qwen3-32B-FP8 is an FP8-quantized version of Alibaba's Qwen3 32B large language model. It supports advanced capabilities including tool calling, long-context understanding, and multilingual performance while significantly reducing VRAM usage compared to the original. The model is designed for local and self-hosted inference by developers building AI applications.
In the Foundation models & chat space, Qwen3 32B takes a focused approach. It focuses on running a high-performance 32-billion-parameter language model with reduced memory requirements via FP8 quantization. It is built as an open-source project for developers. Qwen3 32B is open source under the Open Source license. The product ships for the web, the command line, and API, and it can be self-hosted.
Behind Qwen3 32B is Qwen, and the product first shipped in 2024. Development happens publicly on GitHub with 27.4k stars. The category is crowded — PulseGate's index counts 23 comparable apps. Key capabilities include Large Language Model, FP8 Quantization, and Tool Calling.
Latest indexed changes and source events
Qwen/Qwen3-32B-FP8 verified by the PulseGate indexer
Other apps tracked under the same category.