nvidia/Qwen3-14B-FP8 is a quantized version of the Qwen3 14 billion parameter model using FP8 precision. It includes support for tool calling and advanced chat templates. The model is hosted on Hugging Face and is designed for efficient inference on NVIDIA hardware while maintaining high performance.
Qwen3 14B sits in PulseGate's Foundation models & chat category. It focuses on running large language models efficiently on NVIDIA GPUs with reduced memory footprint using FP8 precision. It is built as an open-source project for developers and AI practitioners. The project is open source (Apache-2.0). It runs on the web and API.
Behind Qwen3 14B is NVIDIA, and it first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 342 commits in the last 90 days. It competes in a saturated segment with 25 similar apps in PulseGate's index. Key capabilities include quantized weights, tool calling support, and chat template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do