This repository contains an NVIDIA-optimized FP4 (NVFP4) quantized version of the Qwen3 14B model. It includes specialized chat templates and tool-calling support optimized for NVIDIA inference stacks. The quantization enables faster and more memory-efficient inference while preserving model capabilities.
Qwen3 14B sits in PulseGate's Foundation models & chat category. Efficiently running the Qwen3 14B model on NVIDIA hardware using advanced 4-bit quantization. Qwen3 14B is an open-source project aimed at AI developers optimizing for NVIDIA GPUs. The project is open source (Apache-2.0). Qwen3 14B is available on the web, API, and the command line.
It is developed by NVIDIA, and the product first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 355 commits in the last 90 days. PulseGate's similarity index places it among 17 comparable tools. Among its 3 catalogued features are quantized inference, tool calling, and chat template.
Latest indexed changes and source events
nvidia/Qwen3-14B-NVFP4 verified by the PulseGate indexer
Other apps tracked under the same category.