This repository contains an NVIDIA-optimized FP4 (NVFP4) quantized version of the Qwen3 14B model. It includes specialized chat templates and tool-calling support optimized for NVIDIA inference stacks. The quantization enables faster and more memory-efficient inference while preserving model capabilities.
Qwen3 14B sits in PulseGate's Foundation models & chat category. Efficiently running the Qwen3 14B model on NVIDIA hardware using advanced 4-bit quantization. It is built as an open-source project for AI developers optimizing for NVIDIA GPUs. Qwen3 14B is open source under the Apache-2.0 license. It ships for the web, API, and the command line.
Behind Qwen3 14B is NVIDIA, and it first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 355 commits in the last 90 days. PulseGate's similarity index places it among 17 comparable projects. Key capabilities include quantized inference, tool calling, and chat template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do