NVIDIA-Nemotron-3-Nano-4B-FP8 is a 4 billion parameter language model from NVIDIA, provided in FP8 precision for optimized performance on NVIDIA hardware. It supports text generation and chat applications using a custom chat template. The model is designed for developers building efficient AI applications that leverage NVIDIA GPUs for local or on-premise inference.
NVIDIA Nemotron 3 Nano 4B is a Foundation models & chat project. It focuses on deploying efficient small language models on NVIDIA GPUs with reduced precision for faster inference. NVIDIA Nemotron 3 Nano 4B is an open-source project aimed at developers. NVIDIA Nemotron 3 Nano 4B is open source under the Apache-2.0 license. It ships for the command line and API.
Behind NVIDIA Nemotron 3 Nano 4B is NVIDIA, based in the United States, and it first shipped in 2024. The project is developed in the open on GitHub with 1k stars and 67 commits in the last 90 days. Among its 4 catalogued features are 4B parameters, FP8 precision, and optimized for NVIDIA.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do