NVIDIA-Nemotron-3-Nano-4B-FP8 is a 4 billion parameter language model from NVIDIA, provided in FP8 precision for optimized performance on NVIDIA hardware. It supports text generation and chat applications using a custom chat template. The model is designed for developers building efficient AI applications that leverage NVIDIA GPUs for local or on-premise inference.
NVIDIA Nemotron 3 Nano 4B is a Foundation models & chat product. It focuses on deploying efficient small language models on NVIDIA GPUs with reduced precision for faster inference. It is built as an open-source project for developers. NVIDIA Nemotron 3 Nano 4B is open source under the Apache-2.0 license. The product ships for the command line and API.
NVIDIA builds and maintains NVIDIA Nemotron 3 Nano 4B, and the product first shipped in 2024. Development happens publicly on GitHub with 1k stars and 67 commits in the last 90 days. Key capabilities include 4B parameters, FP8 precision, and optimized for NVIDIA.
Latest indexed changes and source events
nvidia/NVIDIA-Nemotron-3-Nano-4B-FP8 verified by the PulseGate indexer
Other apps tracked under the same category.