NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 is an FP8-quantized version of NVIDIA's Nemotron-3 language model. It includes optimized chat templates and is designed for efficient inference. Hosted on Hugging Face, it targets developers and researchers needing high-performance language models that can run with lower memory and compute requirements on compatible NVIDIA systems.
In the Foundation models & chat space, NVIDIA Nemotron 3 Nano 30B A3B takes a focused approach. It focuses on deploying efficient large language models with reduced precision for faster inference on NVIDIA hardware. It is built as an open-source project for developers. NVIDIA Nemotron 3 Nano 30B A3B is open source under the Apache-2.0 license. NVIDIA Nemotron 3 Nano 30B A3B is available on the command line and API.
Behind NVIDIA Nemotron 3 Nano 30B A3B is NVIDIA, based in the United States, and the product first shipped in 2026. Development happens publicly on GitHub with 314 stars and 72 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 14 similar tools. Key capabilities include Quantized Weights, FP8 Precision, and Chat Template. It exposes integrations via a public API.
Latest indexed changes and source events
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 verified by the PulseGate indexer
Other apps tracked under the same category.