NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 is an FP8-quantized version of NVIDIA's Nemotron-3 language model. It includes optimized chat templates and is designed for efficient inference. Hosted on Hugging Face, it targets developers and researchers needing high-performance language models that can run with lower memory and compute requirements on compatible NVIDIA systems.
In the Foundation models & chat space, NVIDIA Nemotron 3 Nano 30B A3B takes a focused approach. It focuses on deploying efficient large language models with reduced precision for faster inference on NVIDIA hardware. NVIDIA Nemotron 3 Nano 30B A3B is an open-source project aimed at developers. NVIDIA Nemotron 3 Nano 30B A3B is open source under the Apache-2.0 license. NVIDIA Nemotron 3 Nano 30B A3B is available on the command line and API.
NVIDIA builds and maintains NVIDIA Nemotron 3 Nano 30B A3B, and it first shipped in 2026. The project is developed in the open on GitHub with 314 stars and 72 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 14 similar projects. Among its 3 catalogued features are Quantized Weights, FP8 Precision, and Chat Template. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do