This is an NVIDIA-maintained quantized version of a Llama 3.3-based Nemotron Super model with 49B parameters in FP8 format. It includes a custom chat template and is designed for efficient text generation and conversational AI. The model is hosted on Hugging Face and can be used with the Transformers library.
Llama 3 3 Nemotron Super 49B V1 5 is a Foundation models & chat product. It focuses on running large-scale language model inference efficiently with reduced precision quantization. Llama 3 3 Nemotron Super 49B V1 5 is an open-source project aimed at AI developers and researchers. The project is open source (Apache-2.0). Llama 3 3 Nemotron Super 49B V1 5 is available on the web, the command line, and API.
NVIDIA builds and maintains Llama 3 3 Nemotron Super 49B V1 5, and the product first shipped in 2024. The project is developed in the open on GitHub with 1k stars and 66 commits in the last 90 days.
Latest indexed changes and source events
nvidia/Llama-3_3-Nemotron-Super-49B-v1_5-FP8 verified by the PulseGate indexer