This is a community-quantized 4-bit AWQ version of NVIDIA's Nemotron-3 Super 120B parameter language model. It enables efficient local or self-hosted inference of a powerful open-weights LLM using significantly less GPU memory than the original. The repository provides model weights, tokenizer configuration, and chat templates for integration with popular inference frameworks like Transformers or vLLM.
NVIDIA Nemotron 3 Super 120B A12B is a Foundation models & chat project. It focuses on running large 120B parameter language models locally with reduced memory requirements. NVIDIA Nemotron 3 Super 120B A12B is an open-source project aimed at developers. The project is open source (Open Source). NVIDIA Nemotron 3 Super 120B A12B is available on the command line and API.
Behind NVIDIA Nemotron 3 Super 120B A12B is cyankiwi, and it first shipped in 2024. Among its 3 catalogued features are quantized weights, chat template, and JSON chat format.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do