This is a community-quantized 4-bit AWQ version of NVIDIA's Nemotron-3 Super 120B parameter language model. It enables efficient local or self-hosted inference of a powerful open-weights LLM using significantly less GPU memory than the original. The repository provides model weights, tokenizer configuration, and chat templates for integration with popular inference frameworks like Transformers or vLLM.
NVIDIA Nemotron 3 Super 120B A12B sits in PulseGate's Foundation models & chat category. It focuses on running large 120B parameter language models locally with reduced memory requirements. It is built as an open-source project for developers. NVIDIA Nemotron 3 Super 120B A12B is open source under the Open Source license. NVIDIA Nemotron 3 Super 120B A12B is available on the command line and API.
Behind NVIDIA Nemotron 3 Super 120B A12B is cyankiwi, and the product first shipped in 2024. Key capabilities include quantized weights, chat template, and JSON chat format.
Latest indexed changes and source events
cyankiwi/NVIDIA-Nemotron-3-Super-120B-A12B-AWQ-4bit verified by the PulseGate indexer
Other apps tracked under the same category.