NVIDIA's Llama-3.3-70B-Instruct-NVFP4 provides a precision-optimized (NVFP4) quantized checkpoint of Meta's Llama 3.3 70B Instruct model. It is designed for high-performance inference on NVIDIA GPUs while preserving instruction-following capabilities. The model uses a custom chat template and supports tool use.
Llama 3.3 70B Instruct sits in PulseGate's Foundation models & chat category. It focuses on running the large Llama 3.3 70B model efficiently on NVIDIA hardware with reduced memory footprint. Llama 3.3 70B Instruct is an open-source project aimed at developers. The project is open source (Apache-2.0). Llama 3.3 70B Instruct is available on the web and API.
Behind Llama 3.3 70B Instruct is NVIDIA, and it first shipped in 2024. Development happens publicly on GitHub with 3.3k stars and 347 commits in the last 90 days. The category is crowded — PulseGate's index counts 20 comparable apps. Key capabilities include 4-bit quantization, instruct tuned, and NVIDIA optimized. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do