nvidia/Llama-3.3-70B-Instruct-FP8 provides an FP8-quantized version of Meta's Llama 3.3 70B Instruct model. It includes optimized inference support and a chat template with built-in tool calling capabilities. The model is designed for efficient local or self-hosted deployment using NVIDIA hardware and inference stacks.
In the Foundation models & chat space, Llama 3.3 70B Instruct takes a focused approach. It focuses on running a high-performance 70B parameter language model efficiently on GPU hardware with reduced memory requirements. Llama 3.3 70B Instruct is an open-source project aimed at developers. Llama 3.3 70B Instruct is open source under the Apache-2.0 license. Llama 3.3 70B Instruct is available on the web and the command line, and it can be self-hosted.
It is developed by NVIDIA (United States), and it first shipped in 2024. Development happens publicly on GitHub with 3.3k stars and 354 commits in the last 90 days. The category is crowded — PulseGate's index counts 21 comparable apps. Key capabilities include Instruction Tuning, Quantized Inference, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do