Llama-3.1-8B-Instruct-FP8 is an FP8 quantized version of Meta's Llama 3.1 8B instruct model, optimized by NVIDIA for high-performance inference. It maintains strong instruction-following capabilities while benefiting from reduced memory bandwidth and faster execution on compatible NVIDIA GPUs. The model is hosted on Hugging Face and is intended for developers building efficient LLM-powered applications.
In the Foundation models & chat space, Llama 3.1 8B Instruct takes a focused approach. It focuses on running Llama 3.1 efficiently using reduced-precision FP8 format on NVIDIA hardware. Llama 3.1 8B Instruct is an open-source project aimed at developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
NVIDIA builds and maintains Llama 3.1 8B Instruct, and it first shipped in 2024. Development happens publicly on GitHub with 3.3k stars and 355 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 8 similar projects. Key capabilities include FP8 precision, optimized inference, and instruction tuning.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do