Llama-3.1-8B-Instruct-FP8 is an FP8 quantized version of Meta's Llama 3.1 8B instruct model, optimized by NVIDIA for high-performance inference. It maintains strong instruction-following capabilities while benefiting from reduced memory bandwidth and faster execution on compatible NVIDIA GPUs. The model is hosted on Hugging Face and is intended for developers building efficient LLM-powered applications.
Llama 3.1 8B Instruct sits in PulseGate's Foundation models & chat category. It focuses on running Llama 3.1 efficiently using reduced-precision FP8 format on NVIDIA hardware. Llama 3.1 8B Instruct is an open-source project aimed at developers. The project is open source (Apache-2.0). The product ships for the web, the command line, and API.
It is developed by NVIDIA (United States), and the product first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 355 commits in the last 90 days. PulseGate's similarity index places it among 8 comparable tools. Among its 3 catalogued features are FP8 precision, optimized inference, and instruction tuning.
Latest indexed changes and source events
nvidia/Llama-3.1-8B-Instruct-FP8 verified by the PulseGate indexer
Other apps tracked under the same category.