An FP8 dynamically quantized version of Meta's Llama 3.1 8B Instruct model created by RedHatAI. It maintains high performance while significantly reducing memory footprint for local or self-hosted inference. The model supports standard chat templates and is optimized for efficient deployment.
Meta Llama 3.1 8B Instruct FP8 Dynamic is a Text generation project. It focuses on running the Llama 3.1 8B Instruct model with reduced memory and compute requirements via quantization. It is built as an open-source project for developers. The project is open source (Apache-2.0). Meta Llama 3.1 8B Instruct FP8 Dynamic is available on the web and the command line, and it can be self-hosted.
It is developed by RedHatAI, and it first shipped in 2019. The project is developed in the open on GitHub with 3.6k stars and 167 commits in the last 90 days. Among its 3 catalogued features are Quantized Model, Instruction Tuned, and Dynamic FP8. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do