An FP8 dynamically quantized version of Meta's Llama 3.1 8B Instruct model created by RedHatAI. It maintains high performance while significantly reducing memory footprint for local or self-hosted inference. The model supports standard chat templates and is optimized for efficient deployment.
Meta Llama 3.1 8B Instruct FP8 Dynamic is a Foundation models & chat product. It focuses on running the Llama 3.1 8B Instruct model with reduced memory and compute requirements via quantization. It is built as an open-source project for developers. Meta Llama 3.1 8B Instruct FP8 Dynamic is open source under the Apache-2.0 license. The product ships for the web and the command line, and it can be self-hosted.
It is developed by RedHatAI, and the product first shipped in 2019. Development happens publicly on GitHub with 3.6k stars and 167 commits in the last 90 days. Key capabilities include Quantized Model, Instruction Tuned, and Dynamic FP8. It exposes integrations via a public API.
Latest indexed changes and source events
RedHatAI/Meta-Llama-3.1-8B-Instruct-FP8-dynamic verified by the PulseGate indexer
Other apps tracked under the same category.