This is an FP8-quantized version of Meta's Llama 3.1 70B Instruct model, published by RedHatAI. It maintains strong instruction-following capabilities while using a more memory-efficient format suitable for self-hosted deployment. The model includes a detailed chat template and is designed for production inference environments.
In the Foundation models & chat space, Meta Llama 3.1 70B Instruct takes a focused approach. It focuses on running a high-performance 70B parameter language model with reduced memory requirements through quantization. It is built as an open-source project for developers and enterprises deploying LLMs. Meta Llama 3.1 70B Instruct is open source under the Apache-2.0 license. It runs on the web, the command line, and API, and it can be self-hosted.
It is developed by RedHatAI, and the product first shipped in 2019. Development happens publicly on GitHub with 3.6k stars and 165 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 17 similar tools. Key capabilities include instruction tuning, Quantized FP8, and chat template.
Latest indexed changes and source events
RedHatAI/Meta-Llama-3.1-70B-Instruct-FP8 verified by the PulseGate indexer
Other apps tracked under the same category.