This is an FP8-quantized version of Meta's Llama 3.1 70B Instruct model, published by RedHatAI. It maintains strong instruction-following capabilities while using a more memory-efficient format suitable for self-hosted deployment. The model includes a detailed chat template and is designed for production inference environments.
In the Text generation space, Meta Llama 3.1 70B Instruct takes a focused approach. It focuses on running a high-performance 70B parameter language model with reduced memory requirements through quantization. Meta Llama 3.1 70B Instruct is an open-source project aimed at developers and enterprises deploying LLMs. The project is open source (Apache-2.0). It ships for the web, the command line, and API, and it can be self-hosted.
It is developed by RedHatAI, and it first shipped in 2019. Development happens publicly on GitHub with 3.6k stars and 165 commits in the last 90 days. PulseGate's similarity index places it among 17 comparable projects. Among its 3 catalogued features are instruction tuning, Quantized FP8, and chat template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do