This is an FP8 dynamically quantized version of Meta's Llama-3.3-70B-Instruct model. It maintains strong reasoning and instruction-following capabilities while significantly reducing memory requirements compared to the original 16-bit model. The model is suitable for local or self-hosted inference using compatible quantization and serving frameworks.
In the Text generation space, Llama 3.3 70B Instruct FP8 Dynamic takes a focused approach. It focuses on running a high-performance 70B instruction model with reduced memory footprint. Llama 3.3 70B Instruct FP8 Dynamic is an open-source project aimed at developers. The project is open source (MIT). It runs on the web and API.
cortecs builds and maintains Llama 3.3 70B Instruct FP8 Dynamic, and it first shipped in 2020. The project is developed in the open on GitHub with 13.3k stars and 53 commits in the last 90 days. Among its 3 catalogued features are FP8 quantization, instruction tuning, and dynamic quantization.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do