This is an FP8 dynamically quantized version of Meta's Llama-3.3-70B-Instruct model. It maintains strong reasoning and instruction-following capabilities while significantly reducing memory requirements compared to the original 16-bit model. The model is suitable for local or self-hosted inference using compatible quantization and serving frameworks.
Llama 3.3 70B Instruct FP8 Dynamic is a Foundation models & chat product. It focuses on running a high-performance 70B instruction model with reduced memory footprint. It is built as an open-source project for developers. Llama 3.3 70B Instruct FP8 Dynamic is open source under the MIT license. The product ships for the web and API.
Behind Llama 3.3 70B Instruct FP8 Dynamic is cortecs, and the product first shipped in 2020. Development happens publicly on GitHub with 13.3k stars and 53 commits in the last 90 days. Key capabilities include FP8 quantization, instruction tuning, and dynamic quantization.
Latest indexed changes and source events
cortecs/Llama-3.3-70B-Instruct-FP8-Dynamic verified by the PulseGate indexer
Other apps tracked under the same category.