EAGLE3 is a speculative decoding method that extrapolates hidden states from LLMs to achieve significant speedups in token generation. This 8B model is based on Llama 3.1 Instruct and implements the EAGLE-3 technique. It has been shown to provide 2-3x faster inference compared to standard decoding while preserving output quality. The approach is trainable and compatible with various inference engines.
EAGLE3 LLaMA3.1 Instruct 8B is a Foundation models & chat project. It focuses on accelerating the generation speed of large language models while maintaining output distribution consistency. It is built as an open-source project for developers. The project is open source (Open Source). EAGLE3 LLaMA3.1 Instruct 8B is available on the web and API.
yuhuili builds and maintains EAGLE3 LLaMA3.1 Instruct 8B, and it first shipped in 2023. Development happens publicly on GitHub with 2.5k stars.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do