This repository contains the FP8 quantized version of Meta's Llama-4-Maverick-17B-128E-Instruct model. It is an open-weights large language model optimized for instruction following, tool calling, and long-context tasks. The model can be used with the Hugging Face Transformers library, local inference engines, or cloud providers.
Llama 4 Maverick 17B 128E Instruct is a Text generation project. It focuses on deploying a capable, open-weight instruction-tuned language model with efficient memory usage for local or cloud inference. It is built as an open-source project for developers. The project is open source (Open Source). It runs on the web, the command line, and API, and it can be self-hosted.
It is developed by Meta (United States), and it first shipped in 2024. The project is developed in the open on GitHub with 7.7k stars. Key capabilities include Instruction Following, Tool Use, and Long Context. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do