This repository hosts an AWQ (Activation-aware Weight Quantization) INT4 quantized variant of Google's Gemma-2-9B-IT model. It enables efficient local inference of a capable 9-billion-parameter instruction-tuned LLM on hardware with limited VRAM. The model is distributed via Hugging Face and is compatible with transformers and vLLM inference backends.
In the Foundation models & chat space, Gemma 2 9b It takes a focused approach. It focuses on running large language models efficiently on consumer or edge hardware with reduced memory usage. Gemma 2 9b It is an open-source project aimed at AI developers. The project is open source (MIT). It runs on the command line and API.
hugging-quants builds and maintains Gemma 2 9b It, and it first shipped in 2023. The GitHub repository has been archived. It operates in a well-populated space: PulseGate tracks 14 similar projects. Among its 3 catalogued features are Instruction Tuned, AWQ Quantized, and INT4 Precision.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match