This repository hosts an AWQ (Activation-aware Weight Quantization) INT4 quantized variant of Google's Gemma-2-9B-IT model. It enables efficient local inference of a capable 9-billion-parameter instruction-tuned LLM on hardware with limited VRAM. The model is distributed via Hugging Face and is compatible with transformers and vLLM inference backends.
Gemma 2 9b It sits in PulseGate's Foundation models & chat category. It focuses on running large language models efficiently on consumer or edge hardware with reduced memory usage. It is built as an open-source project for AI developers. Gemma 2 9b It is open source under the MIT license. Gemma 2 9b It is available on the command line and API.
Behind Gemma 2 9b It is hugging-quants, and the product first shipped in 2023. The GitHub repository has been archived. It operates in a well-populated space: PulseGate tracks 14 similar tools. Key capabilities include Instruction Tuned, AWQ Quantized, and INT4 Precision.
Latest indexed changes and source events
hugging-quants/gemma-2-9b-it-AWQ-INT4 verified by the PulseGate indexer
⚠ Archived
Other apps tracked under the same category.