This repository contains GGUF format conversions of Google's Gemma models, making them compatible with llama.cpp and other local LLM tools. It enables efficient CPU and GPU inference on consumer hardware. Multiple quantization options are typically provided for different performance and size tradeoffs.
Gemma is a Foundation models & chat project. It focuses on running Gemma language models locally with reduced memory requirements. It is built as an open-source project for developers. The project is open source (Open Source). Gemma is available on the command line.
rahul7star builds and maintains Gemma. It operates in a well-populated space: PulseGate tracks 12 similar projects. Among its 3 catalogued features are Quantized Models, Local LLM, and GGUF Format.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do