This is a post-training quantized (QAT) version of Google's Gemma 4 model with 4-bit weights and 16-bit activations (w4a16). It is designed for efficient inference while preserving model quality. The model is published on Hugging Face and intended for developers seeking to deploy Gemma 4 with reduced memory footprint.
In the Foundation models & chat space, Gemma 4 E4B It Qat W4a16 Ct takes a focused approach. It focuses on running large Gemma 4 models efficiently on hardware with limited memory using quantization. Gemma 4 E4B It Qat W4a16 Ct is an open-source project aimed at developers. The project is open source (Open Source). Gemma 4 E4B It Qat W4a16 Ct is available on the web, the command line, and API.
It is developed by Google, and the product first shipped in 2025. PulseGate's similarity index places it among 12 comparable tools.
Latest indexed changes and source events
google/gemma-4-E4B-it-qat-w4a16-ct verified by the PulseGate indexer