This is a community-quantized 8-bit version of the Gemma 4 instruction-tuned model optimized for the MLX framework. It enables efficient local inference on Apple devices. The model includes a custom chat template and is distributed via Hugging Face for use with Python and MLX tooling.
In the Quantised & converted weights space, Gemma 4 E2B It takes a focused approach. It focuses on running the Gemma 4 model efficiently on Apple silicon hardware with reduced memory usage. Gemma 4 E2B It is an open-source project aimed at developers. Gemma 4 E2B It is open source under the MIT license. It ships for the web, the command line, and API.
It is developed by lmstudio-community, and it first shipped in 2024. Development happens publicly on GitHub with 5.2k stars and 419 commits in the last 90 days. PulseGate's similarity index places it among 14 comparable projects. Key capabilities include Quantized Inference, Instruction Tuned, and MLX Optimization.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do