This is a community-quantized version of an open-source 20 billion parameter GPT model using MLX format (MXFP4-Q8). It is designed for efficient inference on Apple silicon devices. The model supports chat templates, tool calling, and can be used via Hugging Face Transformers or the MLX framework for local execution.
In the Quantised & converted weights space, Gpt Oss 20b takes a focused approach. It focuses on running large language models efficiently on Apple silicon hardware with reduced memory usage. Gpt Oss 20b is an open-source project aimed at developers. Gpt Oss 20b is open source under the Open Source license. It runs on the web, the command line, and API.
mlx-community builds and maintains Gpt Oss 20b, and it first shipped in 2024. Key capabilities include Model Quantization, MXFP4 Format, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do