Qwen3-4B-AWQ is a 4-billion parameter quantized version of the Qwen3 large language model optimized for efficient inference. It supports text generation, tool calling, and can be used with the Hugging Face Transformers library or vLLM. The model is designed for developers who need a capable yet memory-efficient open-weights LLM that can run locally or be self-hosted.
In the Quantised & converted weights space, Qwen3 4B takes a focused approach. It focuses on running large language models efficiently on consumer hardware with reduced memory requirements. It is built as an open-source project for developers. Qwen3 4B is open source under the Open Source license. It runs on the web, the command line, and API.
Behind Qwen3 4B is Qwen, based in China, and it first shipped in 2024. The project is developed in the open on GitHub with 27.4k stars. The category is crowded — PulseGate's index counts 25 comparable projects. Among its 4 catalogued features are Text Generation, Quantized Model, and Tool Calling. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do