This is a community-quantized 4-bit version of Alibaba's Qwen2.5-3B-Instruct model, optimized for the MLX framework on Apple silicon. It supports instruction following, tool calling, and can be used locally via Hugging Face Transformers or MLX. The model is designed for efficient on-device inference with significantly reduced memory requirements while maintaining strong performance.
Qwen2.5 3B Instruct is a Foundation models & chat project. It focuses on running large language models efficiently on Apple silicon hardware with reduced memory usage. It is built as an open-source project for developers. The project is open source (Open Source). It runs on the web, the command line, and API.
mlx-community builds and maintains Qwen2.5 3B Instruct, and it first shipped in 2024. It competes in a saturated segment with 25 similar apps in PulseGate's index. Among its 4 catalogued features are 4-bit Quantized, Instruction Tuned, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do