A 4 billion parameter version of the Qwen 3.5 model converted to GGUF format with support for multimodal inputs and speculative decoding (MTP). Maintained by Unsloth, it is optimized for fast local inference using tools such as llama.cpp. The model includes comprehensive tokenization and chat templates.
In the Foundation models & chat space, Qwen3.5 4B MTP takes a focused approach. It focuses on enabling efficient local inference of Qwen 3.5 models using GGUF quantization and Unsloth optimizations. Qwen3.5 4B MTP is an open-source project aimed at developers. Qwen3.5 4B MTP is open source under the Apache-2.0 license. It ships for the web, the command line, and API, and it can be self-hosted.
Behind Qwen3.5 4B MTP is Unsloth, and it first shipped in 2023. The project is developed in the open on GitHub with 68.7k stars and 1.2k commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 7 similar projects. Key capabilities include GGUF Format, Multimodal Support, and MTP Architecture.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do