A 4 billion parameter version of the Qwen 3.5 model converted to GGUF format with support for multimodal inputs and speculative decoding (MTP). Maintained by Unsloth, it is optimized for fast local inference using tools such as llama.cpp. The model includes comprehensive tokenization and chat templates.
In the Foundation models & chat space, Qwen3.5 4B MTP takes a focused approach. It focuses on enabling efficient local inference of Qwen 3.5 models using GGUF quantization and Unsloth optimizations. It is built as an open-source project for developers. Qwen3.5 4B MTP is open source under the Apache-2.0 license. It runs on the web, the command line, and API, and it can be self-hosted.
Unsloth builds and maintains Qwen3.5 4B MTP, and the product first shipped in 2023. Development happens publicly on GitHub with 68.7k stars and 1.2k commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 7 similar tools. Key capabilities include GGUF Format, Multimodal Support, and MTP Architecture.
Latest indexed changes and source events
unsloth/Qwen3.5-4B-MTP-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.