This is a GGUF-quantized version of the Voxtral-Mini-4B model optimized for real-time automatic speech recognition across 13 languages. It supports streaming transcription and can be used with tools like transcribe.cpp. The model is designed for local, on-device audio understanding and is available for download from Hugging Face.
In the Voice, TTS & speech space, Voxtral Mini 4B Realtime 2602 takes a focused approach. It focuses on running efficient real-time speech recognition and audio-language modeling locally. Voxtral Mini 4B Realtime 2602 is an open-source project aimed at developers. The project is open source (MIT). The product ships for the web and the command line.
Behind Voxtral Mini 4B Realtime 2602 is Handy Computer, and the product first shipped in 2026. The project is developed in the open on GitHub with 1.5k stars and 412 commits in the last 90 days. Among its 3 catalogued features are speech-to-Text, Streaming ASR, and Multilingual Support.
Latest indexed changes and source events
handy-computer/Voxtral-Mini-4B-Realtime-2602-gguf verified by the PulseGate indexer