VibeVoice-ASR is an open-source automatic speech recognition model designed for long-form audio transcription and diarization. It supports multiple languages and is suitable for developers and researchers needing customizable ASR solutions.
VibeVoice ASR is a Voice, TTS & speech project. It focuses on providing accurate, open-source speech-to-text transcription for long-form audio in multiple languages. It is built as an open-source project for developers and researchers. The project is open source (MIT). It runs on the web, the command line, and API.
It is developed by Microsoft, and it first shipped in 2025. Development happens publicly on GitHub with 49.4k stars and 49 commits in the last 90 days. Among its 5 catalogued features are speech recognition, long-form audio, and multi-language support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do