VibeVoice-1.5B is a 1.5 billion parameter text-to-speech model released by Microsoft. It produces high-quality, natural speech in both English and Chinese. The model is available on Hugging Face and can be used with the Transformers pipeline for straightforward integration into applications requiring speech synthesis.
VibeVoice 1.5B sits in PulseGate's Voice, TTS & speech category. It focuses on generating natural-sounding speech from text in English and Chinese using a compact open model. It is built as an open-source project for developers. VibeVoice 1.5B is open source under the MIT license. It ships for the web, the command line, and API.
It is developed by Microsoft, and it first shipped in 2025. Development happens publicly on GitHub with 51.8k stars and 6 commits in the last 90 days. Key capabilities include text-to-speech, multilingual support, and high quality synthesis.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Other projects tracked alongside this one