An automatic speech recognition (ASR) model developed by Microsoft that converts audio input into text. It is designed to output structured JSON and includes a specialized chat template for handling audio tokens. The model is published on Hugging Face and can be used with the transformers library or compatible inference frameworks.
VibeVoice ASR HF is an Other AI project. It focuses on converting spoken audio into accurate, structured text transcriptions. It is built as an open-source project for developers building voice applications. VibeVoice ASR HF is open source under the MIT license. VibeVoice ASR HF is available on the web and API.
Microsoft builds and maintains VibeVoice ASR HF, and it first shipped in 2025. The project is developed in the open on GitHub with 50.1k stars and 2 commits in the last 90 days. Key capabilities include Automatic Speech Recognition, JSON Output, and Audio Transcription.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do