microsoft/speecht5_asr is a SpeechT5 model trained for automatic speech recognition (ASR). It converts audio input into text and was trained on the LibriSpeech dataset. The model is available on Hugging Face and can be used with the Transformers library.
Speecht5 Asr is a Speech to text project. It focuses on converting spoken audio into text using a pre-trained speech model. Speecht5 Asr is an open-source project aimed at developers. The project is open source (MIT). It ships for the web and API.
Behind Speecht5 Asr is Microsoft, based in the United States, and it first shipped in 2022. Development happens publicly on GitHub with 1.4k stars.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do