microsoft/speecht5_asr is a SpeechT5 model trained for automatic speech recognition (ASR). It converts audio input into text and was trained on the LibriSpeech dataset. The model is available on Hugging Face and can be used with the Transformers library.
In the Voice, TTS & speech space, Speecht5 Asr takes a focused approach. It focuses on converting spoken audio into text using a pre-trained speech model. Speecht5 Asr is an open-source project aimed at developers. The project is open source (MIT). It runs on the web and API.
Behind Speecht5 Asr is Microsoft, based in the United States, and the product first shipped in 2022. The project is developed in the open on GitHub with 1.4k stars.
Latest indexed changes and source events
microsoft/speecht5_asr verified by the PulseGate indexer
Other apps tracked under the same category.