This model is a fine-tuned version of the wav2vec2-large-robust model on the LibriTTS and VoxPopuli datasets. It is designed for high-quality automatic speech recognition in English and is compatible with the Hugging Face Transformers library for easy local inference.
Wav2vec2 Large Robust Ft Libritts Voxpopuli sits in PulseGate's Speech to text category. It focuses on performing high-accuracy English speech-to-text transcription on diverse or noisy audio using a robust fine-tuned model. It is built as an open-source project for developers. The project is open source (Open Source). It ships for the web and the command line.
It is developed by jbetker, and it first shipped in 2021. The GitHub repository has been archived. Among its 4 catalogued features are Automatic Speech Recognition, fine-tuned on LibriTTS, and robust to Noise.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do