Wav2vec2 Large Xlsr 53 Dutch Alternatives
The jonatasgrosman/wav2vec2-large-xlsr-53-dutch model performs automatic speech recognition for Dutch. Below are 24 voice, tts & speech apps with similar functionality to Wav2vec2 Large Xlsr 53 Dutch, matched by what each product actually does — not ranked or scored. Explore each to find the closest fit for your use case.
- Wav2vec2 Large Xlsr 53 Japanesehuggingface.co
wav2vec2-large-xlsr-53-japanese is an open-source speech recognition model fine-tuned for Japanese audio. It enables developers to transcribe spoken Japanese into text, supporting research and application development in speech AI and language processing.
- Wav2vec2 Large Xlsr 53 Hungarianhuggingface.co
Wav2vec2 Large Xlsr 53 Hungarian is an automatic speech recognition model available on Hugging Face. It is a fine-tuned version of the XLSR-53 wav2vec2 architecture adapted specifically for processing Hungarian audio and producing transcriptions. The model accepts audio input and outputs text transcriptions in Hungarian. It is implemented in the Transformers library and supports both high-level pipeline usage and direct loading of the processor and model components. The underlying architecture relies on PyTorch and JAX frameworks. The repository indicates it was trained using data from the Common Voice corpus. Developers and researchers working with Hungarian speech data can integrate the model into applications for transcription tasks. It is delivered as a downloadable model on the Hugging Face platform, with example code provided for loading via the Transformers pipeline or through AutoProcessor and AutoModelForCTC classes. Notebooks for Google Colab and Kaggle are referenced for experimentation. The model carries an Apache-2.0 license, making it available for open use, modification, and distribution. A DOI identifier links to associated metadata for the fine-tuning work.
- Wav2vec2 Large Xlsr Lithuanianhuggingface.co
This model is a fine-tuned version of the XLS-R wav2vec2 large model specifically for Lithuanian automatic speech recognition. Hosted on Hugging Face, it can be used with the Transformers library. It is aimed at developers and researchers working on multilingual or low-resource language speech applications.
- Wav2vec2 Large Xlsr Catalahuggingface.co
This is a fine-tuned version of the large XLS-R wav2vec2 model adapted specifically for Catalan speech recognition. Available on Hugging Face, it integrates with the Transformers library. The model was developed to support the Catalan language community and researchers working on low-resource language technologies.
- Wav2vec2 Large Xlsr 53 Amharichuggingface.co
This is a fine-tuned version of the XLS-R wav2vec2 large model (53 languages) further specialized for Amharic automatic speech recognition. Built with the Hugging Face Transformers library, it enables developers to perform ASR on Amharic speech data with improved accuracy over the base multilingual model.
- Wav2vec2 Large Xlsr Latvian Cvhuggingface.co
wav2vec2-large-xlsr-latvian-cv is a fine-tuned version of the XLS-R model for automatic speech recognition in Latvian. It was trained on the Common Voice Latvian dataset and is available on Hugging Face. The model can be used with the Transformers library for transcribing Latvian audio into text.
- Wav2vec2 Large Xlsr Mvc Swahilihuggingface.co
This model is a fine-tuned version of wav2vec2-large-xlsr for automatic speech recognition in Swahili. It accepts audio input and outputs transcriptions. It is intended for developers building voice applications or transcription tools for East African languages.
- Wav2vec2 Large Xlsr 53 Icelandic Ep30 967hhuggingface.co
wav2vec2-large-xlsr-53-icelandic-ep30-967h is a large XLS-R wav2vec2 model fine-tuned on approximately 967 hours of Icelandic speech data. It supports automatic speech recognition for Icelandic and is published on Hugging Face for use with the Transformers library. The model was developed by the Language and Voice Lab.
- Wav2vec2 Large Xlsr 53 Bretonhuggingface.co
wav2vec2-large-xlsr-53-breton is a fine-tuned version of the XLS-R wav2vec2 model specialized for Breton speech recognition. It was trained on the Common Voice Breton corpus and is distributed on Hugging Face for use with the Transformers library. The model enables developers to build voice applications for the Breton language.
- Wav2vec2 Cv Behuggingface.co
This model is a fine-tuned version of Facebook's wav2vec2 for automatic speech recognition on the Common Voice dataset for Belgian Dutch. It is available on Hugging Face for integration into speech-to-text pipelines.
- Wav2vec2 Large Xlsr 53 Mongolianhuggingface.co
wav2vec2-large-xlsr-53-mongolian is a fine-tuned version of the XLS-R wav2vec2 model trained on Mongolian speech data from the Common Voice dataset. It supports automatic speech recognition tasks and is available on Hugging Face for use with the Transformers library. The model was created by Anton Lozhkov to improve speech-to-text performance for low-resource languages like Mongolian.
- Wav2vec2 Large Xlsr Galicianhuggingface.co
A large wav2vec2 model fine-tuned on Galician speech data using the XLS-R framework. It performs automatic speech recognition for the Galician language and is available through the Hugging Face transformers library. The model was created to expand speech technology support for lower-resource languages like Galician.
- Wav2vec2 Large Robust Ft Libritts Voxpopulihuggingface.co
This model is a fine-tuned version of the wav2vec2-large-robust model on the LibriTTS and VoxPopuli datasets. It is designed for high-quality automatic speech recognition in English and is compatible with the Hugging Face Transformers library for easy local inference.
- Wav2vec2 Large Robust L2 English Phoneme Recognitionhuggingface.co
This is a large wav2vec2 model fine-tuned specifically for robust English phoneme recognition. Hosted on Hugging Face, it uses the Transformers library and is suitable for research and development in speech processing. The model has been trained to handle varied acoustic conditions.
- Wav2vec2 Large Xlsr 53 Punjabihuggingface.co
wav2vec2-large-xlsr-53-punjabi is a large wav2vec 2.0 model fine-tuned on Punjabi speech data from the XLS-R multilingual framework. It performs automatic speech recognition for Punjabi audio and is available via the Hugging Face transformers library. The model supports researchers and developers creating voice interfaces, transcription tools, or multilingual speech applications for low-resource languages.
- Wav2vec2 Xls R Juznevesti Srhuggingface.co
This wav2vec2 model is fine-tuned on the Juznevesti Serbian dataset for automatic speech recognition. It is part of the CLASSLA initiative to support South Slavic languages. The model can be used with the Hugging Face transformers pipeline for transcribing Serbian speech.
- Wav2vec2 Indonesian Javanese Sundanesehuggingface.co
wav2vec2-indonesian-javanese-sundanese is an open-source speech recognition model supporting Indonesian, Javanese, and Sundanese languages. It enables developers to transcribe audio in these languages, facilitating accessibility and language processing for regional users.
- Wav2vec2 Large Xlsr Kazakhhuggingface.co
This model is a fine-tuned version of wav2vec2-large-xlsr-53 for automatic speech recognition in Kazakh. It was trained on the Kazakh Speech Corpus and achieves strong performance on Kazakh ASR tasks. It can be used via the Transformers library for transcribing Kazakh audio to text.
- Wav2vec2 Lv 60 Espeak Cv Fthuggingface.co
facebook/wav2vec2-lv-60-espeak-cv-ft is a fine-tuned wav2vec 2.0 model for automatic speech recognition and phonetic transcription. Trained on 60 languages from the Common Voice and espeak datasets, it is published as open weights on Hugging Face. It is intended for developers building multilingual speech applications using the Transformers library.
- Wav2vec2 Xlsr Khmerhuggingface.co
A fine-tuned version of the wav2vec2 XLS-R model specialized for Khmer automatic speech recognition. It was trained on the OpenSLR Khmer dataset and achieves strong performance on speech-to-text tasks for the Khmer language. The model can be used via the Hugging Face Transformers library for both inference and further fine-tuning.
- Wav2vec2 Large Ru Goloshuggingface.co
A fine-tuned wav2vec2-large model trained on the Russian GOLOS dataset for automatic speech recognition. The model achieves strong performance on Russian speech-to-text tasks and is distributed via Hugging Face for easy integration into Python applications using the Transformers library.
- Wav2vec2 Large Robust 12 Ft Emotion Msp Dimhuggingface.co
A fine-tuned wav2vec2 model trained on the MSP-Podcast corpus for dimensional emotion recognition. It classifies audio into emotional attributes such as arousal, dominance, and valence. Useful for researchers and developers working on affective computing, call center analytics, or human-computer interaction.
- Wav2vec2 Large Mms 1b Azerbaijani Common Voicehuggingface.co
This model is a fine-tuned version of the MMS (Massively Multilingual Speech) wav2vec2-large model on the Common Voice 15.0 Azerbaijani dataset. It performs automatic speech recognition for Azerbaijani audio. The model is available on Hugging Face and can be used with the Transformers library for local inference.
- Wav2vec2 BERT Cantonesehuggingface.co
This model is a fine-tuned wav2vec2-BERT checkpoint for Cantonese automatic speech recognition. It was trained on the Common Voice dataset and other Cantonese speech corpora. The model is available on the Hugging Face Hub and supports inference through the Transformers pipeline for automatic-speech-recognition.