Wav2vec2 Large Xlsr 53 Punjabi Alternatives
wav2vec2-large-xlsr-53-punjabi is a large wav2vec 2.0 model fine-tuned on Punjabi speech data from the XLS-R multilingual framework. Below are 20 voice, tts & speech apps with similar functionality to Wav2vec2 Large Xlsr 53 Punjabi, matched by what each product actually does — not ranked or scored. Explore each to find the closest fit for your use case.
- Wav2vec2 Large Xlsr Mvc Swahilihuggingface.co
This model is a fine-tuned version of wav2vec2-large-xlsr for automatic speech recognition in Swahili. It accepts audio input and outputs transcriptions. It is intended for developers building voice applications or transcription tools for East African languages.
- Wav2vec2 Large Xlsr 53 Dutchhuggingface.co
The jonatasgrosman/wav2vec2-large-xlsr-53-dutch model performs automatic speech recognition for Dutch. It converts spoken audio into text and forms part of the wav2vec 2.0 family of models that have been fine-tuned on Dutch data from the Common Voice corpus. The model is implemented in the Transformers library and supports direct loading through AutoProcessor and AutoModelForCTC classes or via a high-level pipeline for automatic-speech-recognition tasks. It carries an Apache 2.0 license and is hosted on the Hugging Face platform, where it can be used with PyTorch or JAX backends. The training drew from the mozilla-foundation/common_voice_6_0 dataset and participated in the xlsr-fine-tuning-week and related speech recognition evaluation events. Developers and researchers integrate the model into applications that require transcription of Dutch audio. Access occurs through the Hugging Face ecosystem, including inference providers, notebooks, and local installations of the Transformers library. The page provides code examples for both pipeline and direct model loading to enable immediate use in speech-to-text workflows.
- Wav2vec2 Large Xlsr 53 Amharichuggingface.co
This is a fine-tuned version of the XLS-R wav2vec2 large model (53 languages) further specialized for Amharic automatic speech recognition. Built with the Hugging Face Transformers library, it enables developers to perform ASR on Amharic speech data with improved accuracy over the base multilingual model.
- Wav2vec2 Large Xlsr 53 Japanesehuggingface.co
wav2vec2-large-xlsr-53-japanese is an open-source speech recognition model fine-tuned for Japanese audio. It enables developers to transcribe spoken Japanese into text, supporting research and application development in speech AI and language processing.
- Wav2vec2 Large Xlsr 53 Bretonhuggingface.co
wav2vec2-large-xlsr-53-breton is a fine-tuned version of the XLS-R wav2vec2 model specialized for Breton speech recognition. It was trained on the Common Voice Breton corpus and is distributed on Hugging Face for use with the Transformers library. The model enables developers to build voice applications for the Breton language.
- Wav2vec2 Large Xlsr Latvian Cvhuggingface.co
wav2vec2-large-xlsr-latvian-cv is a fine-tuned version of the XLS-R model for automatic speech recognition in Latvian. It was trained on the Common Voice Latvian dataset and is available on Hugging Face. The model can be used with the Transformers library for transcribing Latvian audio into text.
- Wav2vec2 Large Xlsr Lithuanianhuggingface.co
This model is a fine-tuned version of the XLS-R wav2vec2 large model specifically for Lithuanian automatic speech recognition. Hosted on Hugging Face, it can be used with the Transformers library. It is aimed at developers and researchers working on multilingual or low-resource language speech applications.
- Wav2vec2 Large Xlsr 53 Mongolianhuggingface.co
wav2vec2-large-xlsr-53-mongolian is a fine-tuned version of the XLS-R wav2vec2 model trained on Mongolian speech data from the Common Voice dataset. It supports automatic speech recognition tasks and is available on Hugging Face for use with the Transformers library. The model was created by Anton Lozhkov to improve speech-to-text performance for low-resource languages like Mongolian.
- Wav2vec2 Large Xlsr 53 Hungarianhuggingface.co
Wav2vec2 Large Xlsr 53 Hungarian is an automatic speech recognition model available on Hugging Face. It is a fine-tuned version of the XLSR-53 wav2vec2 architecture adapted specifically for processing Hungarian audio and producing transcriptions. The model accepts audio input and outputs text transcriptions in Hungarian. It is implemented in the Transformers library and supports both high-level pipeline usage and direct loading of the processor and model components. The underlying architecture relies on PyTorch and JAX frameworks. The repository indicates it was trained using data from the Common Voice corpus. Developers and researchers working with Hungarian speech data can integrate the model into applications for transcription tasks. It is delivered as a downloadable model on the Hugging Face platform, with example code provided for loading via the Transformers pipeline or through AutoProcessor and AutoModelForCTC classes. Notebooks for Google Colab and Kaggle are referenced for experimentation. The model carries an Apache-2.0 license, making it available for open use, modification, and distribution. A DOI identifier links to associated metadata for the fine-tuning work.
- Wav2vec2 Large Xlsr Catalahuggingface.co
This is a fine-tuned version of the large XLS-R wav2vec2 model adapted specifically for Catalan speech recognition. Available on Hugging Face, it integrates with the Transformers library. The model was developed to support the Catalan language community and researchers working on low-resource language technologies.
- Wav2vec2 Large Robust L2 English Phoneme Recognitionhuggingface.co
This is a large wav2vec2 model fine-tuned specifically for robust English phoneme recognition. Hosted on Hugging Face, it uses the Transformers library and is suitable for research and development in speech processing. The model has been trained to handle varied acoustic conditions.
- Wav2vec2 Xlsr Khmerhuggingface.co
A fine-tuned version of the wav2vec2 XLS-R model specialized for Khmer automatic speech recognition. It was trained on the OpenSLR Khmer dataset and achieves strong performance on speech-to-text tasks for the Khmer language. The model can be used via the Hugging Face Transformers library for both inference and further fine-tuning.
- Wav2vec2 Large Xlsr 53 Icelandic Ep30 967hhuggingface.co
wav2vec2-large-xlsr-53-icelandic-ep30-967h is a large XLS-R wav2vec2 model fine-tuned on approximately 967 hours of Icelandic speech data. It supports automatic speech recognition for Icelandic and is published on Hugging Face for use with the Transformers library. The model was developed by the Language and Voice Lab.
- Wav2vec2 Large Robust Ft Libritts Voxpopulihuggingface.co
This model is a fine-tuned version of the wav2vec2-large-robust model on the LibriTTS and VoxPopuli datasets. It is designed for high-quality automatic speech recognition in English and is compatible with the Hugging Face Transformers library for easy local inference.
- Wav2vec2 Large Xlsr Galicianhuggingface.co
A large wav2vec2 model fine-tuned on Galician speech data using the XLS-R framework. It performs automatic speech recognition for the Galician language and is available through the Hugging Face transformers library. The model was created to expand speech technology support for lower-resource languages like Galician.
- Wav2vec2 Large Xlsr Kazakhhuggingface.co
This model is a fine-tuned version of wav2vec2-large-xlsr-53 for automatic speech recognition in Kazakh. It was trained on the Kazakh Speech Corpus and achieves strong performance on Kazakh ASR tasks. It can be used via the Transformers library for transcribing Kazakh audio to text.
- Wav2vec2 Xls R Juznevesti Srhuggingface.co
This wav2vec2 model is fine-tuned on the Juznevesti Serbian dataset for automatic speech recognition. It is part of the CLASSLA initiative to support South Slavic languages. The model can be used with the Hugging Face transformers pipeline for transcribing Serbian speech.
- Wav2vec2 Indonesian Javanese Sundanesehuggingface.co
wav2vec2-indonesian-javanese-sundanese is an open-source speech recognition model supporting Indonesian, Javanese, and Sundanese languages. It enables developers to transcribe audio in these languages, facilitating accessibility and language processing for regional users.
- Vakyansh Wav2vec2 Sanskrit Sam 60huggingface.co
vakyansh-wav2vec2-sanskrit-sam-60 is a fine-tuned wav2vec2 model developed for automatic speech recognition in the Sanskrit language. It was created to support low-resource language ASR and is hosted on Hugging Face. The model can be used with the Transformers library for converting Sanskrit speech to text.
- Wav2vec2 Lv 60 Espeak Cv Fthuggingface.co
facebook/wav2vec2-lv-60-espeak-cv-ft is a fine-tuned wav2vec 2.0 model for automatic speech recognition and phonetic transcription. Trained on 60 languages from the Common Voice and espeak datasets, it is published as open weights on Hugging Face. It is intended for developers building multilingual speech applications using the Transformers library.