Wav2vec2 Large Xlsr 53 Hungarian is an automatic speech recognition model available on Hugging Face. It is a fine-tuned version of the XLSR-53 wav2vec2 architecture adapted specifically for processing Hungarian audio and producing transcriptions.
The model accepts audio input and outputs text transcriptions in Hungarian. It is implemented in the Transformers library and supports both high-level pipeline usage and direct loading of the processor and model components. The underlying architecture relies on PyTorch and JAX frameworks. The repository indicates it was trained using data from the Common Voice corpus.
Developers and researchers working with Hungarian speech data can integrate the model into applications for transcription tasks. It is delivered as a downloadable model on the Hugging Face platform, with example code provided for loading via the Transformers pipeline or through AutoProcessor and AutoModelForCTC classes. Notebooks for Google Colab and Kaggle are referenced for experimentation.
The model carries an Apache-2.0 license, making it available for open use, modification, and distribution. A DOI identifier links to associated metadata for the fine-tuning work.
Wav2vec2 Large Xlsr 53 Hungarian sits in PulseGate's Speech to text category. It enables accurate automatic speech recognition for Hungarian language audio recordings. It is built as an open-source project for speech technology researchers and developers. Wav2vec2 Large Xlsr 53 Hungarian is open source under the Open Source license. It runs on the web, the command line, and API, and it can be self-hosted.
It is developed by jonatasgrosman, and it first shipped in 2022. Among its 5 catalogued features are speech recognition, hungarian language, and audio transcription.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do