GigaAM Multilingual is a family of open-source Conformer-based speech-recognition models with 220M and 600M parameter variants. It supports multilingual audio transcription and can be fine-tuned for additional languages by machine-learning developers and speech researchers.
GigaAM Multilingual sits in PulseGate's Voice, TTS & speech category. It focuses on transcribing multilingual speech into text without relying on proprietary speech-recognition services. It is built as an open-source project for machine-learning developers and speech researchers. GigaAM Multilingual is open source under the MIT license. GigaAM Multilingual is available on the web, API, and the command line, and it can be self-hosted.
ai-sage builds and maintains GigaAM Multilingual, and it first shipped in 2024. Development happens publicly on GitHub with 764 stars and 4 commits in the last 90 days. Key capabilities include automatic transcription, multilingual speech recognition, and CTC decoding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do