NVIDIA's Parakeet CTC 1.1B is a large-scale automatic speech recognition model trained for high-accuracy transcription across various domains. Built with the NeMo toolkit, it uses Connectionist Temporal Classification (CTC) and supports integration with Transformers and other inference providers. The model is publicly available on Hugging Face and is used by developers building voice interfaces, transcription services, and audio analysis tools.
In the Voice, TTS & speech space, Parakeet Ctc 1.1b takes a focused approach. It focuses on converting spoken audio into accurate text transcripts with high performance and low latency. It is built as an open-source project for developers and researchers. The project is open source (Apache-2.0). It ships for the web, the command line, and API.
It is developed by NVIDIA (United States), and it first shipped in 2019. Development happens publicly on GitHub with 17.8k stars and 133 commits in the last 90 days. Among its 5 catalogued features are automatic speech recognition, CTC decoding, and high accuracy. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do