Vocos-mel-24khz is a neural vocoder available on Hugging Face that synthesizes audio waveforms from mel-spectrograms. Developed by Charactr Inc., it addresses the need for rapid high-quality audio reconstruction in synthesis pipelines by generating spectral coefficients instead of modeling individual time-domain samples.
The model operates as a Fourier-based vocoder trained with a Generative Adversarial Network objective. It produces waveforms in a single forward pass and reconstructs audio through an inverse Fourier transform. This approach closes the gap between time-domain and Fourier-based methods while maintaining synthesis speed and output quality. The repository provides an associated paper with audio samples for evaluation.
It is delivered as a PyTorch model hosted on the Hugging Face platform. Installation for inference uses the command pip install vocos, while training support requires the additional flag pip install vocos[train]. The repository includes usage instructions for reconstructing audio directly from mel-spectrograms.
Vocos-mel-24khz carries an MIT license and is accompanied by the arXiv preprint 2306.00814. It belongs to the class of neural vocoders intended for audio developers and researchers engaged in speech synthesis and audio generation tasks.
Vocos Mel 24khz is an Other AI product. It focuses on synthesizing high-quality audio waveforms from mel-spectrograms for speech and audio applications. Vocos Mel 24khz is an open-source project aimed at audio developers and researchers. The project is open source (Open Source). The product ships for the web, the command line, and API.
Charactr Inc. builds and maintains Vocos Mel 24khz, and the product first shipped in 2023. Among its 4 catalogued features are audio synthesis, neural vocoder, and mel-spectrogram input.
Latest indexed changes and source events
charactr/vocos-mel-24khz verified by the PulseGate indexer
Other apps tracked under the same category.