BigVGAN v2 22khz 80band 256x is a neural vocoder developed by NVIDIA and hosted on the Hugging Face Hub. It belongs to the class of neural vocoders for audio-to-audio tasks and supports audio generation from intermediate representations such as mel-spectrograms.
The model originates from the paper BigVGAN: A Universal Neural Vocoder with Large-Scale Training by Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. It is implemented in PyTorch and carries an MIT license. The repository provides installation instructions, usage examples, and a custom fused CUDA kernel that combines anti-aliased activation with upsampling and downsampling to improve inference speed. An interactive local demo built with Gradio is included.
Integration with the Hugging Face Hub allows straightforward loading of the pretrained checkpoint for inference. The specific checkpoint name indicates operation at 22 kHz with an 80-band mel representation and a 256x upsampling factor. The project page, code repository, weights, and demonstration materials are linked from the model card.
Bigvgan V2 22khz 80band 256x sits in PulseGate's Voice, TTS & speech category. It focuses on generating high-quality audio and speech from neural network models. It is built as an open-source project for speech synthesis researchers and developers. Bigvgan V2 22khz 80band 256x is open source under the Open Source license. The product ships for the web, API, and the command line.
It is developed by NVIDIA, and the product first shipped in 2024. Key capabilities include neural vocoder, audio generation, and speech synthesis.
Latest indexed changes and source events
nvidia/bigvgan_v2_22khz_80band_256x verified by the PulseGate indexer
Other apps tracked under the same category.