XTTS-v2 is a voice generation model hosted on Hugging Face that performs text-to-speech synthesis and voice cloning from a short audio sample. Developed under the coqui organization, it addresses the need for multilingual speech output and rapid voice replication without requiring extensive training datasets.
The model supports 17 languages and performs voice cloning using only a 6-second audio clip. It transfers emotion and speaking style through cloning, enables cross-language voice cloning, and produces multi-lingual speech at a 24 kHz sampling rate. Compared with its predecessor XTTS-v1, the v2 release adds Hungarian and Korean, incorporates architectural changes for improved speaker conditioning, accepts multiple speaker references with interpolation, and delivers gains in stability, prosody, and overall audio quality.
It is the same or similar model that powers Coqui Studio and the Coqui API. The model card lists a coqui-public-model-license and is provided as a downloadable asset on the Hugging Face platform for integration into speech applications.
In the Voice, TTS & speech space, XTTS takes a focused approach. It focuses on generating natural-sounding speech and cloning voices in multiple languages from short audio samples. XTTS is an open-source project aimed at developers building speech synthesis or voice cloning applications. The project is open source (MPL-2.0). XTTS is available on the web, the command line, and API, and it can be self-hosted.
Behind XTTS is Coqui.ai, and the product first shipped in 2018. The project is developed in the open on GitHub with 45.8k stars. Among its 6 catalogued features are voice cloning, multilingual support, and emotion transfer.
Latest indexed changes and source events
coqui/XTTS-v2 verified by the PulseGate indexer
Other apps tracked under the same category.