Skip to content
Alternatives
Software like Voicebox
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- Voiceboxmetademolab.comVoicebox is an open-source AI model developed by Meta AI for advanced speech generation and editing. It supports multilingual text-to-speech, speech infilling, noise removal, and cross-lingual voice transfer. Designed for researchers and developers, it enables flexible, high-quality audio synthesis and editing across multiple languages.
- Pocket Voiceh3manth.comPocket Voice enables users to clone voices and generate speech directly within their web browser. Designed to operate fully client-side, it leverages ONNX Runtime and WebAssembly to process all data locally, ensuring that no audio or text leaves the user's machine. This approach removes the need for server interaction or API keys, emphasizing privacy and user control over their data. The tool allows individuals to record their own voice through a microphone or upload an audio clip, from which it extracts voice embeddings using the mimi encoder. Once a voice is cloned, users can type text and have it spoken in the cloned voice, with audio generated and streamed in real time via AudioWorklet. Pocket Voice also supports saving cloned voices for later reuse, providing a personalized voice library within the browser session. Additionally, it offers built-in voices and language options including English, German, Italian, Portuguese, and Spanish, enabling text-to-speech generation in multiple languages. A feature called Voice Studio lets users select a preset, type text, and generate speech using either a newly cloned or a saved voice. The process requires only about five seconds of recorded speech or a short uploaded audio clip to create a voice clone. All ONNX models are loaded from the HuggingFace CDN, but inference and processing remain entirely within the browser, maintaining strict data privacy. Pocket Voice is intended for legitimate, ethical use and cautions against using the tool for voice cloning without consent or for any harmful, illegal, or deceptive purposes. The tool is built by Hemanth HM and is powered by Pocket TTS. Its client-side architecture makes it suitable for users seeking privacy-preserving voice cloning and text-to-speech capabilities without relying on external servers or cloud services.
- VoiceStudiovoicestudio.shVoiceStudio is an open-source desktop application for local text-to-speech, voice cloning, dubbing, dictation, and transcription. It supports 646 languages, offline operation, voice presets, and an OpenAI-compatible local API for developers.
- LocalVoicehajeklabs.comLocalVoice is a fully offline macOS application for text-to-speech and voice cloning. It generates natural, expressive speech locally on your device, allows voice style control through descriptions or reference audio, and exports to WAV. Ideal for video creators, podcasters, and developers who need private, subscription-free voice generation.
- Voice Cloning Studiohuggingface.coVoice Cloning Studio is a Hugging Face Space built on XTTS-v2 that converts text to speech in multiple languages. It automatically scans input for emojis, URLs, numbers, and other elements that could cause pronunciation problems, lists detected issues, and generates high-quality audio output. The tool is designed for users needing clean, natural voice synthesis.
- Voice Clone Multilingualhuggingface.coVoice Clone Multilingual is a web application that enables users to create spoken audio that mimics a specific person's voice in multiple languages. Users upload a voice sample, enter text, and receive AI-generated speech, ideal for content creators and voiceover artists.
- Voice Clone AI Podcasthuggingface.coVoice Clone AI Podcast enables users to generate podcasts by cloning voices using AI. Users can input scripts or upload audio, and the app produces narrated podcasts with synthetic voices. It is designed for podcasters and content creators seeking automated voice production.
- VoxBoostervoxbooster.comVoxBooster is a Windows application designed for real-time AI-powered voice changing and voice cloning. It enables users to transform their voice instantly during live calls or recordings, offering neural voice cloning, a variety of voice effects, a customizable soundboard, and dictation features. The tool operates directly through the user's microphone, ensuring compatibility with any application that uses mic input, including platforms such as Discord, Teams, Meet, Zoom, OBS, Streamlabs, Twitch, and games. Key features include the ability to select from a library of preset voices or create custom voice clones, apply over 20 stackable voice effects (such as Villain, Cartoon, Demon, Helium, Robot, Alien, and more), and save custom effect presets. The soundboard allows users to add their own audio files, assign hotkeys, and play sounds system-wide so they are broadcast through the microphone in any app. Dictation is supported through offline speech-to-text using Whisper, and live translation is available via Google for several language pairs. Users can also type text and have it spoken aloud in their cloned voice, which is useful for streams and interactive content. Additional capabilities include studio noise suppression to minimize background sounds, floating overlays for on-screen controls, usage statistics tracking, and a fully translated user interface supporting ten languages. Hotkeys can be assigned to a wide range of actions, including muting, changing voices, playing sounds, and more. The app provides live meters for monitoring input and output levels, and does not require installation of a virtual audio driver—any app that accesses the microphone receives the transformed output automatically. VoxBooster is available exclusively for Windows 10 and 11. It offers a 3-day free trial with full access to all features and does not require a credit card to begin. The pricing model is subscription-based, with monthly, quarterly, and annual payment options, and all plans include every feature without restrictions based on tier.
- Voice Cloninghuggingface.coVoice Cloning is a web app that generates synthetic voices by combining user-provided text and sample audio files. It supports multiple languages and is designed for voiceover artists and content creators.
- Voice Clonevoice-clone.orgVoice Clone is a web app that allows users to clone their own voice or others and generate natural-sounding speech from text. It supports voice sample uploads, instant TTS, and is designed for content creators, educators, and anyone needing custom voiceovers.
- Voice Clonervoicecloner.orgVoiceCloner lets users upload or record a voice sample, create a private voice profile, and generate speech from text in the cloned voice. It supports multilingual speech, voice cues, speed control, generated-result history, and optional signed-in profile storage.
- Voicvvoicv.comVoicv is an AI-powered platform that enables users to clone voices, generate natural speech from text, and transcribe speech to text in multiple languages. It offers features like emotion control, AI avatars, and API access, making it suitable for content creators, businesses, and developers seeking advanced audio transformation tools.
- Voice Clonehuggingface.coVoice Clone is a web application that allows users to generate audio of typed text spoken in the style of an uploaded voice recording. It is useful for content creators, voiceover artists, and developers seeking custom voice synthesis.
- TTS Voice Clonerhuggingface.coTTS Voice Cloner allows users to upload a short WAV sample and enter text to generate speech in that cloned voice. It supports multiple languages and outputs a new audio file, making voice-over production faster.
- Voice Clone Multilingualhuggingface.coThis tool allows users to upload a short audio sample of a speaker and enter text. It then synthesizes the text in the chosen language using a voice that closely resembles the reference speaker. The output is a downloadable audio file suitable for dubbing, accessibility, or creative projects.
- Narration Boxnarrationbox.comNarration Box is a browser-based AI voice production platform for creating voiceovers, audiobooks, tutorials, podcasts, and ads. It provides multilingual text-to-speech, voice cloning, emotion and accent direction, pronunciation controls, and editable audio projects for creators and teams.
- Dubvoxcoderluii.devDubvox is an AI voice translation tool that uses voice cloning to return uploaded audio in the speaker’s own voice. It is aimed at creators and other users who want translated audio that keeps the original voice rather than replacing it with a generic robot voice. The service says it preserves emotion, tone, personality, pitch, and speaking style. Its workflow is described in three steps: upload an audio file, let the system clone the voice and translate the content, and then download the translated audio. The output is delivered in minutes, and the site describes the result as the same voice in a new language. It also says the tool supports any language pair, with examples such as Spanish to English, Japanese to French, and Arabic to German. Supported input formats named on the site are MP3, WAV, M4A, and OGG, and it also allows a link to be pasted instead of uploading a file directly. The product is described as built for creators who think globally, with use cases including podcasts, YouTube videos, and course content. The page also says it is designed for translating a user’s own content or content for which permission has been granted, and it refers to this as a consent-first architecture. Dubvox is presented as an invite-only beta with early access pricing and priority onboarding for the first 500 users. It is a CoderLuii project.
- Voice Cloning Demohuggingface.coVoice Cloning Demo is a web application that converts user-inputted text into spoken audio in various languages. It enables users to generate audio files with cloned voices, making it useful for content creators and developers needing realistic voiceovers or speech synthesis. The app is free to use and accessible via browser.
- Chatterbox TTS APIchatterboxtts.comChatterbox TTS API provides a local text-to-speech API server that is compatible with the OpenAI TTS API, allowing users to generate multilingual, voice-cloned speech. Designed as a drop-in replacement for the OpenAI TTS API, it enables integration with applications that already use the OpenAI interface, such as Open WebUI and AnythingLLM, without requiring code changes. The platform addresses the need for customizable, on-premises speech synthesis with support for a wide range of languages and personalized voice options. A notable feature of Chatterbox TTS API is its native multilingual voice cloning, supporting 22 languages including Arabic, Danish, German, Greek, English, Spanish, Finnish, French, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Dutch, Norwegian, Polish, Portuguese, Russian, Swedish, Swahili, and Turkish. Users can upload voice samples and assign them to specific languages for optimal language-aware voice synthesis. The system automatically detects and applies the assigned language during speech generation, ensuring accurate and natural-sounding results across different linguistic contexts. The tool offers a set of features aimed at facilitating the management and deployment of text-to-speech capabilities. These include a voice library management system for uploading and organizing custom voices by name, smart text processing with automatic chunking for handling long texts, and a real-time status monitor to track TTS progress, statistics, and history. Chatterbox TTS API also provides a clean and intuitive web interface for generating speech and managing voices. For deployment, Chatterbox TTS API supports local installation and is Docker-ready, allowing for containerization with persistent voice storage. The installation process is accessible through standard commands, and the platform is available via its GitHub repository. Chatterbox TTS API is not affiliated with Resemble AI.
- KOVOXkovox.ioKovox is an AI-powered web platform that enables users to generate ultra-realistic voiceovers, clone voices, and convert text to expressive audio. It is ideal for creators, brands, and developers seeking lifelike speech synthesis for various applications.
- MorVoicemorvoice.comMorVoice is a platform for generating lifelike AI voices, cloning unique vocal identities, and producing audio content. It offers text-to-speech, voice cloning, and a marketplace for buying and selling AI-generated voices, with API access for developers and tools for creators and businesses.
- Voxabotvoxabot.comVoxabot is a web-based text-to-speech platform that allows users to create voiceovers using leading TTS engines like Google, Azure, and AWS. It features an SSML editor, supports over 150 languages, and provides waveform previews and export options for content creators and businesses.
- AI Voice Generator and Text to Speechvoicecreator.proVoice Creator Pro provides AI voices for text-to-speech, voice cloning from short audio samples, custom voice design, speech-to-text transcription, video dubbing, and subtitle generation. It supports over 600 languages and is used for audiobooks, YouTube voiceovers, podcasts, e-learning, and game development. A desktop app is also available.
- Vocallab AIvocallab.aiVocallab AI is a browser-based voice studio for creating AI voiceovers from text. It offers voice cloning and design, speech-to-text, audiobooks, dubbing, karaoke-style captions, and MP3 or SRT export for creators publishing videos, ads, podcasts, and other media.
Ranked by how close each one sits to Voicebox in the index, not by popularity. Back to Voicebox →