Soniqo is an open-source, fully offline speech AI toolkit designed for on-device applications such as voice agents, transcription, and speech generation. It enables users to perform diarized transcription, zero-shot voice cloning, and long-form speech synthesis directly on devices including Apple Silicon, Android, Windows, and embedded Linux, with no reliance on cloud APIs or external servers. The platform emphasizes privacy and cost control, ensuring that all processing occurs locally and that there are no per-minute pricing models or data leaving the device.
The toolkit offers a wide range of components and models that can be combined to address varied speech AI tasks. These include real-time and batch transcription with speaker diarization, live captions, dictation, and structured text output from audio. For content creation, Soniqo supports synthesizing speech in multiple voices, rapid voice cloning, audiobook narration, and multi-speaker podcast generation, all performed offline. Its text-to-speech models cover dozens of languages and provide features such as emotion and tempo controls, streaming and batch processing, and zero-shot voice cloning. Voice cloning capabilities are benchmarked across multiple languages and engines, with reference audio available for quality assessment.
Soniqo’s architecture comprises over thirty models addressing tasks such as speech-to-text (including support for up to 1,672 languages), text-to-speech, speaker diarization, voice activity detection, wake-word detection, music and audio production, source separation, speech enhancement, restoration, audio super-resolution, and speech-driven avatar animation. Notable models and pipelines include Qwen3-ASR, Whisper, Parakeet, CosyVoice, VoxCPM2, IndexTTS2, Chatterbox Flash, OmniVoice, and others, with support for technologies such as CoreML, ONNX, MLX, and LiteRT. Performance benchmarks are provided for various devices, demonstrating real-time and low-latency operation across platforms.
Delivery options include a command-line interface, a speech server, and integration through package managers such as Homebrew for Apple platforms and Gradle for Android. Soniqo is distributed under the Apache 2.0 license, reinforcing its open-source nature and suitability for integration into real-world products that require on-device speech AI.
Soniqo sits in PulseGate's Other voice & speech category. It enables developers to run speech recognition, voice cloning, and TTS locally without cloud APIs or data leaving the device. Soniqo is an open-source project aimed at developers building speech-enabled applications. The project is open source (Apache-2.0). It ships for the web, the command line, macOS, Windows, and Linux, and it can be self-hosted.
Soniqo contributors builds and maintains Soniqo, and it first shipped in 2026. The project is developed in the open on GitHub with 945 stars and 445 commits in the last 90 days. Key capabilities include on-device transcription, voice cloning, and speaker diarization. The interface is available in 14 languages, including Arabic, German, and English.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do