Voice, TTS & speech Tools & Software
Voice, TTS & speech is part of AI on PulseGate. PulseGate tracks 1,371 Voice, TTS & speech products — 13 indexed in the past week, most recently Musicgen Large.
musicgen-large is a large text-to-music generation model developed by Meta.
Speaker Lite converts WordPress page content into human-like speech using Google Cloud TTS.
nb-wav2vec2-1b-nynorsk is a Norwegian (Nynorsk) speech recognition model based on wav2vec2.
LocalVoice is an offline AI text-to-speech and voice cloning app for macOS.
Ainnate Text To Speech provides advanced AI voice synthesis technology for generating natural-sounding speech from text.
kokoro-cli is an offline-first text-to-speech CLI tool and localhost service powered by Kokoro.
MOSS-TTS is an open-source text-to-speech model developed by the OpenMOSS team.
Nemotron-3.5-ASR-Streaming is an open-source multilingual automatic speech recognition model in GGUF format.
Voxtral-Mini-4B-Realtime-2602-gguf is a quantized GGUF model for real-time multilingual speech-to-text and audio understanding.
VieNeu-TTS-v3-Turbo is a Vietnamese text-to-speech model supporting voice cloning, emotion control, and high-fidelity 48kHz audio.
Quantized GGUF version of the Canary 1B multilingual speech transcription and translation model.
Finnish speech recognition model based on XLS-R wav2vec2.
Qwen3-TTS-12Hz-0.6B-CustomVoice is an open-source text-to-speech model supporting custom voice cloning across multiple languages.
ovos-skill-andersen-tales is a provider skill that reads Hans Christian Andersen fairy tales aloud for voice assistants.
ovos-skill-grimm-tales is a provider skill that reads Brothers Grimm fairy tales aloud for voice assistants.
ovos-skill-ovosblog is a provider skill that reads the OpenVoiceOS blog aloud for voice assistants.
ovos-skill-arxiv-papers is a provider skill that reads arXiv paper abstracts aloud for OpenVoiceOS voice assistants.
GGUF quantized version of OpenAI's Whisper Large V3 for speech-to-text transcription.
Large wav2vec2 Conformer model with rotary embeddings fine-tuned for speech recognition.
Converted Whisper small.en model optimized for CTranslate2 inference.
musicgen-small is a small MusicGen model by Meta for text-to-music generation.
faster-whisper-base.en is a CTranslate2-converted Whisper base model for fast English speech recognition.
MOSS-TTS-v1.5 is an open-source text-to-speech model developed by the OpenMOSS team.
On-device automatic speech recognition models for WhisperKit and Argmax Pro SDK.
A medical automatic speech recognition model developed by Google for radiology and clinical domains.
speaker-diarization-3.0 is a pipeline for determining who spoke when in an audio recording.
Facebook's Massively Multilingual Speech model for language identification.
wav2vec2-lv-60-espeak-cv-ft is an open-source speech recognition model for low-resource languages and phonetic transcription.
wav2vec2-large-xlsr-catala is a speech recognition model fine-tuned for the Catalan language.
faster-whisper-medium is an optimized Whisper medium model for fast speech transcription using CTranslate2.
whisper-bemba-stt is a fine-tuned Whisper model for automatic speech recognition in the Bemba language.
Whisper large-v3-turbo model converted for CTranslate2.
PersonaPlex-7B-MLX-4bit is a quantized speech-to-speech model optimized for Apple Silicon.
tts-daemon is a local HTTP/WebSocket Text-to-Speech gateway with pluggable providers, starting with Piper.
Tokenizer for Qwen3-TTS enabling extreme bitrate reduction and low-latency speech processing.
GGUF quantized versions of Qwen3-TTS text-to-speech models for local inference.
GGUF quantized version of NVIDIA's Parakeet TDT 0.6B automatic speech recognition model.
NVIDIA's streaming speaker diarization model using Sortformer for up to 4 speakers.
wav2vec2-large-robust-ft-libritts-voxpopuli is a fine-tuned speech recognition model for English audio.
faster-distil-whisper-medium.en is a CTranslate2-optimized English speech-to-text model based on Distil-Whisper.
granite-4.0-1b-speech is an IBM Granite model for automatic speech recognition tasks.
Qwen3-ASR-1.7B is an open-source speech recognition model provided in GGUF format for local inference.
Voxtral-Small-24B-2507-gguf provides GGUF quantized versions of an open audio-language model for speech recognition and transcription.
nemotron-speech-streaming-en-0.6b is an English automatic speech recognition model from NVIDIA.
facebook/mms-lid-126 is a massively multilingual speech model for language identification supporting 126 languages.
GGUF quantized version of OpenAI's Whisper medium model for speech recognition.
Speaker diarization pipeline that identifies who spoke when in audio.
A fine-tuned Whisper model for Kazakh speech recognition.
Qwen3-ASR-0.6B-gguf is a quantized automatic speech recognition model in GGUF format.
S2T Small Librispeech ASR is a speech-to-text model trained on the LibriSpeech dataset for automatic speech recognition in English.