pyannote/segmentation-3.0 is an MIT-licensed neural model for segmenting speech in audio. It supports voice activity detection, speaker change detection, overlapped-speech detection, and resegmentation through pyannote.audio for developers building speech and diarization workflows.
Segmentation sits in PulseGate's Speech to text category. It focuses on detecting speech activity, speaker changes, overlapping speech, and speaker segments in audio files. It is built as an open-source project for audio and machine learning developers. Segmentation is open source under the MIT license. It runs on the web, the command line, and API, and it can be self-hosted.
It is developed by pyannote, and it first shipped in 2016. The project is developed in the open on GitHub with 10.5k stars and 12 commits in the last 90 days. Among its 9 catalogued features are Voice Activity Detection, Speaker Segmentation, and Speaker Change Detection.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do