This NVIDIA model provides streaming speaker diarization using a Sortformer architecture, supporting up to 4 speakers. It is evaluated on datasets like NOTSOFAR1 and is designed for low-latency, real-time applications. The model is part of NVIDIA's NeMo toolkit and can be integrated into automatic speech recognition pipelines.
Diar Streaming Sortformer 4spk is a Voice, TTS & speech product. It focuses on performing real-time speaker diarization on streaming audio with multiple speakers. Diar Streaming Sortformer 4spk is an open-source project aimed at developers building speech analytics and meeting transcription systems. The project is open source (Apache-2.0). Diar Streaming Sortformer 4spk is available on the web, the command line, and API.
It is developed by NVIDIA (United States), and the product first shipped in 2019. The project is developed in the open on GitHub with 17.8k stars and 132 commits in the last 90 days. Among its 3 catalogued features are Speaker Diarization, Streaming Inference, and 4 Speaker Support.
Latest indexed changes and source events
nvidia/diar_streaming_sortformer_4spk-v2.1 verified by the PulseGate indexer
Other apps tracked under the same category.