Qwen3-ForcedAligner-0.6B is a specialized 0.6 billion parameter model from the Qwen team designed for forced alignment tasks. It aligns audio recordings with text transcripts at the word or phoneme level, providing precise timestamps. The model uses custom audio tokens and is part of the broader Qwen3 multimodal family.
Qwen3 ForcedAligner 0.6B is a Voice, TTS & speech project. Precise timestamp alignment between spoken audio and corresponding text transcripts. It is built as an open-source project for speech technology developers and researchers. The project is open source (BSD-3-Clause). It runs on the web, the command line, and API.
Qwen builds and maintains Qwen3 ForcedAligner 0.6B, and it first shipped in 2022. The project is developed in the open on GitHub with 24.5k stars and 83 commits in the last 90 days. Among its 3 catalogued features are forced alignment, audio-text alignment, and multimodal tokens.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do