Voxtral is an open-source AI-powered platform that converts audio and video files into text, supporting over 100 languages. It offers high-accuracy transcription, an API for integration, and is community-driven, making it suitable for developers, researchers, and businesses requiring scalable speech-to-text solutions.
In the Speech to text space, Voxtral takes a focused approach. It focuses on transcribing audio and video files to text accurately and efficiently using AI. Voxtral is an open-source project aimed at developers, researchers, content creators, businesses needing transcription. Voxtral is open source under the Open Source license. Voxtral is available on the web and API, and it can be self-hosted.
Voxtral first shipped in 2024. Key capabilities include audio transcription, video transcription, and multi-language support. The interface is available in 13 languages, including German, English, and Spanish. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do