Alibaba's Qwen Team Ships Qwen-Audio 3.1, Slashes Voice API Prices by Up to 95%
Alibaba's Qwen team has released Qwen-Audio 3.1, a five-model voice stack spanning speech recognition, synthesis and real-time conversation, alongside steep price cuts of up to 95% on ASR, 85% on real-time voice and 70% on text-to-speech, intensifying the price war among AI voice providers.
Alibaba just made building voice features into a product a lot cheaper, and it did so by shipping five models at once. The Qwen team released Qwen-Audio 3.1 on September 23, a complete audio stack covering "understanding, generation, interaction and creation," alongside price cuts the company says reach as high as 95% off its previous API rates.
The cuts aren't uniform across the stack — they scale with how commoditized each task has become. A few specifics define what actually shipped:
- Speech recognition (ASR) pricing drops up to 95%, with a new ASR-Next variant adding multi-speaker diarization, timestamps, emotion detection and ambient sound recognition
- Real-time conversation pricing falls roughly 85%; the model can listen and speak simultaneously, support interruptions mid-conversation, and reportedly slows down and responds more gently when it senses a low mood in the speaker
- Text-to-speech pricing drops about 70%, while a new TTS-Next model generates voice, sound effects and background audio in a single pass for use cases like audiobooks and podcasts
One detail worth flagging for anyone building on this immediately: Alibaba Cloud's own recommended catalog currently still lists the older Qwen-Audio 3.0 TTS Plus as its suggested text-to-speech endpoint, even as the 3.1 recognition and realtime models are promoted — a rollout wrinkle worth testing directly rather than assuming from the announcement alone.
One industry newsletter put the competitive stakes plainly, warning that any team locked into contracts with Deepgram, ElevenLabs or OpenAI's ASR "should rerun their cost-per-hour math before the next renewal window." With voice increasingly treated as the next major interface for AI agents, Alibaba's move signals it intends to compete on price as aggressively in audio as it already has in text models — a pressure that rival voice API providers will now have to answer.

