Alibaba Qwen releases Qwen-Audio-3.1 with five new TTS, ASR, and real-time voice models, slashing prices by up to 95% for HK creators.
Alibaba's Qwen team just dropped Qwen-Audio-3.1 — a lineup of five new AI voice models covering speech recognition, text-to-speech, and real-time conversation. And they cut prices by up to 95%.
For Hong Kong creators producing voiceovers, multilingual content, or interactive voice applications, this is the biggest audio AI launch this quarter. Here's what's new and why it matters.
Five New Models, One Platform
Qwen-Audio-3.1 isn't a single model — it's a full voice pipeline. The five models cover different audio generation and processing tasks:
1. ASR — Automatic speech recognition with multilingual and dialect support. It automatically cleans up filler words and repetitions, giving you cleaner transcripts straight out of the box. For HK creators working with Cantonese or mixed-language content, this is a direct upgrade over earlier ASR tools.
2. ASR-Next — Adds multi-speaker identification with timestamps. It detects emotions, ambient sounds, and machine noise. Useful for podcast editing, interview transcription, or dubbing projects where you need per-speaker timestamps.
3. TTS — Text-to-speech with multilingual synthesis and natural cross-language voice transfer. You control emotion, speed, and style through plain text prompts like "Read this with a sharp, commanding tone." No complicated parameters — just describe the voice you want.
4. TTS-Next — Pairs a language model with diffusion-based generation to produce voice, sound effects, and background audio in a single pass. This is the headline feature: one model, one call, and you get a full audio scene — narration plus ambient sound.
5. Real-time — Simultaneous speaking and listening with instant interruption support. When it detects low mood, it adapts with slower, more empathetic responses. Designed for voice assistants and interactive applications.
Price Cuts Change the Math
Alibaba slashed pricing across the entire lineup. TTS drops roughly 70%, the real-time model about 85%, and ASR up to 95%. At these new prices, AI voice generation becomes cost-viable for workflows that previously relied on human voice actors or expensive TTS APIs.
For a Hong Kong agency running daily voiceover production for social media ads, the savings add up fast. A project that cost HKD 500 in API fees could now run for under HKD 50.
What This Means for HK Creators
Cantonese and multilingual support has been a weak point for many Western TTS models. Qwen's ASR explicitly handles dialect recognition, and the TTS model supports cross-language voice transfer — meaning you can record in Cantonese and generate natural-sounding English or Mandarin voiceovers that preserve your vocal characteristics.
The price cuts make it practical to use AI voice for: - Social media video voiceovers - E-learning and tutorial narration - Interactive voice agents and chatbots - Dubbing and localization projects - Podcast production and editing
Frequently Asked Questions
Q: How does Qwen Audio 3.1 compare to Google's new Flash TTS models? A: Both launched in the same week. Google Flash TTS excels at voice cloning from 30-second samples and stage direction support. Qwen Audio 3.1 offers a broader pipeline with ASR, TTS, and real-time models at lower prices — up to 95% cheaper on ASR.
Q: Can Qwen Audio 3.1 handle Cantonese? A: Yes. The ASR model specifically mentions multilingual and dialect recognition, and the TTS model supports cross-language voice transfer. Test with Cantonese-specific phonetics for best results.
Q: What is TTS-Next and how is it different from standard TTS? A: TTS-Next combines a language model with diffusion to generate voice, sound effects, and background audio in a single pass. Standard TTS produces only speech. TTS-Next generates a full audio scene, useful for video production and immersive content.
Q: Are the new models available through API? A: Yes, through Alibaba's Qwen Cloud platform. Pricing is per-character for TTS and per-hour for ASR. Check the Qwen Cloud console for HK region availability.
Q: Can I use Qwen Audio 3.1 with Cooly Studio? A: Cooly.ai supports AI voice generation in your creative pipeline. Generate voiceovers with Qwen Audio models through API and bring them into Cooly Studio alongside video and image projects for a complete AI production workflow.
