Gradium AI's new flagship TTS model hits 81% hard-case accuracy and 216ms latency — ahead of Cartesia and ElevenLabs. What it means for HK video creators.
Gradium AI released a new TTS model that outperforms Cartesia Sonic-3.6 and ElevenLabs v3 on hard-case accuracy. The model is now the default across Gradium's API and Studio with no migration required.
What Gradium AI Released
On August 31, 2026, Gradium AI switched on a new text-to-speech model as its default across its API and Studio platform. The model delivers an 81.0% human-rated pass rate on a 500-sentence hard-case evaluation set spanning five languages (English, German, French, Spanish, Portuguese). Time to first audio clocks in at 216ms at P50 on Coval — 170ms faster than the model it replaces.
The hard-case evaluation set, which Gradium open-sourced on Hugging Face under CC BY 4.0, covers ten criteria: spelling, acronyms, alphanumeric tokens, dates, regular numbers, large and floating numbers, email, plus three composite criteria (Orders, IT Ticket, Claims) that stack several atomic elements into realistic agent turns. Each sentence is scored by independent native-speaker raters, and a sentence only passes if every element is pronounced correctly — one dropped digit fails the entire sentence.
Existing voices, including custom clones, keep working unchanged. There is no migration step for current users.
How It Stacks Up Against the Competition
Gradium's benchmark results are striking. Pooled across all ten criteria and averaged over five languages with equal weight:
- Gradium TTS: 81.0% - Cartesia Sonic 3.6: 75.1% - ElevenLabs v3 Conversational: 65.4% - Fish Audio S2.1 Pro: 49.5% - Inworld TTS 1.5 Max: 46.5%
All benchmarks were generated in August 2026 with default settings. The audio was loudness-normalized, order randomized, and raters capped at 40 comparisons with an enforced break to prevent fatigue.
The 216ms time-to-first-audio at P50 puts Gradium in the same ballpark as the fastest TTS models available today. For real-time voice agent applications, this matters — but for video voiceover production, the accuracy score is the more relevant metric, since post-production editing can compensate for latency.
What This Means for Hong Kong Creators
Gradium TTS does not currently support Chinese or Cantonese, which limits its appeal for Hong Kong creators producing bilingual content. However, for English-language voiceover work — explainer videos, corporate presentations, international brand content — the accuracy improvement over Cartesia and ElevenLabs is meaningful.
If you're producing AI video through Cooly Studio and need accurate English voiceovers for client deliverables, Gradium is worth testing. The commercial-use licensing is clear, and the open-sourced benchmark adds transparency that most TTS vendors don't offer.
For Cantonese voiceover needs, dedicated tools like ElevenLabs and Cooly Studio's integrated voice features remain the best options. But the Gradium benchmark is a reminder that the TTS accuracy race is accelerating — expect more vendors to publish hard-case evaluations in the coming months.
Frequently Asked Questions
Q: Is Gradium TTS free to use? A: Gradium TTS is available through Gradium's API and Studio platform. Pricing is usage-based, similar to ElevenLabs and Cartesia.
Q: Does Gradium TTS support Chinese or Cantonese? A: Not yet. The current model supports English, German, French, Spanish, and Portuguese. Chinese and Cantonese are not in the initial release.
Q: How does Gradium compare to Cartesia Sonic-3.6 for video voiceovers? A: Gradium leads on hard-case accuracy (81.0% vs 75.1%), which matters for accurate number and spelling pronunciation in corporate videos. Cartesia still offers broader language support and a more established ecosystem.
Q: Can I use Gradium TTS with Cooly Studio? A: Yes. Gradium's API can be integrated into any video production pipeline. Generate your video in Cooly Studio, then use Gradium for voiceover narration.
Q: Is the benchmark dataset publicly available? A: Yes. Gradium open-sourced the 500-sentence evaluation set on Hugging Face under CC BY 4.0, making it one of the most transparent TTS benchmarks available.
