Alibaba launches CosyVoice Studio, Qwen-Audio tops three global rankings
Alibaba launched CosyVoice Studio on August 7 as China's first all-in-one AI voice productivity platform, with Qwen-Audio ranking first globally on ASR, RealTime, and TTS benchmarks.
On August 7, Alibaba officially launched CosyVoice Studio, China's first all-in-one AI voice productivity platform. Built on the proprietary Qwen-Audio model, the platform bundles three core modules: CosyFlow for voice recording, CosyAgent for voice AI agents, and CosyCreative for audio content creation. This is the first time Alibaba has offered its voice model capabilities externally as an aggregated product.
On the Artificial Analysis benchmark, the Qwen-Audio family ranks first globally with a 1.7% character error rate in automatic speech recognition and takes the top spot in both real-time interaction and text-to-speech, surpassing international models including GPT-Realtime-2. On the same day, DingTalk also became the first enterprise collaboration app to support full-scenario AI voice input.
CosyFlow layers semantic understanding on top of speech-to-text, automatically filtering filler words and turning dictation into formatted emails, meeting minutes, and other structured text. It can distinguish speakers by voiceprint. CosyAgent lets users create voice AI agents with enterprise knowledge, tool calling, and real-time voice interaction through natural language, targeting customer service and telemarketing. CosyCreative ships with thousands of voice timbres, supports voice cloning, and generates conversational podcasts or multi-character audiobooks from uploaded documents or links.
CosyVoice is available now on the iOS, Android, Mac, and Windows app stores for a limited-time free trial. CosyAgent and CosyCreative are currently in whitelist-invite beta and will open more broadly in the near term.