Alibaba releases Qwen3-Max-Thinking, setting global reasoning records
Alibaba released Qwen3-Max-Thinking on January 26, trillion-parameter, HLE 58.3 surpassing GPT-5.2, API ¥2.5/¥10 per 1M tokens, 256K context.
On the evening of January 26, Alibaba officially released the Qwen flagship reasoning model Qwen3-Max-Thinking. With over a trillion total parameters and large-scale reinforcement learning post-training, it matches GPT-5.2-Thinking, Claude Opus 4.5, and Gemini 3 Pro across 19 authoritative benchmarks.
Qwen3-Max-Thinking adopts a new Test-time Scaling mechanism that uses experience-extraction reflection to avoid redundant parallel reasoning, achieving more efficient inference within the same context. On the Humanity's Last Exam benchmark it scores 58.3, far above GPT-5.2-Thinking's 45.5 and Gemini 3 Pro's 45.8. The context window is 256K.
The model is live on Qwen Chat with API access at ¥2.5 input and ¥10 output per 1M tokens, offering strong value. On the same day Alibaba open-sourced the Qwen3-TTS speech series supporting voice cloning and anthropomorphic speech generation. Ordinary users can try it free on Qwen PC and web clients.