Alibaba releases Qwen3.7-Max, trillion-parameter model runs autonomously for 35 hours
Alibaba released Qwen3.7-Max at the May 20 cloud summit, a trillion-parameter MoE that can run autonomously for 35 hours with 1158 tool calls on a new chip platform.
Alibaba officially released the new-generation Qwen flagship model Qwen3.7-Max at the 2026 Alibaba Cloud Summit on May 20. This is the third major iteration of the Qwen flagship within three months.
Qwen3.7-Max uses a trillion-parameter MoE architecture with 36 trillion tokens of pre-training data covering 119 languages, natively supporting a 1M token context window. On the Arena global model blind test, Qwen3.7-Max surpasses Kimi-K2.6, DeepSeek-v4-pro, and GLM-5.1, ranking alongside GPT, Claude, and Gemini's strongest models in the top tier, and topping the Chinese model leaderboard.
The model focuses on long-horizon autonomous execution. In a 35-hour fully autonomous kernel optimization experiment with over 1000 tool calls, Qwen3.7-Max started from scratch on the Pingtouge Zhenwu M890 chip, worked autonomously for 35 hours, conducted 432 kernel evaluations and 1158 tool calls, and achieved 10x speedup on the optimized inference kernel over the official reference implementation.
API pricing is ¥12 input and ¥36 output per 1M tokens. The model landed on Alibaba Bailian platform on May 22.