Back to news
launchZ.ai智谱2026-08-26

Zhipu open-sources GLM-5.3-Flash multimodal model, priced at 1/40 of Opus 4.8

Zhipu released and open-sourced GLM-5.3-Flash (320B-A18B), the first natively multimodal model in the GLM-5 series, using a hybrid sparse-linear attention architecture. It scored 57 on the AA Intelligence Index, matching Claude Opus 4.8.

On the evening of August 26, Zhipu officially released and open-sourced GLM-5.3-Flash (320B-A18B), the first natively multimodal model in the GLM-5 series. The model uses a hybrid sparse and linear attention architecture, with the Manifold-Constrained Hyper-Connections (mHC) mechanism improving model scaling capability.

GLM-5.3-Flash scored 57 on the Artificial Analysis Intelligence Index, matching Anthropic Claude Opus 4.8 and entering the frontier model range. In Z.ai's Code Bench test, its coding ability is comparable to Opus 4.8.

Pricing is $0.15 input / $0.50 output / $0.03 cached input per million tokens, one-tenth of GLM-5.3, one-twentieth during the promo discount, and just 1/40 of Claude Opus 4.8.

Before formal release, Zhipu tested it anonymously as Ox-Alpha on OpenCode and OpenRouter. The model processed more than 20 trillion tokens in six days and became the most popular model of the week, setting new traffic records on both platforms.

All traffic was served on domestic Chinese chip clusters connected via Zhipu's self-developed high-bandwidth interconnect, with a dedicated inference engine built on top of SGLang. End-to-end service performance improved 3x over the initial baseline on the same hardware.

GLM-5.3-Flash natively integrates vision capability, allowing the model to autonomously decide when to observe and use visual feedback to guide the next action. It targets front-end development, game development, and 3D simulation tasks. Weights are released under the MIT license.

open-sourcemultimodalMOEfrontier-modellow-cost