Artificial Analysis scores Zhipu GLM-5.3 at 60 on Intelligence Index, tying Kimi K3 for open-weights lead
On Aug 22 Artificial Analysis published a Zhipu GLM-5.3 evaluation scoring 60 on the Intelligence Index, tying Moonshot Kimi K3 for the open-weights lead. The model hit 1770 Elo on GDPval-AA v2, up 246 points from its predecessor, while remaining 19% cheaper per task than Kimi K3.
On August 22, independent AI evaluation firm Artificial Analysis published its full benchmark suite for Zhipu GLM-5.3. The Intelligence Index composite score of 60 ties Moonshot Kimi K3 as the leading open-weights model today. On the agent-focused GDPval-AA v2 evaluation, GLM-5.3 reached 1770 Elo, a 246-point improvement over its predecessor.
On cost, GLM-5.3 remains 19% cheaper per task than Kimi K3, though token usage rose roughly 20% relative to the prior generation. This trade-off indicates that Zhipu used post-training scaling to lift the capability ceiling substantially, while inference-compute requirements also grew. Model weights are expected to open-source within the next week, meaning developers can soon test the model locally or in private clouds.
Zhipu had already run a publicity push around GLM-5.3 on August 14-17, highlighting a roughly 50% programming uplift, an open-source number one on Terminal-Bench, and emerging white-box cybersecurity capabilities: 84.5% on CyberGym with 2,436 real vulnerabilities surfaced. The Artificial Analysis score confirms Zhipu's broader progression in general intelligence and agent tasks.
Artificial Analysis further notes that GLM-5.3 and Kimi K3 tying for the open-weights lead shows Chinese frontier AI labs are using post-training scaling to compensate for relatively smaller base models. That posture contrasts with the base-first approach favored by US frontier labs.