AI SpaceAug 29, 20266 min read

GLM-5.3 vs GLM-5.2: The Same Model, Trained Better

Same base model, same parameter count, same price. Zhipu attributes the whole of GLM-5.3's gain to post-training, and the reasoning benchmark moved eight points.

On this page
  1. The one number that compares directly
  2. Why the coding numbers cannot be lined up
  3. What else changed
  4. The licence question
  5. Should you move
  6. Sources

GLM-5.3 arrived on 14 August, two months after GLM-5.2. The interesting detail is in how Zhipu describes it: the same base model, with the gains attributed to post-training.

Same 753 billion parameters. Same $1.40 in and $4.40 out. What changed is what the model learned after pre-training finished.

The one number that compares directly

The two cards share exactly one benchmark, and it is a good one:

Humanity's Last Exam, with tools — %
Humanity's Last Exam, with tools — %
ModelHLE (with tools)
GLM-5.254.7%
GLM-5.362.5%

Sources: Zhipu GLM-5.2 launch blog (Hugging Face), GLM-5.3 model card (Hugging Face)

54.7 to 62.5. Just under eight points on a hard reasoning test with tool use, from post-training alone, at no change in price.

That is a large move for a point release, and it is the cleanest evidence in either card that the upgrade is real rather than a version bump.

Why the coding numbers cannot be lined up

You would expect a coding comparison here, and it is unavailable, for a reason worth understanding.

GLM-5.2GLM-5.3
Terminal-Bench81.0 (Terminus-2 harness)
Terminal-Bench 2.188.2
Terminal-Bench 3.028.3
SWE-bench Pro62.1
DeepSWE v1.166.9

GLM-5.2 reported Terminal-Bench under the Terminus-2 harness and SWE-bench Pro. GLM-5.3 reports Terminal-Bench 2.1 and 3.0, DeepSWE, FrontierSWE and AutomationBench. The benchmark landscape moved underneath the two releases, so 81.0 and 88.2 are answers to different questions and subtracting them produces a number with no meaning.

Zhipu's own framing is a 50% improvement on its internal Z.ai Code Bench, which is the vendor's private measure. It is a real claim about a real test and there is no way for anyone outside Zhipu to check it, so take it as directional.

For the general case of this problem, see why a benchmark score without its version is not a score.

What else changed

MetricGLM-5.2GLM-5.3
Released2026-06-172026-08-14
Parameters753B (MoE)753B (MoE)
Context window1M tokens1M tokens
Native visionNoNo
LicenseOpen weights (MIT)Open weights (GLM-5.3 licence)
Input $/1M$1.40$1.40
Output $/1M$4.40$4.40

Sources: Zhipu GLM-5.2 launch blog (Hugging Face), GLM-5.3 model card (Hugging Face)

Beyond the reasoning gain, GLM-5.3 adds capabilities GLM-5.2 did not report at all: FrontierSWE 78.1, AutomationBench 48.2, a GDPval-AA v2 rating of 1769, and a cybersecurity result of 84.5 on CyberGym, which Zhipu positions ahead of the frontier proprietary models it compared against.

Context stays at a million tokens. Both remain text-only, which is worth noting because GLM-5.3-Flash, released twelve days later, is the multimodal member of the family.

The licence question

GLM-5.2 shipped open weights under MIT. GLM-5.3 was announced as open-weights with the release to follow roughly two weeks later, after safety evaluation.

If your plan depends on holding the weights, confirm you can download the checkpoint for the specific version you intend to run. An announcement of intent and a published artefact are different things, and the gap between them is exactly the period in which a self-hosting plan can quietly become a hosted-API plan.

Should you move

If you are on GLM-5.2 through an API, yes. Same price, same context, a substantially better reasoning score and a broader set of published capabilities. There is no cost argument for staying.

If you are self-hosting GLM-5.2, check the weights first. GLM-5.2's MIT checkpoint is downloadable today. Confirm GLM-5.3's is too, and under terms that work for you, before planning the migration.

Either way, look at Flash before you assume you need the big model. GLM-5.3-Flash is around nine times cheaper, natively multimodal, MIT, and within four points on the coding benchmarks. For the routine majority of work it is the better default, with GLM-5.3 as the escalation target. That is the arrangement in running open models at frontier level.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…