AI SpaceAug 28, 20267 min read

Grok 4.6 vs GLM-5.3: Where the Open Model Lands

Two flagships released two days apart, and three benchmarks both of them published. On all three the open model comes out marginally ahead, at a lower price.

On this page
  1. Holding those numbers correctly
  2. The specifications
  3. Where Grok 4.6 is the better answer
  4. Where GLM-5.3 is the better answer
  5. What we would actually do
  6. Sources

xAI released Grok 4.6 on 12 August 2026. Zhipu released GLM-5.3 two days later. One is proprietary, one ships weights, and unusually for this season they reported three of the same benchmarks, which makes a real comparison possible instead of the usual exercise in mismatched version numbers.

Where both models published a number — %
Where both models published a number — %
ModelTerminal-Bench 3.0DeepSWE v1.1
Grok 4.626.5%65.9%
GLM-5.328.3%66.9%

Sources: BenchLM: Grok 4.6 benchmarks & pricing, GLM-5.3 model card (Hugging Face)

GLM-5.3 is ahead on both, and on the third shared measure it scores 1769 against Grok 4.6's 1753. That one, GDPval-AA v2, is a rating rather than a percentage.

Three benchmarks, three narrow wins for the open model.

Holding those numbers correctly

Before drawing conclusions, the caveats, because margins this thin deserve them.

The gaps are small. 28.3 against 26.5, 66.9 against 65.9, 1769 against 1753. On three independent measures the same model is ahead each time, which is more interesting than any single result, and it still describes two models of broadly equal capability rather than a gap you would feel.

The sources differ in kind. GLM-5.3's figures come from Zhipu's own model card. Grok 4.6's come from a third-party tracker rather than xAI's own publication. Different harnesses, different scaffolds, no referee.

Terminal-Bench 3.0 is a hard benchmark and these are low scores. Both models sit in the twenties, because 3.0 was built to separate frontier models from near-frontier ones. Do not compare either number to an 88 from a 2.1 leaderboard, see why the version is part of the score.

The specifications

MetricGrok 4.6GLM-5.3
VendorxAIZhipu AI
Released2026-08-122026-08-14
ParametersNot disclosed753B (MoE)
Context window500K tokens1M tokens
Native visionYesNo
LicenseProprietaryOpen weights (GLM-5.3 licence)
Input $/1M$2$1.40
Output $/1M$6$4.40

Sources: BenchLM: Grok 4.6 benchmarks & pricing, GLM-5.3 model card (Hugging Face)

The differences that matter are in the bottom three rows.

Context. Grok 4.6 has 500K, GLM-5.3 has 1M. Grok's pricing also steps up above 200K prompt tokens, and it applies the higher rate to *every* token in the request, so a 205K-token prompt is billed at the higher tier throughout rather than only on the overage. On long-context agentic work that is a meaningful cliff to design around.

Licence. GLM-5.3 ships weights; Grok 4.6 does not. One caveat worth keeping: Zhipu announced GLM-5.3 as open-weights with the release to follow after safety evaluation, so confirm you can download the checkpoint before planning around self-hosting.

Price. $1.40 and $4.40 against $2.00 and $6.00. GLM-5.3 is roughly 30% cheaper on both sides before the context surcharge.

Where Grok 4.6 is the better answer

The benchmark table above covers coding and agentic work. Grok 4.6's stronger showing is elsewhere: it leads on knowledge-work evaluations covering long-horizon analyst tasks, professional work and legal evaluation, while losing the coding benchmarks to GPT-5.6 Sol.

That is a real distinction and it should guide the choice. If your workload is research, analysis and drafting rather than repository work, Grok 4.6 is playing to its strength and the comparison above measures the wrong thing.

It also has native image input, where GLM-5.3 is text-only. If your pipeline reads screenshots, diagrams or scanned documents, that decides it, though note that GLM-5.3-Flash *is* natively multimodal at a tenth of GLM-5.3's price, so within the same family there is a vision option.

Where GLM-5.3 is the better answer

For agentic coding, GLM-5.3 wins on the numbers both vendors published, costs about 30% less, holds twice the context without a pricing cliff, and can be downloaded.

That last property is worth more than the benchmark margins. Weights you hold mean inference happens where you choose, which for regulated work turns a procurement conversation into a deployment decision, and it means the model cannot be deprecated or repriced on someone else's schedule.

What we would actually do

Neither of these is the default choice for most work, and that is the honest conclusion.

For the routine majority, GLM-5.3-Flash does the job at $0.15 and $0.50, a tenth of GLM-5.3 and a thirteenth of Grok 4.6, while landing within four points on Terminal-Bench 2.1. Reach for GLM-5.3 when a task has earned it, and for Grok 4.6 when the work is analysis rather than code.

That is the routing pattern from running open models at frontier level, and it is why the interesting question about these two flagships is which one to escalate *to*, rather than which one to standardise on.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…