On this page
xAI released Grok 4.6 on 12 August 2026. Zhipu released GLM-5.3 two days later. One is proprietary, one ships weights, and unusually for this season they reported three of the same benchmarks, which makes a real comparison possible instead of the usual exercise in mismatched version numbers.
| Model | Terminal-Bench 3.0 | DeepSWE v1.1 |
|---|---|---|
| Grok 4.6 | 26.5% | 65.9% |
| GLM-5.3 | 28.3% | 66.9% |
Sources: BenchLM: Grok 4.6 benchmarks & pricing, GLM-5.3 model card (Hugging Face)
GLM-5.3 is ahead on both, and on the third shared measure it scores 1769 against Grok 4.6's 1753. That one, GDPval-AA v2, is a rating rather than a percentage.
Three benchmarks, three narrow wins for the open model.
Holding those numbers correctly
Before drawing conclusions, the caveats, because margins this thin deserve them.
The gaps are small. 28.3 against 26.5, 66.9 against 65.9, 1769 against 1753. On three independent measures the same model is ahead each time, which is more interesting than any single result, and it still describes two models of broadly equal capability rather than a gap you would feel.
The sources differ in kind. GLM-5.3's figures come from Zhipu's own model card. Grok 4.6's come from a third-party tracker rather than xAI's own publication. Different harnesses, different scaffolds, no referee.
Terminal-Bench 3.0 is a hard benchmark and these are low scores. Both models sit in the twenties, because 3.0 was built to separate frontier models from near-frontier ones. Do not compare either number to an 88 from a 2.1 leaderboard, see why the version is part of the score.
The specifications
| Metric | Grok 4.6 | GLM-5.3 |
|---|---|---|
| Vendor | xAI | Zhipu AI |
| Released | 2026-08-12 | 2026-08-14 |
| Parameters | Not disclosed | 753B (MoE) |
| Context window | 500K tokens | 1M tokens |
| Native vision | Yes | No |
| License | Proprietary | Open weights (GLM-5.3 licence) |
| Input $/1M | $2 | $1.40 |
| Output $/1M | $6 | $4.40 |
Sources: BenchLM: Grok 4.6 benchmarks & pricing, GLM-5.3 model card (Hugging Face)
The differences that matter are in the bottom three rows.
Context. Grok 4.6 has 500K, GLM-5.3 has 1M. Grok's pricing also steps up above 200K prompt tokens, and it applies the higher rate to *every* token in the request, so a 205K-token prompt is billed at the higher tier throughout rather than only on the overage. On long-context agentic work that is a meaningful cliff to design around.
Licence. GLM-5.3 ships weights; Grok 4.6 does not. One caveat worth keeping: Zhipu announced GLM-5.3 as open-weights with the release to follow after safety evaluation, so confirm you can download the checkpoint before planning around self-hosting.
Price. $1.40 and $4.40 against $2.00 and $6.00. GLM-5.3 is roughly 30% cheaper on both sides before the context surcharge.
Where Grok 4.6 is the better answer
The benchmark table above covers coding and agentic work. Grok 4.6's stronger showing is elsewhere: it leads on knowledge-work evaluations covering long-horizon analyst tasks, professional work and legal evaluation, while losing the coding benchmarks to GPT-5.6 Sol.
That is a real distinction and it should guide the choice. If your workload is research, analysis and drafting rather than repository work, Grok 4.6 is playing to its strength and the comparison above measures the wrong thing.
It also has native image input, where GLM-5.3 is text-only. If your pipeline reads screenshots, diagrams or scanned documents, that decides it, though note that GLM-5.3-Flash *is* natively multimodal at a tenth of GLM-5.3's price, so within the same family there is a vision option.
Where GLM-5.3 is the better answer
For agentic coding, GLM-5.3 wins on the numbers both vendors published, costs about 30% less, holds twice the context without a pricing cliff, and can be downloaded.
That last property is worth more than the benchmark margins. Weights you hold mean inference happens where you choose, which for regulated work turns a procurement conversation into a deployment decision, and it means the model cannot be deprecated or repriced on someone else's schedule.
What we would actually do
Neither of these is the default choice for most work, and that is the honest conclusion.
For the routine majority, GLM-5.3-Flash does the job at $0.15 and $0.50, a tenth of GLM-5.3 and a thirteenth of Grok 4.6, while landing within four points on Terminal-Bench 2.1. Reach for GLM-5.3 when a task has earned it, and for Grok 4.6 when the work is analysis rather than code.
That is the routing pattern from running open models at frontier level, and it is why the interesting question about these two flagships is which one to escalate *to*, rather than which one to standardise on.
Sources
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project