On this page
Open the GLM-5.3 model card and you will find these two lines a few rows apart:
| Benchmark | GLM-5.3 |
|---|---|
| Terminal-Bench 2.1 | 88.2 |
| Terminal-Bench 3.0 | 28.3 |
Same model. Same card. Same day. A sixty-point spread.
Zhipu did nothing wrong here. Publishing both is the honest choice, and more vendors should. The spread is real, and it is the most useful thing in the card, because it shows exactly how much a benchmark's version number carries.
| Model | Terminal-Bench 2.1 | Terminal-Bench 3.0 | Terminal-Bench 4.0 |
|---|---|---|---|
| GLM-5.3 | 88.2% | 28.3% | — |
| Claude Fable 5 | 84.6% | — | 42% |
Sources: GLM-5.3 model card (Hugging Face), MarkTechPost: Alibaba releases Qwen3.8-Max, MarkTechPost: Claude Fable 5.1 and Mythos 5.1
Why the gap exists
Terminal-Bench 3.0 is a harder test with a wider spread of tasks, built deliberately to separate frontier models from near-frontier ones, a job 2.1 had stopped doing well once everyone clustered in the mid-eighties.
The effect is visible in how differently the two versions rank the same pair of models. On 2.1, Claude Fable 5 and Opus 4.8 sit 4.9 points apart. On 3.0, the same pair is separated by 12.7. The harder test spreads the field out, which is the entire point of shipping it.
So a score near 28 on 3.0 and a score near 88 on 2.1 can describe identical ability. Neither is wrong. They are answers to different questions.
The state of the field right now
Here is every published Terminal-Bench figure across the current wave, laid out by version:
| Metric | GLM-5.3 | GLM-5.3-Flash | Qwen3.8-Max | Grok 4.6 | Claude Fable 5 |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.2%† | 84.3%† | 86.6%† | — | 84.6% |
| Terminal-Bench 3.0 | 28.3%†‡ | — | — | 26.5% | — |
| Terminal-Bench 4.0 | — | — | — | — | 42% |
† vendor-reported (self-reported by the model's vendor, not an independent harness). ‡ Terminal-Bench 3.0 is a harder, separate benchmark from 2.1 — the two scores are not comparable.
Sources: GLM-5.3 model card (Hugging Face), MarkTechPost: GLM-5.3-Flash release, MarkTechPost: Alibaba releases Qwen3.8-Max, BenchLM: Grok 4.6 benchmarks & pricing, MarkTechPost: Claude Fable 5.1 and Mythos 5.1
Read the dashes. There is no single version every one of these models has reported. Terminal-Bench 4.0 arrived with Fable 5.1 at the start of September; 3.0 is what GLM-5.3 and Grok 4.6 report; 2.1 is what most of the open models still lead with.
A table like that is the reason "model X beats model Y on Terminal-Bench" is usually an unfinished sentence.
Four rules that make the numbers usable
Carry the version, always. "88.2 on Terminal-Bench" is not a fact. "88.2 on Terminal-Bench 2.1, vendor-reported" is. We changed our own benchmark data to enforce this: each version is a separate field, and our charting code refuses to plot two of them as one series.
Compare down a column, never across. If two models both published 2.1, compare those. When one has 2.1 and the other has 3.0, the comparison is unavailable, and a dash is the correct thing to publish. The arithmetic is easy to do and meaningless.
Treat the harness as part of the score. Terminal-Bench numbers depend on the scaffold that drives them, and vendors run their own. Zhipu's 88.2 came out of Zhipu's harness. Anthropic's numbers come out of Claude Code. Same benchmark, different machinery around the model.
Read a new version as a new instrument. Terminal-Bench 3.0 introduced continuous integration, semantic versioning, and result migrations. The benchmark is now maintained software with releases. Expect the numbers to move when the instrument changes, the same way you would expect a scale to read differently after recalibration.
What earns trust instead
The four rules keep you from being misled by other people's numbers. Getting a number you can act on takes something else: twenty tasks from your own backlog.
Take real tickets your team already closed, with the diffs that resolved them. Run each candidate model against them in the harness you actually use. Score with the test suite you already trust. It takes an afternoon and it produces the only success rate that predicts your bill, because it measures your codebase, your conventions, and your tooling.
That local number does the job every published benchmark is standing in for. It tells you whether a model finishes your work, which is the input to cost per solved task and the only figure worth putting in front of a budget conversation.
Published benchmarks stay useful for what they are good at: a shortlist. They will tell you that GLM-5.3, Qwen3.8-Max and Fable 5 are all in the same broad tier and that a model scoring far below them is unlikely to surprise you. Narrowing five candidates to three is real value. Choosing between the three is a job for your own twenty tasks.
The short version
A benchmark score is three things: a number, a version, and a harness. Publish all three, compare only where all three match, and keep a private eval for the decision that actually costs money.
For how we put open models into production behind that discipline, see running open models at frontier level.
Sources
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project