On this page
- Is Jev better than GLM-5.2?
- At a glance: what each model does
- How they did on our tests
- Why the rebuild won: ask for facts, keep the policy in code
- Who should decide the items Jev is unsure about?
- What does each cost per month?
- Does the newer GLM-5.3 beat Jev on clean tasks?
- When is GLM-5.2 still the right choice?
- FAQ
- Sources
Jev beat GLM-5.2 at sorting our operations inbox, but only after we rebuilt the job. We ran the tests from 21 to 23 September 2026. On GLM-5.2's own prompt, Jev lost, 89.0% to 90.1%. Asked plain yes/no facts instead, it won 90.7% to 84.0% on 150 new items. Each item cost $0.0004 on Jev, against $0.0051.
Picking a label from a fixed list? Rebuild the job for Jev. Writing anything? Keep GLM-5.2.
- Jev is a decision model from TypeSafe that answers questions and never writes text. It reached Cloudflare AI Gateway on 18 September 2026.
- GLM-5.2 is a writing model from Zhipu AI's General Language Model (GLM) family. You can download it and run it yourself.
- Rebuilt, Jev caught 48 of 50 client items; GLM-5.2 caught 42.
- Speed: Jev answered in under a second per item, GLM-5.2 in several.
- With Claude checking the items Jev was unsure about, the rebuilt setup caught all 50 client items.
Is Jev better than GLM-5.2?
Jev is better at choosing labels, once the task is shaped for it. Jev cannot do anything that needs written output, because it writes none.
Pick Jev if you sort, route or flag items against a fixed list and can restructure the prompt around it. Pick GLM-5.2 if you need summaries, extracted fields or written reasons. It also suits you if you want to host the model yourself or feed it long documents. Use both if your pipeline does some of each. That is what we run.
For a primer on how Jev differs from a chat model, see what is Jev.
At a glance: what each model does
Jev picks answers; GLM-5.2 writes them. Everything else in the table follows from that. A token is about three quarters of a word.
| Jev | GLM-5.2 | |
|---|---|---|
| Who makes it | TypeSafe | Zhipu AI (Z.ai) |
| What it does | Answers questions you set: yes or no, pick one of up to 255 options, or a score | Reads, reasons and writes free text |
| Writes text? | No | Yes |
| Price per million tokens | $0.042 to read, output free | $1.40 to read, $4.40 to write |
| Speed | 70 to 500 milliseconds (TypeSafe's own figure); 0.3 seconds an item in our tests | 7 to 10 seconds an item in our tests |
| How much it reads at once | 32,000 tokens | 1 million tokens |
| Run it yourself? | No: hosted API through TypeSafe early access or Cloudflare AI Gateway | Yes: open weights under the open-source MIT licence, 753 billion parameters |
Jev's price and speed come from TypeSafe's launch post, its reading limit from Cloudflare's model page. The other column comes from Z.ai's pricing page and the GLM-5.2 launch post. TypeSafe's direct API is still early access with a waitlist.
How they did on our tests
On our own inbox, Jev lost a straight swap by about a point. Rebuilt, it won by 6.7 points on items neither setup had seen.
| Jev | GLM-5.2 | |
|---|---|---|
| Same prompt, 303 build items | 89% | 90.1% |
| Rebuilt, 150 new items | 90.7% | 84% |
Source: Our tests, 21 to 23 September 2026. Correct labels written by Claude against a written policy.
The task is our operations inbox. It labels every incoming item as client, internal, noise or unknown. Items include emails, calendar invites, call recordings, WhatsApp messages and AI coding sessions. It also tags a signal, such as a request, a commitment or money. Missing a client item is the costly mistake, so we count those separately.
Round one: Jev dropped into the old prompt
We gave Jev the instructions GLM-5.2 already used. Both ran on the 303 items we built the system with.
| Jev | GLM-5.2 | |
|---|---|---|
| Items labelled right | 89.0% | 90.1% |
| Client items caught | 88.6% | 94.3% |
| Seconds per item | 0.28 | 7.4 |
| Cost per item | $0.00008 | $0.0051 |
Jev was cheaper and faster, and it lost on the number that matters most. A test on one shared prompt mostly measures which model follows that prompt better. This prompt was written for the incumbent.
Round two: Jev rebuilt, tested blind
We then rebuilt the job around Jev. Both models ran on 150 items held back from all the building work, 50 of them from clients.
| Rebuilt Jev | GLM-5.2 | |
|---|---|---|
| Items labelled right | 90.7% | 84.0% |
| Client items caught (of 50) | 48 | 42 |
| Items wrongly called client | 0 | 0 |
| Signal tagged right | 85 to 87% | 82.5% |
| Cost per item | $0.0004 | $0.0051 |
| Seconds per item | 0.3 | 7 to 10 |
Three limits apply. First, Claude wrote the correct labels for the 150 items, working from a written labelling policy. A human check of 20 of them is still pending. Second, GLM-5.2 changes its own label on 16% of re-runs, so its score moves between runs. Third, this is one inbox with one policy; a different job could land differently.
Why the rebuild won: ask for facts, keep the policy in code
The rebuild won because Jev answers small factual questions well. Our code then applies the rules that turn those facts into a label. It is the approach we recommend for any Jev job: every question in one call, no prose to parse, and low-confidence answers go to a person or a rule.
The first version asked Jev to read our policy and name the label. The rebuilt version asks three kinds of question in one call:
- Who sent it, picked from a list of the 78 organisations we work with. A "new company" flag caught every sender not on it.
- About 12 yes-or-no facts about the item. TypeSafe calls this question type a "noul".
- Its own best guess at the label, used as one more input rather than the answer.
Two small scoring models turn those answers into the label and the signal. We trained them on 300 labelled items. The gain shows most on the signal: 64% right when asked directly, 85 to 87% through the scoring model.
We also leave our own companies off the sender list. The rebuild also sends far more of each item, up to 36,000 characters. The old prompt was cut much shorter. How to use Jev covers the question design in detail.
Who should decide the items Jev is unsure about?
Claude should, with 90.2% of the unsure items right. GLM-5.2, as it runs in production, got 70.6%. Jev alone did better than that, at 76.5%.
Every answer comes back with a confidence score. About 69% of our traffic scores 0.8 or higher, and those answers were 98.1% right. The rest, plus anything from a new company, goes to a second decider. We tested six candidates:
| Who decides the unsure items | Right, unseen items (51) | Right, build items (111) |
|---|---|---|
| GLM-5.2 as it runs in production | 70.6% | 65.8% |
| Jev on its own | 76.5% | 77.5% |
| GLM-5.2 given the organisation list and policy | 80.4% | 75.7% |
| gpt-oss-120b | not tested | 60.4% |
| GPT-5.4 mini | not tested | 72.7% (22 items only) |
| Claude | 90.2% | 87.4% |
With Claude on the unsure third, the whole system got 95.3% of the 150 unseen items right. It caught all 50 client items. The labelling caveat bites harder here, since Claude wrote the correct labels. A person still reviews what the system flags; see human-in-the-loop AI for how that fits. For a much larger labelling job, see Jev vs Claude.
What does each cost per month?
Jev costs about $1.50 a month for our inbox. GLM-5.2 costs about $18.
Our inbox sees about 120 items a day, and the monthly figures come from the per-item costs in round two. The Claude lane for unsure items is left out of the per-item figure.
At our volume, neither bill decides anything; the accuracy gap does. The cost gap matters at higher volumes. Jev pricing works through larger jobs and the billing details.
Does the newer GLM-5.3 beat Jev on clean tasks?
Yes, by a little. On 25 September 2026 we ran GLM-5.3, the next GLM release, against Jev on a fresh set of pick-one tasks. Each task had a short list of described options.
The newer model got 99.0% right and Jev 97.2%. GLM-5.3 cost 41 times as much per task. It also took about 6 times as long at the median, 2.9 seconds against 0.46.
| Accuracy | Both halves of a pair right | |
|---|---|---|
| Jev | 97.2% | 95% |
| GLM-5.3 | 99% | 98% |
Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
The inbox result above still stands, because the tests differ. The inbox win came from rebuilding the job around Jev: small yes/no facts, our organisation list, and scoring models in code. The benchmark asked both models the same single question, with a clear line describing each option.
When the task is that clean, the large model is slightly ahead. When the job is messy and you can reshape it, Jev can pull ahead.
You do not have to choose one. Jev answered every task, and GLM-5.3 re-answered only when Jev gave one of three answers it often got wrong. That mix got 98.1% right at $86 per million tasks, against 99.0% and $809 for GLM-5.3 alone. LLM cascade explains how we picked the three.
Claude wrote the 400 tasks, and 400 is a small set. GLM-5.3 ran with its default reasoning on. Decision model vs LLM covers when 1.8 points is worth 41 times the price, and Jev vs Tev has the full table.
When is GLM-5.2 still the right choice?
GLM-5.2 is the right choice for any step that has to write: pulling out fields, summarising, or explaining a decision. We kept extraction on it.
Jev cannot produce a name, a date string or a sentence. TypeSafe says so plainly: it gives up text generation. Pydantic AI's integration already plans for that split. It hands text fields to a second model.
Long input is the other case. The writing model reads up to a million tokens; Jev reads 32,000. GLM-5.2 is also the only one of the two you can download, which matters if your data cannot leave your servers. For the writing work, compare it with its successor in GLM-5.3 vs GLM-5.2.
FAQ
Is Jev a replacement for GLM-5.2?
Only for decisions from a fixed list, like routing or flagging. Jev does not write text, so summaries, extraction and written explanations still need a writing model. In our inbox, Jev now picks the labels and GLM-5.2 still pulls out the details. Treat Jev as a replacement for one step of a pipeline, and keep a writing model for the rest.
Why did Jev lose when we first swapped it in?
The prompt was written for the other model. On that shared prompt, Jev scored 89.0% to 90.1% and caught fewer client items. We then changed the questions to small facts and moved the labelling rules into code. After that, Jev won on unseen items, 90.7% to 84.0%. A straight swap mostly tests prompt-following, not what the model can do.
How much cheaper is Jev than GLM-5.2?
In our rebuilt setup, Jev cost $0.0004 per inbox item against $0.0051. At about 120 items a day, that is about $1.50 a month against $18. On list prices, Jev charges $0.042 per million tokens read, and output is free. GLM-5.2 charges $1.40 per million tokens read and $4.40 per million written.
Can I run Jev or GLM-5.2 on my own servers?
GLM-5.2, yes: its weights are published under a permissive open-source licence. Jev is a hosted service, available through TypeSafe's early access programme and Cloudflare AI Gateway. Cloudflare lists Jev with zero data retention, meaning the provider does not keep what you send.
How reliable are these test results?
Our results come from one inbox, tested 21 to 23 September 2026. Claude wrote the correct labels for the 150 unseen items against a written policy. A human check of 20 of them is pending. GLM-5.2 changes its own answer on 16% of re-runs. Treat the gaps as a strong signal for this kind of job, not a universal ranking.
Sources
- TypeSafe, Introducing System One Models and Jev, 15 September 2026
- AI Engineer Guide, TypeSafe AI's Jev on Cloudflare AI Gateway, 18 September 2026
- Cloudflare, Jev model documentation, accessed 25 September 2026
- Z.ai, API pricing, accessed 25 September 2026
- Zhipu AI on Hugging Face, GLM-5.2 launch post, 17 June 2026
- Pydantic, TypeSafe models in Pydantic AI, accessed 25 September 2026
- Du et al., GLM: General Language Model Pretraining with Autoregressive Blank Infilling, ACL 2022
- Sean Goedecke, Jev means structured output is interesting again, 16 September 2026
- Cloudflare, GLM-5.3 on Workers AI, accessed 25 September 2026
- XY Space, Jev vs Tev benchmark results, 25 September 2026
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project