AI ComparisonsSep 23, 2026Updated Sep 25, 20269 min read

Jev vs GLM-5.2: It Lost Our Inbox Test, Then Won Once Rebuilt

Jev vs GLM-5.2 on our inbox: on GLM-5.2's prompt, Jev lost. Asked plain yes/no facts instead, it won 90.7% to 84.0% on 150 fresh items, at $0.0004 each.

On this page
  1. Is Jev better than GLM-5.2?
  2. At a glance: what each model does
  3. How they did on our tests
  4. Why the rebuild won: ask for facts, keep the policy in code
  5. Who should decide the items Jev is unsure about?
  6. What does each cost per month?
  7. Does the newer GLM-5.3 beat Jev on clean tasks?
  8. When is GLM-5.2 still the right choice?
  9. FAQ
  10. Sources

Jev beat GLM-5.2 at sorting our operations inbox, but only after we rebuilt the job. We ran the tests from 21 to 23 September 2026. On GLM-5.2's own prompt, Jev lost, 89.0% to 90.1%. Asked plain yes/no facts instead, it won 90.7% to 84.0% on 150 new items. Each item cost $0.0004 on Jev, against $0.0051.

Picking a label from a fixed list? Rebuild the job for Jev. Writing anything? Keep GLM-5.2.

  • Jev is a decision model from TypeSafe that answers questions and never writes text. It reached Cloudflare AI Gateway on 18 September 2026.
  • GLM-5.2 is a writing model from Zhipu AI's General Language Model (GLM) family. You can download it and run it yourself.
  • Rebuilt, Jev caught 48 of 50 client items; GLM-5.2 caught 42.
  • Speed: Jev answered in under a second per item, GLM-5.2 in several.
  • With Claude checking the items Jev was unsure about, the rebuilt setup caught all 50 client items.

Is Jev better than GLM-5.2?

Jev is better at choosing labels, once the task is shaped for it. Jev cannot do anything that needs written output, because it writes none.

Pick Jev if you sort, route or flag items against a fixed list and can restructure the prompt around it. Pick GLM-5.2 if you need summaries, extracted fields or written reasons. It also suits you if you want to host the model yourself or feed it long documents. Use both if your pipeline does some of each. That is what we run.

For a primer on how Jev differs from a chat model, see what is Jev.

At a glance: what each model does

Jev picks answers; GLM-5.2 writes them. Everything else in the table follows from that. A token is about three quarters of a word.

JevGLM-5.2
Who makes itTypeSafeZhipu AI (Z.ai)
What it doesAnswers questions you set: yes or no, pick one of up to 255 options, or a scoreReads, reasons and writes free text
Writes text?NoYes
Price per million tokens$0.042 to read, output free$1.40 to read, $4.40 to write
Speed70 to 500 milliseconds (TypeSafe's own figure); 0.3 seconds an item in our tests7 to 10 seconds an item in our tests
How much it reads at once32,000 tokens1 million tokens
Run it yourself?No: hosted API through TypeSafe early access or Cloudflare AI GatewayYes: open weights under the open-source MIT licence, 753 billion parameters

Jev's price and speed come from TypeSafe's launch post, its reading limit from Cloudflare's model page. The other column comes from Z.ai's pricing page and the GLM-5.2 launch post. TypeSafe's direct API is still early access with a waitlist.

How they did on our tests

On our own inbox, Jev lost a straight swap by about a point. Rebuilt, it won by 6.7 points on items neither setup had seen.

Items given the right label on our inbox
Items given the right label on our inbox
JevGLM-5.2
Same prompt, 303 build items89%90.1%
Rebuilt, 150 new items90.7%84%

Source: Our tests, 21 to 23 September 2026. Correct labels written by Claude against a written policy.

The task is our operations inbox. It labels every incoming item as client, internal, noise or unknown. Items include emails, calendar invites, call recordings, WhatsApp messages and AI coding sessions. It also tags a signal, such as a request, a commitment or money. Missing a client item is the costly mistake, so we count those separately.

Round one: Jev dropped into the old prompt

We gave Jev the instructions GLM-5.2 already used. Both ran on the 303 items we built the system with.

JevGLM-5.2
Items labelled right89.0%90.1%
Client items caught88.6%94.3%
Seconds per item0.287.4
Cost per item$0.00008$0.0051

Jev was cheaper and faster, and it lost on the number that matters most. A test on one shared prompt mostly measures which model follows that prompt better. This prompt was written for the incumbent.

Round two: Jev rebuilt, tested blind

We then rebuilt the job around Jev. Both models ran on 150 items held back from all the building work, 50 of them from clients.

Rebuilt JevGLM-5.2
Items labelled right90.7%84.0%
Client items caught (of 50)4842
Items wrongly called client00
Signal tagged right85 to 87%82.5%
Cost per item$0.0004$0.0051
Seconds per item0.37 to 10

Three limits apply. First, Claude wrote the correct labels for the 150 items, working from a written labelling policy. A human check of 20 of them is still pending. Second, GLM-5.2 changes its own label on 16% of re-runs, so its score moves between runs. Third, this is one inbox with one policy; a different job could land differently.

Why the rebuild won: ask for facts, keep the policy in code

The rebuild won because Jev answers small factual questions well. Our code then applies the rules that turn those facts into a label. It is the approach we recommend for any Jev job: every question in one call, no prose to parse, and low-confidence answers go to a person or a rule.

The first version asked Jev to read our policy and name the label. The rebuilt version asks three kinds of question in one call:

  • Who sent it, picked from a list of the 78 organisations we work with. A "new company" flag caught every sender not on it.
  • About 12 yes-or-no facts about the item. TypeSafe calls this question type a "noul".
  • Its own best guess at the label, used as one more input rather than the answer.
The rebuilt inbox triageEach incoming item gets one Jev call asking for the sender from a list of 78 organisations, about a dozen yes/no facts and a direct guess. Two small scoring models turn the answers into a label. Items Jev is confident about keep that label; unsure items and new senders go to Claude.0.8 OR ABOVEUNSURE OR NEWIncoming itememail, call, messageJEVJev, one callsender from 78 organisationsabout 12 yes/no facts + a guessTwo scoring modelstrained on 300 labelled itemsConfident?Label + signalClaude decidesLEGENDStepConfidence gateStronger model

Two small scoring models turn those answers into the label and the signal. We trained them on 300 labelled items. The gain shows most on the signal: 64% right when asked directly, 85 to 87% through the scoring model.

We also leave our own companies off the sender list. The rebuild also sends far more of each item, up to 36,000 characters. The old prompt was cut much shorter. How to use Jev covers the question design in detail.

Who should decide the items Jev is unsure about?

Claude should, with 90.2% of the unsure items right. GLM-5.2, as it runs in production, got 70.6%. Jev alone did better than that, at 76.5%.

Every answer comes back with a confidence score. About 69% of our traffic scores 0.8 or higher, and those answers were 98.1% right. The rest, plus anything from a new company, goes to a second decider. We tested six candidates:

Who decides the unsure itemsRight, unseen items (51)Right, build items (111)
GLM-5.2 as it runs in production70.6%65.8%
Jev on its own76.5%77.5%
GLM-5.2 given the organisation list and policy80.4%75.7%
gpt-oss-120bnot tested60.4%
GPT-5.4 mininot tested72.7% (22 items only)
Claude90.2%87.4%

With Claude on the unsure third, the whole system got 95.3% of the 150 unseen items right. It caught all 50 client items. The labelling caveat bites harder here, since Claude wrote the correct labels. A person still reviews what the system flags; see human-in-the-loop AI for how that fits. For a much larger labelling job, see Jev vs Claude.

What does each cost per month?

Jev costs about $1.50 a month for our inbox. GLM-5.2 costs about $18.

Our inbox sees about 120 items a day, and the monthly figures come from the per-item costs in round two. The Claude lane for unsure items is left out of the per-item figure.

At our volume, neither bill decides anything; the accuracy gap does. The cost gap matters at higher volumes. Jev pricing works through larger jobs and the billing details.

Does the newer GLM-5.3 beat Jev on clean tasks?

Yes, by a little. On 25 September 2026 we ran GLM-5.3, the next GLM release, against Jev on a fresh set of pick-one tasks. Each task had a short list of described options.

The newer model got 99.0% right and Jev 97.2%. GLM-5.3 cost 41 times as much per task. It also took about 6 times as long at the median, 2.9 seconds against 0.46.

Right answers on 400 new pick-one tasks
Right answers on 400 new pick-one tasks
AccuracyBoth halves of a pair right
Jev97.2%95%
GLM-5.399%98%

Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev

The inbox result above still stands, because the tests differ. The inbox win came from rebuilding the job around Jev: small yes/no facts, our organisation list, and scoring models in code. The benchmark asked both models the same single question, with a clear line describing each option.

When the task is that clean, the large model is slightly ahead. When the job is messy and you can reshape it, Jev can pull ahead.

You do not have to choose one. Jev answered every task, and GLM-5.3 re-answered only when Jev gave one of three answers it often got wrong. That mix got 98.1% right at $86 per million tasks, against 99.0% and $809 for GLM-5.3 alone. LLM cascade explains how we picked the three.

Claude wrote the 400 tasks, and 400 is a small set. GLM-5.3 ran with its default reasoning on. Decision model vs LLM covers when 1.8 points is worth 41 times the price, and Jev vs Tev has the full table.

When is GLM-5.2 still the right choice?

GLM-5.2 is the right choice for any step that has to write: pulling out fields, summarising, or explaining a decision. We kept extraction on it.

Jev cannot produce a name, a date string or a sentence. TypeSafe says so plainly: it gives up text generation. Pydantic AI's integration already plans for that split. It hands text fields to a second model.

Long input is the other case. The writing model reads up to a million tokens; Jev reads 32,000. GLM-5.2 is also the only one of the two you can download, which matters if your data cannot leave your servers. For the writing work, compare it with its successor in GLM-5.3 vs GLM-5.2.

FAQ

Is Jev a replacement for GLM-5.2?

Only for decisions from a fixed list, like routing or flagging. Jev does not write text, so summaries, extraction and written explanations still need a writing model. In our inbox, Jev now picks the labels and GLM-5.2 still pulls out the details. Treat Jev as a replacement for one step of a pipeline, and keep a writing model for the rest.

Why did Jev lose when we first swapped it in?

The prompt was written for the other model. On that shared prompt, Jev scored 89.0% to 90.1% and caught fewer client items. We then changed the questions to small facts and moved the labelling rules into code. After that, Jev won on unseen items, 90.7% to 84.0%. A straight swap mostly tests prompt-following, not what the model can do.

How much cheaper is Jev than GLM-5.2?

In our rebuilt setup, Jev cost $0.0004 per inbox item against $0.0051. At about 120 items a day, that is about $1.50 a month against $18. On list prices, Jev charges $0.042 per million tokens read, and output is free. GLM-5.2 charges $1.40 per million tokens read and $4.40 per million written.

Can I run Jev or GLM-5.2 on my own servers?

GLM-5.2, yes: its weights are published under a permissive open-source licence. Jev is a hosted service, available through TypeSafe's early access programme and Cloudflare AI Gateway. Cloudflare lists Jev with zero data retention, meaning the provider does not keep what you send.

How reliable are these test results?

Our results come from one inbox, tested 21 to 23 September 2026. Claude wrote the correct labels for the 150 unseen items against a written policy. A human check of 20 of them is pending. GLM-5.2 changes its own answer on 16% of re-runs. Treat the gaps as a strong signal for this kind of job, not a universal ranking.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…