On this page
An LLM cascade lets a cheap model answer every task and re-asks a larger model on a few. Together AI released Tev on 23 September 2026; two days later we tested it against Jev and GLM-5.3. Our best cascade put Jev first and re-asked GLM-5.3 on 8.2% of tasks. It got 98.1% right for $86 per million tasks. GLM-5.3 alone got 99.0% for $809.
- The rule: Jev answers first. If its answer is one of three it often gets wrong, GLM-5.3 answers the same prompt and its answer is used.
- Accuracy: 98.1% on tasks held out from tuning, against 97.2% for Jev alone and 99.0% for GLM-5.3 alone.
- Cost: $86 per million tasks, 9.4 times cheaper than GLM-5.3 alone. Re-asked tasks pay for both calls.
- Speed: a median of 479 milliseconds per task, because only 8.2% of tasks wait for the second model.
- The surprise: starting with Tev, the cheaper model, made a worse cascade. Its best version got 97.1% and re-asked 39% of tasks.
- The test: 400 new pick-one tasks, run on 25 September 2026. Every number here comes from our public results file.
What is an LLM cascade?
An LLM cascade is a setup where a cheap, fast model answers every task and a larger, dearer model re-answers only some of them. You pay the large model's price on a small share of the work, not all of it. In our test the cheap model was Jev, TypeSafe's decision model, and the large one was GLM-5.3, a chat model from Z.ai.
The idea has a name in the research. FrugalGPT, a 2023 research paper, called it an LLM cascade. It reported matching GPT-4 with up to 98% lower cost on its tasks. The same pattern also goes by model routing, or "cheap model first, then LLM".
The hard part is choosing which tasks to re-ask. There are two common ways:
- By confidence. The cheap model reports how sure it is. Anything below a bar goes to the large model.
- By answer. You look at which answer the cheap model gave. Some answers are ones it often gets wrong, so those go to the large model.
We have used both. Our operations inbox uses the confidence version. This post measures the answer version, because it needs nothing from the cheap model except its answer.
How much does it save?
The cascade cost $86 per million tasks, against $809.08 for GLM-5.3 alone. That is 89% less for 0.9 points of accuracy.
We tested on 400 pick-one tasks. Each gives the model an input, a question and three to six options, and the model returns one option. Examples include ruling on a return request, routing an agent to a tool and grading a bug report. The four setups below run on the same tasks. The cascade rows are scored only on tasks their rule was not built from.
| Cost | |
|---|---|
| Jev only | $20 |
| Cascade (recommended) | $86 |
| Cascade, re-ask every answer Jev ever missed | $161 |
| GLM-5.3 only | $809 |
Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
| Accuracy | |
|---|---|
| Jev only | 97.2% |
| Cascade (recommended) | 98.1% |
| Cascade, re-ask every answer Jev ever missed | 98.5% |
| GLM-5.3 only | 99% |
Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
The arithmetic is simple. A Jev answer costs $0.0000196. A GLM-5.3 answer costs $0.0008091, and the cascade buys one for 8.2% of tasks. That adds $0.0000663, for $0.0000860 a task, or $86 a million.
The wider cascade re-asks 17.5% of tasks. It gains 0.4 points and nearly doubles the bill. Both prices are list rates: Jev and Tev charge $0.042 per million input tokens with output free, and GLM-5.3 lists at $1.40 in and $4.40 out on Cloudflare Workers AI. A token is about three quarters of a word.
Which answers get re-asked, and why those three?
Three of Jev's answers get re-asked: respond_directly, approve_store_credit and approve_full_refund. Each passed two tests. Jev was right less than 90% of the times it gave that answer, and GLM-5.3 did better on those same tasks.
Precision means how often a model is right when it gives a particular answer. It is the number that tells you whether to trust an answer you are holding.
| Task family | Jev's answer | Times Jev gave it | Jev's precision | GLM-5.3 right on those tasks |
|---|---|---|---|---|
| Agent routing | respond_directly | 3 | 66.7% | 100.0% |
| Return requests | approve_store_credit | 8 | 75.0% | 100.0% |
| Return requests | approve_full_refund | 25 | 88.0% | 100.0% |
Two of the three are return rulings. Jev scored 90% on return requests, its weakest family, while GLM-5.3 got every one right.
The second test kept two answers off the list. Jev also gave frontend (bug triage) and remove_harassment (moderation) with under 90% precision. Neither made the list, because GLM-5.3 did no better on those tasks. Re-asking them would add cost and no right answers.
The routing rule reads only Jev's answer, never the correct one, so it works on live traffic. A plain lookup table decides.
Why the cheaper first model made a worse cascade
Tev made a worse first model, even though it costs about half as much per task. With Tev first, the best cascade got 97.1% while re-asking 39% of tasks, for $326 per million. Jev alone scored 97.2% for $20.
The reason is where each model's mistakes land. A cascade by answer works when the mistakes cluster in a few answers. Jev's did. Tev's spread out: it gave 17 different answers with precision under 90%, in seven of the eight task families. Jev gave five.
Some of Tev's weak answers, from the per-answer tables:
- `require_human_review`, 61%. Tev sends safe agent actions to a person. It errs toward caution.
- `deny_outside_window`, 64%. Tev declines returns as too late when a different ruling was right.
- `remove_scam`, 60%, and `escalate_self_harm`, 50%, in moderation.
- `negative`, 67%, in review sentiment.
To catch those, the cascade has to re-ask a large share of Tev's traffic. Each re-asked task pays for both calls, so Tev's lower price is gone long before the accuracy catches up. The first model's price matters less than how neatly its mistakes group. Jev vs Tev has the full head to head.
How to build your own cascade
Build it from your own labelled data. Our list of three answers fits our 400 tasks, not yours. You need a few hundred labelled examples and both models' answers on them.
- Label a few hundred real examples. Write the correct answer for each. Include the rare answers, since those are where the cheap model slips. Our guide to testing a classifier covers writing test items in pairs that differ by one edit.
- Run both models on all of them. Use the prompts you will ship. Record every answer, the cost and the time.
- Work out the cheap model's precision for each answer. For each answer it gave, count how often it was right.
- Pick the risky list on part of the data and score it on the rest. We built the list on four fifths of the task pairs and scored the other fifth, five times over, on three shuffles. Keep the setup that is cheapest within the accuracy you need.
- Price the re-asked tasks at both calls. The cheap call has already happened when you decide to re-ask.
- Re-check as your data changes. New products, new policies and new phrasing move precision. Re-run the labelled set on a schedule.
The confidence version follows the same steps with a confidence bar in place of the list. How to use Jev shows where to set that bar. Our inbox uses it: Jev keeps the items it is sure of, and Claude decides the rest. Jev vs GLM-5.2 has those results.
Honest limits
Four limits apply to these numbers:
- The 98.1% is slightly optimistic. We tried 20 cascade setups and picked the best one using the same results. The rule itself was built and scored on separate tasks, but the choice of threshold was not.
- The set is small. 400 tasks means wide error bars. Some answers appear only a handful of times.
respond_directlywas Jev's answer three times. - Claude wrote the tasks, and the same kind of model checked the labels, with spot checks. None came from public datasets.
- The list is specific to these task families. Yours will differ. That is why step 1 above exists.
Latency has a cost too. A re-asked task waits for Jev and then GLM-5.3. At the two medians that adds up to about 3.3 seconds, against 459 milliseconds for Jev alone. The cascade's median stays at 479 milliseconds because only 8.2% of tasks wait. Jev ran through the AI Space gateway and GLM-5.3 through AI Space with reasoning at its defaults, so both times include that extra hop.
For when a large model alone is worth the price, see decision model vs LLM. For comparing models by cost per right answer, see the cheapest model per solved task. For what a decision model is, see what is Jev.
FAQ
What is the difference between an LLM cascade and model routing?
The words overlap. A cascade usually runs the cheap model first and escalates after seeing its answer. Routing can also mean choosing a model before any call, from the task itself. Our setup is a cascade: Jev always answers, and its answer decides whether GLM-5.3 is asked too.
Does an LLM cascade always save money?
No. It saves money when the cheap model is right most of the time and its mistakes cluster in answers you can spot. Our Tev-first cascade re-asked 39% of tasks and cost $326 per million. That was worse than Jev alone on both accuracy and price.
Should I route by confidence or by answer?
Route by answer when the cheap model's mistakes cluster in a few answers. Route by confidence when it reports a trustworthy confidence and mistakes are spread out. You can test both on the same labelled set. Our inbox routes by confidence; this benchmark routes by answer.
How many examples do I need to build a cascade?
A few hundred labelled examples is a practical start. We used 400. Rare answers need enough examples to measure precision, so collect more of those. Score the rule on examples it was not built from.
Is the large model's answer always right when it is re-asked?
No. On our three risky answers, GLM-5.3 got every re-asked task right. Across all 400 tasks it scored 99.0%. Keep a person on the decisions where a wrong answer is costly.
Sources
- Together AI, How to train your own Jev for $17, 23 September 2026
- TypeSafe, Introducing System One Models and Jev, 15 September 2026
- Cloudflare, GLM-5.3 on Workers AI, accessed 25 September 2026
- Chen, Zaharia and Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, 9 May 2023
- XY Space, Jev vs Tev benchmark results, 25 September 2026
- XY Space, Jev vs Tev benchmark methodology, 25 September 2026
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project