GuidesSep 25, 20269 min read

LLM Cascade: Re-Ask 8% of Tasks, Get 98.1% at a Ninth of the Cost

An LLM cascade lets a cheap model answer first and re-asks a large one only on risky answers. Ours got 98.1% right for $86 per million tasks, against $809.

On this page
  1. What is an LLM cascade?
  2. How much does it save?
  3. Which answers get re-asked, and why those three?
  4. Why the cheaper first model made a worse cascade
  5. How to build your own cascade
  6. Honest limits
  7. FAQ
  8. Sources

An LLM cascade lets a cheap model answer every task and re-asks a larger model on a few. Together AI released Tev on 23 September 2026; two days later we tested it against Jev and GLM-5.3. Our best cascade put Jev first and re-asked GLM-5.3 on 8.2% of tasks. It got 98.1% right for $86 per million tasks. GLM-5.3 alone got 99.0% for $809.

  • The rule: Jev answers first. If its answer is one of three it often gets wrong, GLM-5.3 answers the same prompt and its answer is used.
  • Accuracy: 98.1% on tasks held out from tuning, against 97.2% for Jev alone and 99.0% for GLM-5.3 alone.
  • Cost: $86 per million tasks, 9.4 times cheaper than GLM-5.3 alone. Re-asked tasks pay for both calls.
  • Speed: a median of 479 milliseconds per task, because only 8.2% of tasks wait for the second model.
  • The surprise: starting with Tev, the cheaper model, made a worse cascade. Its best version got 97.1% and re-asked 39% of tasks.
  • The test: 400 new pick-one tasks, run on 25 September 2026. Every number here comes from our public results file.

What is an LLM cascade?

An LLM cascade is a setup where a cheap, fast model answers every task and a larger, dearer model re-answers only some of them. You pay the large model's price on a small share of the work, not all of it. In our test the cheap model was Jev, TypeSafe's decision model, and the large one was GLM-5.3, a chat model from Z.ai.

The idea has a name in the research. FrugalGPT, a 2023 research paper, called it an LLM cascade. It reported matching GPT-4 with up to 98% lower cost on its tasks. The same pattern also goes by model routing, or "cheap model first, then LLM".

The hard part is choosing which tasks to re-ask. There are two common ways:

  • By confidence. The cheap model reports how sure it is. Anything below a bar goes to the large model.
  • By answer. You look at which answer the cheap model gave. Some answers are ones it often gets wrong, so those go to the large model.

We have used both. Our operations inbox uses the confidence version. This post measures the answer version, because it needs nothing from the cheap model except its answer.

One Jev callYour code sends Jev one piece of text and a set of typed questions. Jev returns an answer to each with a probability, writing no text. Your code acts on the answers it is sure of and hands the unsure ones to a person or a larger model.Text + yes/no, choice,score questionsOne answer each + probabilitiesAct on the confident answersHand over the unsure onesDecisionYour codeJevPerson or larger modelLEGENDCallResponse

How much does it save?

The cascade cost $86 per million tasks, against $809.08 for GLM-5.3 alone. That is 89% less for 0.9 points of accuracy.

We tested on 400 pick-one tasks. Each gives the model an input, a question and three to six options, and the model returns one option. Examples include ruling on a return request, routing an agent to a tool and grading a bug report. The four setups below run on the same tasks. The cascade rows are scored only on tasks their rule was not built from.

Cost per million tasks
Cost per million tasks
Cost
Jev only$20
Cascade (recommended)$86
Cascade, re-ask every answer Jev ever missed$161
GLM-5.3 only$809

Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev

Right answers on 400 new tasks
Right answers on 400 new tasks
Accuracy
Jev only97.2%
Cascade (recommended)98.1%
Cascade, re-ask every answer Jev ever missed98.5%
GLM-5.3 only99%

Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev

The arithmetic is simple. A Jev answer costs $0.0000196. A GLM-5.3 answer costs $0.0008091, and the cascade buys one for 8.2% of tasks. That adds $0.0000663, for $0.0000860 a task, or $86 a million.

The wider cascade re-asks 17.5% of tasks. It gains 0.4 points and nearly doubles the bill. Both prices are list rates: Jev and Tev charge $0.042 per million input tokens with output free, and GLM-5.3 lists at $1.40 in and $4.40 out on Cloudflare Workers AI. A token is about three quarters of a word.

Which answers get re-asked, and why those three?

Three of Jev's answers get re-asked: respond_directly, approve_store_credit and approve_full_refund. Each passed two tests. Jev was right less than 90% of the times it gave that answer, and GLM-5.3 did better on those same tasks.

Precision means how often a model is right when it gives a particular answer. It is the number that tells you whether to trust an answer you are holding.

Task familyJev's answerTimes Jev gave itJev's precisionGLM-5.3 right on those tasks
Agent routingrespond_directly366.7%100.0%
Return requestsapprove_store_credit875.0%100.0%
Return requestsapprove_full_refund2588.0%100.0%

Two of the three are return rulings. Jev scored 90% on return requests, its weakest family, while GLM-5.3 got every one right.

The second test kept two answers off the list. Jev also gave frontend (bug triage) and remove_harassment (moderation) with under 90% precision. Neither made the list, because GLM-5.3 did no better on those tasks. Re-asking them would add cost and no right answers.

The routing rule reads only Jev's answer, never the correct one, so it works on live traffic. A plain lookup table decides.

Why the cheaper first model made a worse cascade

Tev made a worse first model, even though it costs about half as much per task. With Tev first, the best cascade got 97.1% while re-asking 39% of tasks, for $326 per million. Jev alone scored 97.2% for $20.

The reason is where each model's mistakes land. A cascade by answer works when the mistakes cluster in a few answers. Jev's did. Tev's spread out: it gave 17 different answers with precision under 90%, in seven of the eight task families. Jev gave five.

Some of Tev's weak answers, from the per-answer tables:

  • `require_human_review`, 61%. Tev sends safe agent actions to a person. It errs toward caution.
  • `deny_outside_window`, 64%. Tev declines returns as too late when a different ruling was right.
  • `remove_scam`, 60%, and `escalate_self_harm`, 50%, in moderation.
  • `negative`, 67%, in review sentiment.

To catch those, the cascade has to re-ask a large share of Tev's traffic. Each re-asked task pays for both calls, so Tev's lower price is gone long before the accuracy catches up. The first model's price matters less than how neatly its mistakes group. Jev vs Tev has the full head to head.

How to build your own cascade

Build it from your own labelled data. Our list of three answers fits our 400 tasks, not yours. You need a few hundred labelled examples and both models' answers on them.

  1. Label a few hundred real examples. Write the correct answer for each. Include the rare answers, since those are where the cheap model slips. Our guide to testing a classifier covers writing test items in pairs that differ by one edit.
  2. Run both models on all of them. Use the prompts you will ship. Record every answer, the cost and the time.
  3. Work out the cheap model's precision for each answer. For each answer it gave, count how often it was right.
  4. Pick the risky list on part of the data and score it on the rest. We built the list on four fifths of the task pairs and scored the other fifth, five times over, on three shuffles. Keep the setup that is cheapest within the accuracy you need.
  5. Price the re-asked tasks at both calls. The cheap call has already happened when you decide to re-ask.
  6. Re-check as your data changes. New products, new policies and new phrasing move precision. Re-run the labelled set on a schedule.

The confidence version follows the same steps with a confidence bar in place of the list. How to use Jev shows where to set that bar. Our inbox uses it: Jev keeps the items it is sure of, and Claude decides the rest. Jev vs GLM-5.2 has those results.

Honest limits

Four limits apply to these numbers:

  • The 98.1% is slightly optimistic. We tried 20 cascade setups and picked the best one using the same results. The rule itself was built and scored on separate tasks, but the choice of threshold was not.
  • The set is small. 400 tasks means wide error bars. Some answers appear only a handful of times. respond_directly was Jev's answer three times.
  • Claude wrote the tasks, and the same kind of model checked the labels, with spot checks. None came from public datasets.
  • The list is specific to these task families. Yours will differ. That is why step 1 above exists.

Latency has a cost too. A re-asked task waits for Jev and then GLM-5.3. At the two medians that adds up to about 3.3 seconds, against 459 milliseconds for Jev alone. The cascade's median stays at 479 milliseconds because only 8.2% of tasks wait. Jev ran through the AI Space gateway and GLM-5.3 through AI Space with reasoning at its defaults, so both times include that extra hop.

For when a large model alone is worth the price, see decision model vs LLM. For comparing models by cost per right answer, see the cheapest model per solved task. For what a decision model is, see what is Jev.

FAQ

What is the difference between an LLM cascade and model routing?

The words overlap. A cascade usually runs the cheap model first and escalates after seeing its answer. Routing can also mean choosing a model before any call, from the task itself. Our setup is a cascade: Jev always answers, and its answer decides whether GLM-5.3 is asked too.

Does an LLM cascade always save money?

No. It saves money when the cheap model is right most of the time and its mistakes cluster in answers you can spot. Our Tev-first cascade re-asked 39% of tasks and cost $326 per million. That was worse than Jev alone on both accuracy and price.

Should I route by confidence or by answer?

Route by answer when the cheap model's mistakes cluster in a few answers. Route by confidence when it reports a trustworthy confidence and mistakes are spread out. You can test both on the same labelled set. Our inbox routes by confidence; this benchmark routes by answer.

How many examples do I need to build a cascade?

A few hundred labelled examples is a practical start. We used 400. Rare answers need enough examples to measure precision, so collect more of those. Score the rule on examples it was not built from.

Is the large model's answer always right when it is re-asked?

No. On our three risky answers, GLM-5.3 got every re-asked task right. Across all 400 tasks it scored 99.0%. Keep a person on the decisions where a wrong answer is costly.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…