GuidesSep 20, 2026Updated Sep 25, 20268 min read

Jev Pricing: Our $26.63 Bill for Sorting 143,891 Skills

Jev pricing is $0.042 per million input tokens with free output. We sorted 143,891 skills for $26.63. Here is the bill, the estimate, and when it matters.

On this page
  1. How much does Jev cost?
  2. What did a real job cost?
  3. Jev vs an LLM for the same job
  4. What a million decisions cost
  5. Why the bill came in under our estimate
  6. How to estimate a Jev bill before you run it
  7. How billing works on Cloudflare AI Gateway
  8. When the saving does not matter
  9. FAQ
  10. Sources

Jev costs $0.042 per million tokens of input (about 750,000 words), and its answers are free. That was its price when it launched on Cloudflare AI Gateway on 18 September 2026. We used it to sort 143,891 AI agent skills into categories, and the bill was $26.63. By our estimate, Claude Opus 5 would charge about $11,600 for the same work.

  • Who makes it: TypeSafe, whose own API has a waitlist. Cloudflare AI Gateway, a service that routes calls to many AI models, sells it too.
  • Our bills: we ran every test from 21 to 23 September 2026, and the bills came in below our own estimates.
  • Per item: our inbox pays $0.0004 an item, against $0.0051 on the model it replaced.
  • Before you run: a trial on 300 items predicted the full cost almost exactly.
  • Small jobs: at 120 items a day the saving is about $16 a month, so the case for Jev there is speed and accuracy.

How much does Jev cost?

You pay $0.042 per million input tokens, and nothing for output. TypeSafe's launch post of 15 September 2026 puts the same price another way: $42 per billion tokens, with output "too cheap to meter".

A token is a chunk of text, about three quarters of a word. A million input tokens is roughly 750,000 words of text sent to the model, for a little over four cents.

Output is free because Jev does not write text. It answers typed questions about a piece of text you send it, which TypeSafe calls the "state". There are three question types: a yes/no question (TypeSafe calls these "noul"), a choice from up to 255 options, and a score on a scale. Each answer comes back with probabilities and a confidence figure, per Cloudflare's model page. Our pillar post, what is Jev, covers how the model works.

Two limits shape the bill:

  • 32,000 tokens of context. That is the most text Jev reads in one call, per Cloudflare's model page.
  • One state per call. Each call judges one piece of text. We found no way to send 25 items in one request, so every item pays for its own copy of your questions and options.

What did a real job cost?

A full pass over our Skill Atlas cost $26.63 on the bill. The atlas is a catalogue of 143,891 AI agent skills sorted into 276 categories. Our own count of the tokens sent put the cost higher, at $30.97.

We ran the atlas pass and the tests below from 21 to 23 September 2026, all through Cloudflare. The method:

  • Two calls per skill. A choice question takes at most 255 options, and we have 276 categories. So the first call picks one of 58 teams, and the second picks a category inside that team (21 options at most).
  • About 4,400 input tokens per skill. Each skill's description plus the option lists.
  • Median answer time 0.67 seconds, with 95% of calls done inside 2.9 seconds.
  • Zero answers outside the list. Every answer was a category we offered.

The run took about 17 hours against the 9 we planned. Rate throttling, bursts of refused calls, and running out of prepaid credit account for the gap. Speed is covered in how to use Jev.

The second job is our operations inbox. It reads every incoming email, calendar invite, call recording, WhatsApp message and AI coding session. It labels each one as client, internal, noise or unknown, and tags a signal such as a request, a commitment or money.

The inbox used to run on GLM-5.2, a large chat model from Zhipu. That cost $0.0051 per item. The rebuilt version asks Jev about a dozen yes/no questions per item and sends it up to 36,000 characters of text. It costs $0.0004 per item. The accuracy results are in Jev vs GLM-5.2.

Jev vs an LLM for the same job

An ordinary AI chat model costs far more. By our estimate, Nemotron 3 120B would charge about $1,065 for the same atlas pass. Claude Opus 5 would charge about $11,600. Those models, called LLMs (large language models), bill for every word they read and every word they write back.

One pass over 143,891 skills
One pass over 143,891 skills
Cost of the pass
Jev, billed$26.63
Nemotron 3 120B, est.$1,065
Claude Opus 5, est.$11,600

Source: Jev from Cloudflare's bill; the others are our estimates at the rates we measured, one skill per request. September 2026.

JobJevAlternativeAlternative cost
Full atlas pass, 143,891 skills$26.63 billedNemotron 3 120Babout $1,065 (estimate)
Full atlas pass, 143,891 skills$26.63 billedClaude Opus 5, one skill per requestabout $11,600 (estimate)
Inbox triage, per item$0.0004GLM-5.2$0.0051

The Opus figure assumes one skill per request. In production we send Claude about 25 skills in each prompt, which cuts its input roughly 20 times. Jev cannot group items that way, so the fair gap is smaller than the table shows.

Cost only means something next to accuracy. On a 60-skill hard set run through the same gateway, here is how often each model's top answer matched the category Claude had assigned:

ModelMatched Claude's categoryCost per 1,000 skillsMedian time
Claude Opus 5 (11 skills only)63.6%$80.902.1s
Nemotron 3 120B52.3%$7.407.6s
gpt-oss-120b40.4%$3.746.0s
Jev38.3%$0.180.92s
Claude Sonnet 5 (15 skills only)40.0%$32.622.1s
Qwen3 30B33.3%$0.804.4s
Llama 3.3 70B27.7%$3.424.1s

Three caveats apply. The score is agreement with Claude's labels, not a verified right answer. The two Claude runs stopped early, too few skills to rank them. And the set was built from hard cases, so every figure is lower than it would be on typical skills. The full picture is in Jev vs Claude. For the idea of pricing by finished work rather than by token, see the cheapest model per solved task.

What a million decisions cost

On short pick-one tasks, a million Jev decisions cost about $19.64 at list price. We measured this on 25 September 2026 across 400 new tasks. Each task is a paragraph of text and a short list of options. Tev, Together AI's open model, would do the same million for $10.33. Two large chat models would charge $809 and $1,708.

List-price cost of 1 million pick-one tasks
List-price cost of 1 million pick-one tasks
Cost per million tasks
Tev$10.33
Jev$19.64
Jev, then GLM-5.3 on 3 answers$86
GLM-5.3$809
Claude Opus 5.5$1,708

Source: Our Jev vs Tev benchmark, 400 tasks, 25 September 2026, github.com/choyiny/jev-vs-tev

Jev and Tev have the same list price, $0.042 per million input tokens with free output. The gap comes from counting. For the same text, Jev's API billed 468 input tokens per task and Tev's billed 246. TypeSafe probably wraps the input in its own prompt.

That is a different gap from the one in the next section. There, Cloudflare billed less than our own count of the same Jev calls. Here, two APIs count the same text differently.

Accuracy moves the other way. Jev got 97.2% of the tasks right and Tev 90.0%. Both chat models got about 99%. The chat models ran with their default reasoning on, and reasoning tokens count as output. Decision model vs LLM works through when those last two points are worth 41 to 87 times the price, and Jev vs Tev compares the two small models. Claude wrote these 400 tasks, and 400 is a small set.

You can also pay for the chat model only some of the time. Jev answers every task, and GLM-5.3 re-answers only when Jev gives one of three answers it often gets wrong. That mix got 98.1% right for $86 per million tasks. GLM-5.3 alone got 99.0% for $809.08. Only 8% of tasks paid for the second call. LLM cascade shows how we chose the three answers.

Why the bill came in under our estimate

Every run we checked billed less than list price times our token count. The atlas pass counted out to $30.97 and billed $26.63. On one benchmark run the gap was wider: our estimate said $23.07 and the bill said $9.29.

Our token count against Cloudflare's bill
Our token count against Cloudflare's bill
Our token countCloudflare's bill
Skill Atlas pass$30.97$26.63
One benchmark run$23.07$9.29

Source: Our runs, 21 to 23 September 2026

We have not found out why the billed tokens run lower. The practical rule is to read spend from Cloudflare, not from your own arithmetic. We pull it from two GraphQL datasets in the Cloudflare analytics API: aiGatewayRequestsAdaptiveGroups for calls routed through AI Gateway, and aiInferenceAdaptiveGroups for Workers AI models.

How to estimate a Jev bill before you run it

Run a few hundred items first, then scale the cost up by count. Before the atlas pass we ran 300 skills. That trial projected $31.00 for the full catalogue. Our token count for the real run came to $30.97, within 0.1%.

Two things make this work. Input per item is steady, because the questions and options repeat on every call and only the item text changes. And there is no output to vary. Pick trial items at random from the real set, so long and short items, and every language, turn up in the same mix.

Time is harder to predict than money. Our trial ran at about 4.5 calls a second with 8 calls in flight. The full run sustained 2.5 to 4 a second. Raising the calls in flight to 24 gave 3.1 a second and 8% failures. Plan time from the sustained rate, not the trial.

How billing works on Cloudflare AI Gateway

On Cloudflare AI Gateway, Jev counts as a third-party model, so you pay from prepaid credit, not from your usual invoice. This is what we observed between 21 and 23 September 2026.

  • Prepaid credit. You load credit in the dashboard. Cloudflare's unified billing docs state a 5% fee on every top-up: a $100 credit costs $105.
  • One pot for all third-party models. The same credit pays for Jev and for any other outside provider you call through the gateway.
  • Workers AI bills separately. Models Cloudflare runs itself on Workers AI go on your normal account billing, not the credit.
  • No balance API that we could find. When credit runs out, calls fail with a payment error. That is one reason our atlas run took 17 hours. Turn on auto top-up, which the docs describe as a threshold plus a recharge amount.

Payment errors also appeared for another reason. Some items that contained code, URLs or shell commands failed every retry with the same error, and passed once we replaced that text with a short stand-in label. Short random bursts of the same error cleared on retry.

Jev is also listed in AI Space's model catalog.

When the saving does not matter

At small volumes, price is not the reason to switch. Our operations inbox handles about 120 items a day. That is about $1.50 a month on Jev and $18 a month on the old model. Sixteen dollars is not a budget decision.

The case at that volume is accuracy and speed. On a 150-item test set held back from tuning, the rebuilt inbox labelled 90.7% of items correctly. The old model managed 84.0%. Answers came back in about 0.3 seconds, against 7 to 10 for the old model. The right answers in that test set were written by Claude against a written labelling policy. That may favour answers that agree with Claude, and a human check of 20 items is still pending. GLM-5.2 also changes its own label on 16% of re-runs, which makes its score harder to pin down.

The price does matter when volume is high and the job is a choice. Sorting a whole catalogue, screening every inbound message, or scoring a large archive are all cases where a $26 bill replaces a four-figure one. If your job needs the model to write text back, Jev cannot do it. Pydantic AI's TypeSafe integration hands text fields to an ordinary LLM for exactly that reason.

FAQ

Is Jev free?

No. Output is free, but input costs $0.042 per million tokens, per TypeSafe's launch post of 15 September 2026. Input is the text you send plus your questions and options, so you pay for every call. Our largest job, sorting 143,891 skills in two calls each, billed $26.63. Small jobs cost fractions of a cent.

How much does Jev cost per request?

The cost per request depends on how much text you send. Our Skill Atlas used about 4,400 input tokens per skill across two calls. Our operations inbox sends more text and about a dozen questions per item, and costs $0.0004 per item. Run a few hundred items first and scale up; our 300-skill trial predicted the full cost within 0.1%.

Can I batch several items into one Jev call?

We found no way to. Each call takes one piece of text, which TypeSafe calls the state, and answers questions about it. That means your questions and options are paid for on every item. LLMs can take 25 items in one prompt, which cuts their input cost a lot, so compare Jev against a batched LLM, not an unbatched one.

Is Jev cheaper than GLM-5.2?

On our operations inbox, yes. It costs $0.0004 per item against $0.0051, and it answers far faster. At about 120 items a day that is $1.50 a month against $18. The accuracy comparison, and its caveats, are in Jev vs GLM-5.2.

Where can I buy Jev?

TypeSafe's own API is in early access with a waitlist. It has been live on Cloudflare AI Gateway since 18 September 2026 as typesafe/jev. There you pay from prepaid credit, with a 5% fee on each top-up. We ran all our tests through Cloudflare.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…