AI ComparisonsSep 25, 20269 min read

Jev vs Claude: 143,891 Skills for $27, and When to Trust It

Jev vs Claude on 143,891 skills: Jev cost $27 and matched Claude 46.4% overall, but 93.6% when at least 90% sure. Use Jev first, Claude for the rest.

On this page
  1. Is Jev as accurate as Claude?
  2. At a glance
  3. How they did on our tests
  4. Where do Jev and Claude disagree, and why?
  5. How does Jev compare with other models on a hard set?
  6. What about clean, single-question tasks?
  7. What about non-English text?
  8. Using both: Jev as a second labeller
  9. When not to use Jev
  10. FAQ
  11. Sources

Jev vs Claude is a fast labeller against a model that reasons. From 21 to 23 September 2026 we checked Jev against Claude's labels on our Skill Atlas. The atlas is a catalogue of 143,891 AI agent skills sorted into 276 categories. Jev's bill was $26.63, and it matched Claude's category 46.4% of the time. When Jev was at least 90% sure, it matched 93.6%.

  • What Jev is: a model from TypeSafe, launched 15 September 2026, that picks answers from a list you give it. It cannot write text (TypeSafe).
  • What Claude is: a general model that reads the whole item, reasons about it and can explain its choice in words.
  • Cost of our full pass with Jev: $30.97 by token count, $26.63 on the actual bill.
  • Answers outside the 276 categories: zero.
  • The useful number: Jev's confidence tracks whether it agrees with Claude. That makes it a good first pass.

Is Jev as accurate as Claude?

No, not item for item, but it knows when it is right. If you need one label per item with no review, use Claude. If you have a large pile and a budget, run Jev first, keep what it is sure about, and send the rest to Claude. On our atlas Jev agreed with Claude on 46.4% of categories overall, and on 93.6% of the skills where its confidence was 0.9 or higher.

Jev is weaker as a single labeller and stronger as a filter. A filter that costs $27 for 143,891 items changes what is worth labelling at all.

At a glance

Jev and Claude answer different kinds of request. Jev picks from options and returns a confidence score for each. Claude reads, reasons and writes.

JevClaude
Writes text?No. Picks from up to 255 options per questionYes
Cost for our full atlas$30.97 by tokens ($26.63 billed)About $11,600 at Claude Opus 5 prices, one skill per request (our estimate)
Speed per item0.67 seconds median, 2.9 seconds at the slow end2.1 seconds on our hard test (Claude Opus 5 and Claude Sonnet 5)
How it answersA choice plus a confidence from 0 to 1A label, and reasons if you ask
Where to get itCloudflare AI Gateway today; TypeSafe's own access is early access with a waitlistAnthropic's API and most major clouds

Jev's price is $0.042 per million input tokens, and output is free (Cloudflare AI Gateway write-up). Tokens are chunks of text, about three quarters of a word each. Each skill we sent was about 4,400 tokens. The Jev pricing post covers what that means for other workloads.

The 255-option limit mattered for us. Our list has 276 categories, so we grouped them into 58 teams of at most 21 categories. Jev picked a team first, then a category inside it: two calls per skill.

How they did on our tests

On the 91,306 atlas skills Claude labelled itself, Jev matched Claude on the exact category 46.4% of the time, the team 54.4% and the department 66.9%. When Claude's second choice counts too, category agreement rises to 56.9%.

All of these figures measure agreement with Claude's labels, not accuracy against a human answer key. Claude can be wrong, and on hard items it does not always agree with itself. Re-labelling the same hard skills, Claude Opus 5 matched the original label roughly 60 to 70% of the time. That run stopped early.

Agreement also depended on how sure Claude was. Where Claude marked its own label high confidence, Jev agreed 62.3% of the time. At medium confidence, 36.5%. At low, 19.5%. The items Jev missed were often items Claude found hard too.

Jev's confidence against agreement with Claude

The table below covers the 91,306 skills where we had both Jev's confidence and Claude's label, run 21 to 23 September 2026.

Keep only Jev's confident answers, and agreement with Claude climbs
Keep only Jev's confident answers, and agreement with Claude climbs
Share of skills keptSame category as Claude
Any100%46.4%
0.5+51.6%66.1%
0.6+38.4%73.9%
0.7+27.6%81.7%
0.8+18.6%88.4%
0.9+10.4%93.6%

Source: Our Skill Atlas run, 91,306 skills, 21 to 23 September 2026

Keep only answers where Jev's confidence is at leastShare of skills keptSame category as ClaudeSame team as Claude
Any100%46.4%54.4%
0.551.6%66.1%71.3%
0.638.4%73.9%78.1%
0.727.6%81.7%84.8%
0.818.6%88.4%90.7%
0.910.4%93.6%95.3%

Read the 0.8 row as: Jev settles almost a fifth of the pile, and on that fifth it agrees with Claude nearly nine times in ten. The other four fifths go to Claude.

Method and caveats

  • Dates: 21 to 23 September 2026. Numbers re-checked on 23 September.
  • Task: assign each skill to one of 276 categories, grouped into 58 teams and a smaller set of departments.
  • Reference: Claude's labels, not a human answer key.
  • Run time: about 17 hours against 9 planned. Throttling, bursts of payment errors from the gateway and an overnight pause slowed it.
  • Estimate check: a 300-skill trial projected $31.00 for the full pass. The token count came to $30.97.

Where do Jev and Claude disagree, and why?

Most of the department-level gap came from one department, and the cause was our instruction. Jev placed 48.9% of skills in Skill and Agent Engineering. Claude placed 22.6% there. That one department made up 85% of the department-level disagreements.

Jev put 13,519 skills into a category for skill creation and templates. Among them were a profile builder for real estate investment trusts, an investment thesis tracker and a restaurant finder. Our instructions say skills about writing, installing, evaluating or orchestrating other skills belong in Skill and Agent Engineering, whatever domain they mention. Claude applied that rule only to skills about skills. Jev applied it far more widely, and filed these by their format instead of what they do: property investing, portfolio research, dining.

Jev reads an instruction the way it is written. Claude fills in what the writer meant. The fix is to ask two questions instead of one:

  1. What task does this skill do for its user? Pick the category from that.
  2. Separately, a yes or no: is this skill about building other skills or agents?

TypeSafe calls yes or no questions "noul" (TypeSafe docs). Splitting the task this way keeps a literal reader from folding the topic question into the format question. We have not tested the two-question version yet. A rough correction to the probabilities after the run already lifted department agreement from 66.9% to 73.3%. The how to use Jev post goes deeper on phrasing questions for it.

How does Jev compare with other models on a hard set?

Jev landed in the middle of the open models on our hardest skills, at a small fraction of their cost and time. We picked 60 skills that are hard to place, 47 of them in English, and ran each model against Claude's labels.

ModelMatched Claude's categorySame teamCost per 1,000 skillsMedian seconds per skill
Claude Opus 5 (11 items only)63.6%63.6%$80.902.1
Nemotron 3 120B52.3%56.8%$7.407.6
gpt-oss-120b40.4%48.9%$3.746.0
Claude Sonnet 5 (15 items only)40.0%40.0%$32.622.1
Jev38.3%42.6%$0.180.92
Qwen3 30B33.3%42.9%$0.804.4
gpt-oss-20b29.8%36.2%$2.063.3
Llama 3.3 70B27.7%38.3%$3.424.1

Nemotron 3 120B was the strongest open model, at 41 times Jev's cost per skill. Over the full atlas, the same pass with Nemotron would have cost about $1,065 by our estimate.

The Claude rows are too small to rank. Both Claude runs stopped early, at 11 and 15 items.

Format failures matter as much as the scores. Llama 3.3 70B invented six categories that are not on the list. Qwen3 30B sometimes answered with the skill's own name. Jev cannot do either, because it can only return an option it was given.

What about clean, single-question tasks?

On clean tasks with a short, described option list, the gap nearly closes. On 25 September 2026 we ran both models on 400 new tasks, such as applying a return policy or triaging a bug report. Each task had a short list of described options. Claude Opus 5.5 got 99.2% right and Jev 97.2%. Claude cost 87 times as much per task.

This time the answers were fixed in advance, so the scores are accuracy, not agreement with Claude. Claude did write the 400 tasks, and 400 is a small set, so the 2-point gap is a guide rather than a ranking. Opus 5.5 ran with its default thinking on, which also makes it slower: 2.5 seconds median against 0.46 for Jev.

Jev's confidence held up too. Calibration measures whether a model's stated confidence matches how often it is right. Jev's calibration error was 0.020, meaning its confidence was off from its real hit rate by about 2 points on average. That is why the advice above works: trust Jev when it is sure, and send the rest to Claude.

The atlas is a much harder task: 276 fuzzy categories, many of which overlap. That is why agreement there was 46.4% while accuracy here was 97.2%.

The full results are in Jev vs Tev. Our post on decision model vs LLM covers when the larger model is worth its price.

What about non-English text?

Jev is about 11 points weaker on skills written in Chinese or Japanese characters than on skills in the Latin alphabet. Cyrillic sits between the two, and Korean is lowest.

ScriptSkillsJev matched Claude's category
Latin alphabet77,32848.1%
Cyrillicn/a42.5%
Chinese and Japanese characters12,53236.9%
Koreann/a35.0%

If much of your text is not in English, raise the confidence bar for those items or send more of them to Claude.

Using both: Jev as a second labeller

The best use of Jev next to Claude is as a second, independent opinion. It follows the approach we recommend for any Jev job: ask everything in one call, act on the confident answers, and send the rest to a person or a rule. Two labellers that agree can be accepted. Two that disagree go to a person. For the reviewer, agreeing or disagreeing with one proposed label takes two clicks. Picking 1 of 276 categories from scratch is a chore.

Jev as a second labellerJev and Claude label the same item independently. When they agree, the label is accepted. When they disagree, a person sees one proposed label and agrees or disagrees, which takes two clicks.YESNOItemJevClaudeAgree?Accept labelPersonagree or disagreeLEGENDLabellerCheckHuman decision

LLM as a judge works on the same idea: a cheaper model checks a costlier one, and people check the disagreements. It fits a human on the loop review, where your team samples and corrects rather than approving every item.

We see the same pattern in our operations inbox, which labels every incoming email, call and message. There Jev handles the items it is sure of and Claude decides the uncertain ones. The Jev vs GLM-5.2 post has those numbers. Jev is also listed in AI Space's model catalog.

When not to use Jev

  • Jev cannot write. It cannot explain a label, draft a reply or summarise. Anything that needs words needs another model (TypeSafe).
  • It has no step where it thinks before answering, so instructions have to be literal and precise (Sean Goedecke).
  • It takes up to 32,000 tokens per request (Cloudflare docs) and does not read images yet.
  • Our figures are agreement with Claude, not ground truth. We did not measure either model against a human answer key.
  • TypeSafe's speed and price claims are its own, self-tested. We report only what we measured.

For a fuller picture of the model itself, start with what is Jev.

FAQ

Is Jev better than Claude for classification?

Not item for item. On the 91,306 atlas skills Claude labelled itself, Jev matched Claude's category 46.4% of the time. Its strength is confidence: at 0.9 or above it matched 93.6%, on about a tenth of the skills. For large piles, Jev as a first pass plus Claude for the uncertain items costs far less than Claude alone, and your reviewers see only the disagreements.

How much did Jev cost to label 143,891 items?

$30.97 by token count, and $26.63 on the Cloudflare bill. Each skill was about 4,400 tokens of input and needed two calls, because Jev takes at most 255 options per question. Jev's price is $0.042 per million input tokens with free output. We estimate the same pass at about $11,600 with Claude Opus 5 at list prices, sending one skill per request.

Can Jev replace Claude?

Not for work that needs words. Jev cannot write, explain or summarise. It picks from a list you give it and returns a confidence score. It can take the easy share of a labelling job off Claude, which on our atlas was a fifth of the items at 88.4% agreement when we kept only answers at 0.8 confidence or above.

Does Jev work in languages other than English?

Yes, but less well. On our atlas it matched Claude's category 48.1% of the time for skills in the Latin alphabet, 42.5% for Cyrillic, 36.9% for Chinese and Japanese characters and 35.0% for Korean. Raise the confidence bar for non-English items, or send more of them to Claude.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…