On this page
- What we tested, and how
- How do you call Jev?
- Our recommended approach: one call, no prose, confidence decides
- Ask for facts, not verdicts
- Give it the list to choose from
- What if you have more than 255 options?
- Write instructions for a literal reader
- Send more text, and strip out code first
- Use confidence to decide who decides
- Running it at volume
- Settings that skewed our comparison models
- Where Jev does not fit
- FAQ
- Sources
To use Jev, send it your text plus yes/no, pick-one or score questions; it sends back probabilities instead of prose. Since 18 September 2026 you can call it through Cloudflare as typesafe/jev, and the text you send costs a fraction of a cent per item. It works best when you ask for small facts and let your own code turn them into the final decision.
- What it is: a model that answers questions with a choice, a yes/no probability or a score. It cannot write text. Our explainer on what Jev is covers the background.
- Where to call it: Cloudflare (
typesafe/jev), or TypeSafe's own API directly or through Pydantic AI. TypeSafe's API is in early access with a waitlist. - Limits: up to 255 options per choice question and 32,000 tokens of input (a token is about three quarters of a word).
- Our recommended approach: put all of an item's questions in one call, let Jev answer without prose, and send low-confidence answers to a person or a rule.
- Our tests: run 21 to 23 September 2026 on our operations inbox and our Skill Atlas.
What we tested, and how
We ran Jev on two jobs from 21 to 23 September 2026. The first is our operations inbox, which labels every incoming item (emails, calendar invites, call recordings, WhatsApp messages, AI coding sessions) as client, internal, noise or unknown, and tags a signal such as a request, a commitment or money. The second is our Skill Atlas, a catalogue of 143,891 AI agent skills sorted into 276 categories.
Three caveats apply to every number below. The inbox test labels were written by Claude against a written labelling policy, which flatters any Claude step, and a human check of 20 items is still pending. The Skill Atlas figures measure agreement with Claude's labels, not ground truth. And the tests are ours: one inbox, one catalogue, three days. The full results live in Jev vs GLM-5.2 and Jev vs Claude. This post keeps to what we would do again.
How do you call Jev?
The quickest route is Cloudflare's REST endpoint: post the model name, your text as state, and your questions, and you get JSON back. The request below follows the shape in Cloudflare's model page.
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run \
--header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"model": "typesafe/jev",
"input": {
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "Does this convey urgency?",
"criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" }
},
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations"
}
}
}
}
}'There are three question types. A choice picks one option from your list. A score places the text on a scale you define. A noul is TypeSafe's name for a yes/no question, and it returns a probability between 0 and 1. Choice and score answers also carry a confidence value and the probability of every option.
In Python, Pydantic AI maps a typed class straight onto Jev's questions. Each field becomes a question, and its docstring becomes the instruction:
from typing import Annotated, Literal
from pydantic import BaseModel, ConfigDict
from pydantic_ai import Agent
from pydantic_ai.models.typesafe import TypeSafeModel
from pydantic_ai.output import BoolCriteria
class Ticket(BaseModel):
model_config = ConfigDict(use_attribute_docstrings=True)
area: Literal['billing', 'bug', 'account']
"""Which team owns this ticket?"""
urgent: Annotated[bool, BoolCriteria(
true='The customer is losing money or has a deadline today.',
false='It can wait its turn in the queue.',
)]
"""Should this ticket jump the queue?"""
agent = Agent(TypeSafeModel('jev-latest'), output_type=Ticket)
result = agent.run_sync('My Android timeline is blank and my standup is in ten minutes.')
print(result.output)
print(result.response.provider_details['confidence'])The Pydantic AI route needs a TYPESAFE_API_KEY, so it goes through TypeSafe's own API, which is still early access. If you already pick between agent frameworks, our LangGraph vs Pydantic AI comparison covers the wider trade-off. Jev is also listed in AI Space's model catalog.
Our recommended approach: one call, no prose, confidence decides
Put every question about an item into one call. Jev answers them together and returns no prose, so there is nothing to parse and it cannot invent an option you did not list. Then read the confidence on each answer: act on the confident ones, and let a person or a rule decide the rest.
The rest of this guide covers each part: which questions to ask, how to give Jev a list to choose from, and where to set the confidence bar. On our inbox, the approach scored 90.7% against 84.0% for the model it replaced, and 95.3% once Claude decided the unsure third.
Ask for facts, not verdicts
Ask Jev for small facts it can check in the text, and let your own code turn those facts into the verdict. Asked for our signal label directly, Jev got it right 64% of the time. Asked a dozen yes/no facts and routed through a small trained model, it got 85 to 87% right.
Our first version was a straight copy of the prompt we used with a chat model. It worked, but the rebuild worked better. The rebuild makes one call per item, and that call asks for:
- One choice over our own list: which organisation is this item about.
- About twelve yes/no facts: does it ask for something, does it mention an amount, does it promise a date.
- The direct label as one more input: we still ask for the final label, but only as a hint.
Two small statistical models, each trained on 300 labelled items, then turn those answers into the final label. They are cheap to retrain and easy to inspect. When a label is wrong, you can see which fact tipped it.
The rebuild beat the chat model it replaced, GLM-5.2. On a blind set of 150 items, Jev picked the right label 90.7% of the time and GLM-5.2 84.0%. Jev caught 48 of the 50 client items, GLM-5.2 caught 42, and neither flagged a non-client as a client. That is one inbox and a test set labelled by Claude, so treat it as a direction rather than a promise.
Give it the list to choose from
Jev picks well when the options are yours and each one is described. Our inbox passes a registry of 78 organisations, each described by its people, email domains, projects and a short note, and asks Jev to choose one.
Two details made the list work:
- Leave your own companies off. Every item mentions us somewhere. With our own names on the list, they pull answers towards themselves.
- Add a new-company flag. A yes/no question asks whether the item comes from an organisation not on the list. In our tests it caught every unregistered organisation, which is how new clients get noticed instead of being filed as noise.
The option descriptions matter more than the question wording. On 25 September 2026 we ran Jev on 400 new pick-one tasks five ways. With each option described, Jev got 97.2% right. With bare option names, it fell to 92.8%, a drop of about 4.5 points. Rewording the question, reversing the option order or adding a line of reading guidance moved it by 0.3 points or less. Together AI's open model Tev lost 6.8 points without descriptions.
| Accuracy | |
|---|---|
| Default prompt | 97.2% |
| Plus a line of reading guidance | 97.2% |
| Options in reverse order | 97.2% |
| Generic question | 97% |
| Option names only, no descriptions | 92.8% |
Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
Claude wrote these 400 tasks, and the set is small, so small gaps between versions are noise. The drop without descriptions is the one to act on. Write one line per option, and test a change before you trust it: how to test a classifier covers the method, and Jev vs Tev has the full results.
What if you have more than 255 options?
Split the choice into two calls: first the group, then the option inside that group. A choice question takes at most 255 options, and the Skill Atlas has 276 categories.
We grouped the categories into 58 teams of at most 21 each. The first call picks the team. The second call picks the category from that team's short list. Across the full run of 143,891 skills, Jev never answered with a category outside the list. Agent tools that need a clean enum out of free text are the obvious fit for this shape.
Write instructions for a literal reader
Jev follows your category descriptions word for word, so describe each one by what makes it different, not by what every item has in common. Pydantic AI's docs list literal readings of ambiguous wording as a known weakness, and we hit it.
In the Skill Atlas, Jev put 48.9% of 91,306 skills into the Skill & Agent Engineering department, where Claude put 22.6%. Our instructions put skills about writing, installing, evaluating or orchestrating other skills in that department, whatever domain they mention. Claude applied the rule to skills about skills; Jev applied it far more widely. That one department accounted for 85% of the department-level disagreements. Among the 13,519 skills Jev filed under skill creation and templates were a real estate trust profile builder, an investment thesis tracker and a restaurant finder.
A correction applied afterwards lifted department agreement with Claude from 66.9% to 73.3%. The fix belongs in the descriptions, before the run. Jev vs Claude has the full breakdown.
Send more text, and strip out code first
Send Jev more of the item than you would send a chat model, but replace code, links and shell commands with short stand-ins first. Input is cheap, and our inbox now sends up to 36,000 characters per item, where the GLM-5.2 version sent 4,500.
Command-like text caused the one error we could not retry away. Some items returned a 402 "Payment error" (code 2021) even with credit on the account. Eleven items failed on every retry. All eleven passed once code blocks, links and shell commands were swapped for short stand-ins such as [code]. Separate from those, random bursts of 402s appeared and cleared on their own.
Use confidence to decide who decides
Use Jev's confidence score to split the work: Jev keeps the items it is sure of, and a stronger model decides the rest. On our inbox, 69% of items came back at 0.8 confidence or higher, and Jev was right on 98.1% of them.
We send everything under 0.8, and every new-company flag, to Claude. Together the two lanes got 95.3% right and caught all 50 client items. Claude's cost is not included in Jev's per-item figure.
Measure the fallback on the uncertain items before you choose it. Those items are hard by definition, and GLM-5.2 did worse on them than Jev did:
| Who decides the uncertain items | Right, blind set (51 items) | Right, development set (111 items) |
|---|---|---|
| Claude | 90.2% | 87.4% |
| GLM-5.2 with our registry and policy | 80.4% | 75.7% |
| Jev itself | 76.5% | 77.5% |
| GLM-5.2 as it ran before | 70.6% | 65.8% |
| GPT-5.4 mini | n/a | 72.7% (22 items) |
| gpt-oss-120b | n/a | 60.4% |
The same threshold works for sorting a catalogue. In the Skill Atlas, the 18.6% of skills that Jev scored at 0.8 or higher matched Claude's category 88.4% of the time, against 46.4% across all skills. Jev vs GLM-5.2 covers the inbox numbers, and our guide to agentic workflows covers where a human sits in a lane like this.
A second method routes by which answer Jev gave, not by its confidence. Label a few hundred examples, then measure how often Jev is right each time it gives each answer. Send only the answers it often gets wrong to a stronger model. In our public benchmark, three risky answers went to GLM-5.3: 8% of tasks were re-asked, and the mix got 98.1% right at $86 per million tasks. LLM cascade has the full sweep, and how to test a classifier shows how to build the labelled set.
Running it at volume
Run about eight requests at a time, retry what fails, and stop the job if failures pile up. At eight at a time our calibration run managed about 4.5 items a second, and 2.5 to 4 a second over a long run. At 24 at a time throughput fell to 3.1 a second and 8% of requests failed.
The retry rules we settled on:
- Network errors: retry after a short, fixed pause.
- 429 (too many requests): back off, doubling the wait each time.
- 402 and 401 bursts: put the item back in the queue and move on.
- 300 failures in a row: stop the job. Something upstream has changed, and retrying will not fix it.
The full Skill Atlas pass took about 17 hours against the 9 we predicted, because of throttling, 402 bursts and a credit balance that ran out mid-run. We could not find an API for checking the balance, so turn on automatic top-up before a long job. Read your spend from Cloudflare's bill, not from token estimates: ours ran high twice. Jev pricing covers the costs in detail.
Settings that skewed our comparison models
If you run your own bake-off, check the other models' defaults as well as Jev's settings. Two settings cost us time.
- Claude 5 models: on the gateway, extended thinking is on by default. With a 200-token output cap, the thinking used the budget and the answer came back empty.
- GLM-5.2: setting
reasoning_effortto"none"made its answers worse. Passingchat_template_kwargswithenable_thinking: falseworked, and cut its time from 7.4 seconds to 2.5 per item.
Our runs of Claude Opus 5 and Claude Sonnet 5 on the hard benchmark stopped at 11 and 15 items, so those two rows are thin.
Where Jev does not fit
Jev cannot write a sentence, so anything that needs a summary, a reply or an extracted name goes to a text model. Pydantic AI's docs also list arithmetic, dates, multi-step reasoning and text written to steer the model as weak spots. It does not read images yet. Our tests cover one inbox and one catalogue, mostly in English. In the Skill Atlas, agreement with Claude was 48.1% on skills written in the Latin alphabet, 42.5% in Cyrillic and 35.0% in Korean.
FAQ
Can I use Jev without a Cloudflare account?
Yes, but only with early access. TypeSafe's own API, which Pydantic AI's TypeSafeModel uses through a TYPESAFE_API_KEY, is open to developers coming off a waitlist. Cloudflare lists Jev as typesafe/jev for anyone with an account. The questions and answers have the same shape either way.
Can Jev generate text or extract names?
No. Jev answers typed questions with choices, yes/no probabilities and scores, and it does not produce free text. In Pydantic AI, string fields can be handed to a text model through a fallback. If your workflow needs a written reply or an extracted value, pair Jev with a text model and let Jev handle the sorting and routing it is fast at.
How many options can a Jev choice question have?
Up to 255. If you have more, split the choice into two calls: pick a group first, then pick the option within that group. We sorted 276 categories this way, through 58 teams of at most 21 categories each, and Jev never returned an answer outside the list across 143,891 items.
What confidence threshold should I use with Jev?
Start at 0.8 and check it on your own data. On our inbox, items at 0.8 or higher were 69% of the traffic and 98.1% right. On our Skill Atlas, the same threshold kept 18.6% of items. The right cut-off depends on your job, and on how good the fallback is at the items Jev is unsure about.
How much text can I send to Jev?
Cloudflare lists a context window of 32,000 tokens, roughly 24,000 words. Input is cheap at $0.042 per million tokens, so send more of each item than you would to a chat model. Our inbox sends up to 36,000 characters per item. Replace code, links and shell commands with short stand-ins first, because command-like text triggered errors we could not retry away.
Sources
- Jev model page and API reference, Cloudflare Docs, accessed 25 September 2026
- TIL: TypeSafe AI's Jev on Cloudflare AI Gateway, AI Engineer Guide, 18 September 2026
- Introducing System One Models & Jev, TypeSafe, 15 September 2026
- TypeSafe models, Pydantic AI docs, accessed 25 September 2026
- Jev means structured output is interesting again, Sean Goedecke, 16 September 2026
- Jev vs Tev benchmark results, prompt sensitivity, XY Space, 25 September 2026
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project