On this page
Jev is an AI model from TypeSafe AI that answers questions about your text but never writes a sentence back. You send the text plus yes/no, pick-one or score questions, and it returns each answer with a probability. On 25 September 2026, Bloomberg covered it as the model that "can't chat" taking on bigger rivals.
Key facts
- Who makes it: TypeSafe AI, which came out of stealth on 15 September 2026 with a $40M seed round.
- What it does: answers typed questions about a piece of text. It does not generate text, and it does not take images yet.
- Speed: 70 to 500 milliseconds per call, by TypeSafe's own figures.
- Price: $0.042 per million input tokens (a token is about three quarters of a word). Output is free.
- Where to get it: Cloudflare AI Gateway since 18 September 2026, Pydantic AI, and TypeSafe's own API in early access.
- Our tests: it labelled 143,891 AI agent skills for $26.63 on the bill and never answered outside the list.
What is Jev?
Jev is a decision model. It takes a "state", which is your text, plus a set of typed questions. It returns an answer to each question with a probability attached. It cannot write an email, a summary or a reply. TypeSafe describes the output as "type-safe structured values", meaning every answer is one of the options you defined before the call.
Cloudflare's documentation says Jev "returns calibrated answers with probabilities and confidence." Calibrated means the probability is meant to match how often the answer is right. When Jev says 0.9, it should be right about nine times in ten.
Giving up text is the point. A chat model can answer anything, in any words, and you then have to check that its words fit your software. Jev gives up the words. Your software gets a value it can use directly, such as true, "billing" or a score.
How does Jev work?
You send one piece of text and a list of questions, and each question has one of three types. TypeSafe calls its yes/no type a "noul", a true-or-false answer with a probability. The other two are "choice", which picks one option from a list of up to 255, and "score", which rates the text on a scale.
Here is a small example. The state is an email from a customer: "Our invoice 2291 was charged twice this month. Can someone look before Friday?"
| Question type | The question | What comes back |
|---|---|---|
| Yes/no (noul) | Is the sender asking for a refund or correction? | true or false, with a probability |
| Choice | Which team should handle this: billing, support or sales? | One of the three options, with a probability for each |
| Score | How urgent is this? | A score, with a confidence |
This is the approach we recommend. Put all your questions in one call. Jev answers them together and returns no prose, so there is nothing to parse and no chance of a made-up fourth team. A person or a rule then decides what to do with the low-confidence answers. How to use Jev walks through it on a real job.
What is a "System One" model, and why is it so fast?
"System One" is TypeSafe's name for a model that answers directly instead of reasoning step by step. The term comes from Daniel Kahneman's split between fast, intuitive thinking and slow, deliberate thinking. Jev is built for the fast kind.
Speed comes from how the answer is produced. A chat model writes its reply one token at a time, and each token waits for the one before it. Jev works differently. As Sean Goedecke explains, it can "produce answers to many questions in parallel in a single forward pass". He puts the fastest response at around 70 milliseconds and the slowest at 500, against "a couple of seconds for normal LLMs".
| Seconds per item | |
|---|---|
| Jev | 0.28s |
| GLM-5.2 | 7.4s |
Source: Our tests on 303 operations inbox items, 21 to 23 September 2026
TypeSafe claims Jev is 40 to 200 times faster than frontier chat models. The comparison is TypeSafe's own, run on its own workflow tests, so treat it as a vendor figure.
The cost of that speed is thinking time. Jev cannot reason through a problem before it answers. Goedecke expects this to "cap this kind of model around the strength of non-reasoning LLMs", and doubts TypeSafe has a lasting technical lead. Both points are opinion, but they match what we saw on the hardest items.
How much does Jev cost?
Jev costs $0.042 per million input tokens, and output is free. The context window, meaning how much text it can read in one call, is 32,000 tokens. Each call carries one state, so you pay once per item for however much text you send.
Because output is free, adding more questions to a call costs almost nothing. The bill tracks how much text you send, not how many answers you get. For a worked monthly bill and the difference between the token estimate and the invoice, see Jev pricing in practice.
What we found when we put it to work
In our tests, Jev sorted our operations inbox better than the chat model we used before. It also labelled a catalogue of 143,891 items for $26.63. We ran these tests from 21 to 23 September 2026 on two real jobs.
Our operations inbox. Every incoming item gets labelled as client, internal, noise or unknown, and tagged with a signal such as request, commitment or money. Items include email, calendar invites, call recordings, WhatsApp messages and AI coding sessions. That job ran on GLM-5.2, a general chat model. A straight swap to Jev scored 89.0% against 90.1%. After we rebuilt the task around Jev's question types, it scored 90.7% against 84.0% on a blind holdout of 150 items.
| On 150 held-out inbox items | Jev (rebuilt) | GLM-5.2 |
|---|---|---|
| Right label | 90.7% | 84.0% |
| Client items caught | 48 of 50 | 42 of 50 |
| Cost per item | $0.0004 | $0.0051 |
| Time per item | 0.3 seconds | 7 to 10 seconds |
Jev's confidence score also proved useful. Items where Jev's confidence was 0.8 or higher made up 69% of traffic, and 98.1% of those were right. When Claude decided the less certain third, the combined system reached 95.3% and caught all 50 client items. Claude's cost is not included in the per-item figure. The full comparison is in Jev vs GLM-5.2.
Our Skill Atlas. This is a catalogue of 143,891 AI agent skills sorted into 276 categories. Jev labelled all of them for $26.63 on Cloudflare's bill ($30.97 by our own token count), with zero answers outside the category list. Its labels matched Claude's 46.4% of the time overall, rising to 93.6% on the tenth of skills where Jev was 0.9 confident or more. The detail is in Jev vs Claude.
The inbox and Skill Atlas tests have caveats. Claude wrote the inbox holdout labels against a written labelling policy, which flatters any Claude lane, and a human check of 20 items is still pending. GLM-5.2 changes its own label on 16% of re-runs. The Skill Atlas numbers measure agreement with Claude's labels, not ground truth.
An independent test. On 25 September 2026 we ran Jev on 400 new tasks written for the purpose. One asks whether a return policy allows a refund. Another asks whether an agent's shell command is safe to run. The tasks come in pairs, where a small edit flips the right answer.
Jev got 97.2% of tasks right, and both halves of 95.0% of pairs. Together AI's open model Tev got 90.0% of tasks right. Jev's misses clustered on a few answers, mostly counting days under the return policy. Sending just those answers to GLM-5.3 for a second opinion reached 98.1% at $86 per million tasks; LLM cascade shows how.
Two large chat models, GLM-5.3 from Z.ai and Claude Opus 5.5 from Anthropic, did best at about 99%. They cost 41 to 87 times Jev's price per task. Claude wrote the tasks, and 400 is a small set, so read these as a guide. The full comparison is in Jev vs Tev.
| Accuracy | Both halves of a pair right | |
|---|---|---|
| Tev | 90% | 80% |
| Jev | 97.2% | 95% |
| GLM-5.3 | 99% | 98% |
| Claude Opus 5.5 | 99.2% | 98.5% |
Source: Our Jev vs Tev benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
These tasks are cleaner than our Skill Atlas. Each has a handful of options, and every option is described. That is why the scores here sit far above the 46.4% agreement with Claude on the atlas.
Where Jev fits, and where it does not
Jev fits high-volume jobs with a fixed list of answers, and fails on anything that needs new text or long reasoning. The confidence score is what makes it practical, because it tells you which answers to pass to a person or a larger model.
Good fit
- Labels and routing: which team, which client, which category.
- Yes/no checks on every item, such as "does this mention money?"
- Scores at volume, such as urgency or relevance.
- A screen on items before a larger model or a person sees them, the same shape as LLM-as-a-judge at a fraction of the cost.
- The fast first step inside an agentic workflow.
Poor fit
- Anything that needs written output: replies, summaries, extracted names. In Pydantic AI, a text field escalates to a language model.
- Hard judgement calls. On a set of 60 hard skills, a larger model got more right than Jev: 52.3% against 38.3%. That model, Nemotron 3 120B, cost far more per item.
- Text full of code, URLs or shell commands. Some of ours failed with a payment error every retry, and passed once we swapped those parts for neutral stand-in text.
- Lists longer than 255 options. We split our 276 categories into 58 teams and asked in two calls.
- Text longer than 32,000 tokens, or images.
Jev also reads instructions word for word. A category called "skill and agent engineering" pulled in a restaurant finder and an investment thesis tracker, because they were technically skills. A correction pass afterwards lifted department-level agreement with Claude from 66.9% to 73.3%. How to use Jev covers how to write the questions.
Where can you use Jev?
You can call Jev through Cloudflare AI Gateway as typesafe/jev, through Pydantic AI's TypeSafeModel, or through TypeSafe's own API once you are off the waitlist. The Cloudflare route has been live since 18 September 2026 and offers zero data retention.
Pydantic AI treats Jev as a decision model. You describe the answer as a typed object, each field gets its own confidence, and any free-text field is handed to a language model. That makes the pairing natural: Jev takes the choices, a chat model writes the words, and a person reviews what neither is sure of.
Jev is also listed in AI Space's model catalog.
If you want weights you can run yourself, Together AI released Tev1-4B-experimental on 23 September 2026. Tev is an open model trained to answer the same kind of pick-one question. Together's post shows how it was trained for about $17. In our test it was cheaper and faster than Jev, and less accurate. Jev vs Tev has the numbers.
FAQ
Can Jev chat or write text?
No. Jev answers typed questions (yes/no, pick one from a list, or a score) and returns each answer with a probability. It does not write replies, summaries or any free text. When a task needs written output, pair Jev with a chat model: Jev makes the decision, and the chat model writes the words. Pydantic AI does this automatically for text fields.
Who makes Jev?
TypeSafe AI makes Jev. Three founders started the company in 2024. It worked in stealth for about two years, then announced Jev on 15 September 2026 with a $40M seed round. Jev is its first "System One" model, built for fast structured decisions.
Is Jev better than ChatGPT or Claude?
For fixed-list decisions at volume, it can be. In our operations inbox test, a rebuilt Jev setup beat GLM-5.2 on accuracy at under a tenth of the cost. On hard judgement calls, larger models did better, and Claude still decides our least certain items. Jev is a fast first pass, not a replacement for a reasoning model.
Can Jev hallucinate?
Jev cannot invent an answer outside the options you give it. In our Skill Atlas run, it gave zero answers outside a 276-category list across 143,891 skills. It can still pick the wrong option. That is why the probability matters: low-confidence answers should go to a person or a larger model for review.
How fast is Jev?
TypeSafe puts Jev's response time at 70 to 500 milliseconds per call. In our tests, inbox items took about 0.3 seconds, and the Skill Atlas run had a median of 0.67 seconds per skill. A chat model doing the same inbox job took 7 to 10 seconds per item.
Is there an open-source Jev?
Not from TypeSafe. The closest is Tev1-4B-experimental from Together AI, released with open weights and its training recipe. It answers pick-one questions with a single option letter. In our 400-task test it got 90.0% right against Jev's 97.2%, at about half the cost per task. How to weigh the two is in Jev vs Tev.
Sources
- Bloomberg, "Jev, an AI Model That Can't Chat, Takes On Bigger Rivals", 25 September 2026: bloomberg.com
- TypeSafe AI, "Introducing System One Models & Jev", 15 September 2026: typesafe.ai
- Sean Goedecke, "Jev means structured output is interesting again", 16 September 2026: seangoedecke.com
- Cloudflare, "typesafe/jev" model documentation, accessed 25 September 2026: developers.cloudflare.com
- AI Engineer Guide, "TypeSafe AI's Jev on Cloudflare AI Gateway", 18 September 2026: aiengineerguide.com
- Pydantic, "TypeSafe" model documentation, accessed 25 September 2026: pydantic.dev
- Wikipedia, "Jev (AI model)", accessed 25 September 2026: en.wikipedia.org)
- Together AI, "How to train your own Jev for $17", 23 September 2026: together.ai
- XY Space, "Jev vs Tev benchmark results", 25 September 2026: github.com
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project