On this page
- Which should you pick, Jev or Tev?
- At a glance: Jev and Tev side by side
- How did Jev and Tev do on our tests?
- Is Tev faster and cheaper than Jev?
- Can you trust their confidence scores?
- What is Tev, and how was it made?
- How do frontier models compare on the same tasks?
- What are the limits of this test?
- FAQ
- Sources
Jev is more accurate than Tev; Tev is faster and cheaper. Together AI released Tev1-4B-experimental on 23 September 2026 with no accuracy numbers. We tested it on 25 September against TypeSafe's Jev on 400 new tasks. Jev got 97.2% right and Tev 90.0%. Tev answered faster, in 173 milliseconds against 459. It also cost less: $10.33 per million tasks against $19.64.
Need the right answer on close reading? Pick Jev. Need to run it yourself, or need speed on easy tasks? Pick Tev.
- Tev is Together AI's open-weight model trained to act like Jev. Together says it cost about $17 and 25 minutes to train. Together's launch post gave the recipe but no scores.
- Jev is TypeSafe's decision model, launched on 15 September 2026 as a hosted service. It answers questions about text and writes no text.
- Accuracy: Jev 97.2%, Tev 90.0%, on 400 tasks Tev had never seen.
- Pairs: each task comes with a twin that differs by one small edit. Jev got both twins right 95.0% of the time, Tev 80.0%.
- Speed: Tev's typical answer took 173 milliseconds, Jev's 459.
- Cost: both list $0.042 per million tokens read, with output free. Tev still costs about half as much per task, because Jev bills about twice as many tokens for the same text.
- The whole test cost $1.10. The data, code and results are public on GitHub.
Which should you pick, Jev or Tev?
Pick Jev when a wrong answer costs you something. Pick Tev when you need to host the model, need the fastest reply, or your tasks are easy.
Pick Jev if your task turns on one word: a "not", a date, a number, who is speaking. On our hard tasks Jev got 95.9% right and Tev 84.1%. Pick Tev if your data cannot leave your servers, you want to fine-tune your own version, or you sort plain requests where both models score near 100%. Keep a frontier model in reserve for the cases neither handles; see decision model vs LLM for that trade.
New to Jev? Start with what is Jev.
At a glance: Jev and Tev side by side
Jev and Tev each pick one answer from a list you give them, at the same list price. They differ in who runs them, how fast they answer and how often they are right. A token is a small chunk of text, about three quarters of a word.
| Tev | Jev | |
|---|---|---|
| Who makes it | Together AI | TypeSafe |
| Run it yourself? | Yes: open weights on Hugging Face | No: hosted API |
| List price per million tokens | $0.042 to read, output free | $0.042 to read, output free |
| Cost per million tasks, our test | $10.33 | $19.64 |
| Typical reply time, our test | 173 ms | 459 ms |
| Accuracy, our test | 90.0% | 97.2% |
| How you call it | A normal chat call; it replies with one option letter | A choice question; it returns the option plus a probability for each |
Prices come from Together on X and TypeSafe's launch post; the rest comes from our results.
How did Jev and Tev do on our tests?
Jev got 389 of 400 tasks right and Tev got 360. The gap is too large to be chance.
| Accuracy | Both halves of a pair right | |
|---|---|---|
| Tev | 90% | 80% |
| Jev | 97.2% | 95% |
Source: Our benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
We wrote 400 new tasks in 8 families: support requests, return policies, agent routing, content moderation, review sentiment, fact checking, approving agent actions and bug triage. They come in 200 pairs. The two halves of a pair differ by one small edit that changes the right answer. Here is one pair. A jacket was delivered on 2 March, and the policy allows returns within 30 days. In one half the customer asks on 20 March, so the answer is a full refund. In the other they ask on 15 April, so the answer is to decline. A model that skims for topic words gets one half right and the other wrong.
Pair accuracy, the second bar, catches that skimming: Tev got both halves right on 80.0% of pairs, Jev on 95.0%.
Head to head
Jev and Tev were both right on 356 tasks and both wrong on 7. Jev alone was right on 33. Tev alone was right on 4. A standard test for paired results (McNemar's) puts the chance of a gap this lopsided at about one in a million (p = 1.08e-06).
Hard tasks widen the gap
About 40% of pairs are marked hard: negations, sarcasm, rule edge cases, misleading words.
| Easy (230 tasks) | Hard (170 tasks) | |
|---|---|---|
| Tev | 94.3% | 84.1% |
| Jev | 98.3% | 95.9% |
Source: Our benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
On easy tasks the two are 4 points apart. On hard ones, nearly 12.
Where Tev slipped most
Tev's three weakest families all ask it to apply a rule, not spot a topic:
- Approving agent actions: 82% (Jev 100%). An AI agent proposes a shell command, a database query or an email. The model must approve it, send it to a person or block it. The difference is often one flag, such as a test run marker removed from a command, or an outside email address.
- Content moderation: 84% (Jev 98%). Quoting abuse to report it is allowed; posting it is not. The words are the same.
- Review sentiment: 84% (Jev 98%). Five levels, with "but" clauses and sarcasm.
Tev scored 100% on support requests, and so did Jev. Each family has only 50 tasks, so treat these family scores as a direction, not a precise number. Tev's weights are public, so a team could train it further on its own weak families.
Where each model's mistakes land
Tev leans toward caution and misses the rarer answers. Jev's few misses sit in one place: counting days on return policies.
- Tev rarely asks a clarifying question when it should. Four tasks called for one; Tev asked once.
- Tev sends safe actions to a person. It approved only 62% of the safe agent actions and passed the rest to human review. That errs toward caution, which is the cheaper kind of mistake.
- Tev misses negative reviews. It picked "negative" on only 40% of the reviews that were negative.
- Jev's misses cluster in return-policy day counting. It declined only 71% of the returns that fell outside the window. When it offered a full refund it was right 88% of the time, and store credit 75%.
One balanced score sums this up. It rates each possible answer on its own, then averages them, so a model that ignores rare answers scores lower. Out of 100, Tev scored 88.5 and Jev 95.5. Some answers appear only a handful of times, so read those per-answer rates as a direction.
Is Tev faster and cheaper than Jev?
Yes on both. Tev's typical reply came back 2.7 times faster, and it cost about half as much per task.
| Typical (p50) | Slow end (p95) | |
|---|---|---|
| Tev | 0.173s | 0.215s |
| Jev | 0.459s | 0.954s |
Source: Our benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev
Latency is the wait between sending a task and getting the answer. Half of Tev's replies took 173 milliseconds or less; for Jev it was 459. The slowest 1 in 20 replies show the spread. Tev's took up to 215 milliseconds, Jev's up to 954. Tev is also steadier. One caveat favours Tev: we reached Jev through our AI Space gateway, which adds a hop. We called Tev on Together directly.
Why does Jev cost more at the same price? Both list $0.042 per million tokens read, and neither charges for output. But for the same task, Jev billed 468 input tokens on average and Tev 246. TypeSafe likely wraps its own prompt around your text. So a million tasks cost $19.64 on Jev and $10.33 on Tev. At that scale, both bills are small. Jev pricing works through bigger jobs.
Can you trust their confidence scores?
Yes, for both, and Jev's are closer to the truth. Calibration means a model's stated confidence matches how often it is right: when it says 90% sure, it should be right about 9 times in 10.
Jev returns a probability for every option. For Tev we read the probability from its first reply letter. Both scored well on the usual calibration measure, expected calibration error. Stated confidence missed the real hit rate by 2.0 points on average for Jev. For Tev the miss was 3.1 points. A second measure, the Brier score, also rewards being right more often, so the accuracy gap shows up there too: Jev 0.041, Tev 0.158, where lower is better. In practice, both let you send low-confidence answers to a person. Jev's threshold will hold more reliably. How to use Jev shows that setup.
What is Tev, and how was it made?
Tev is a 4 billion parameter open model that Together AI trained to answer like Jev. Together built it as a small add-on (a LoRA fine-tune) on top of Alibaba's Qwen3.5 4B.
Together trained it on 37,840 examples. The run took about 25 minutes and cost about $17. The examples come from well-known public datasets of yes/no questions, bank support requests, news headlines and film reviews. Together added generated tasks on policies and routing. The recipe is on GitHub. You call it with temperature 0 and a limit of 8 output tokens, and it replies with one option letter. It still writes that letter the way a chat model does. It does not use Jev's method of answering many questions in one pass.
Because Tev learned from those public datasets, testing it on them would reward memory. So we wrote new tasks. Claude drafted all 400, and none come from those datasets. How to test a classifier explains the method.
How do frontier models compare on the same tasks?
Two large chat models scored higher still, at far greater cost. GLM-5.3 is an open model from Z.ai; it got 99.0%. Claude Opus 5.5 from Anthropic got 99.2%. They cost $809.08 and $1,707.74 per million tasks, and took about 2.5 to 2.9 seconds per reply. We ran them only as a reference point.
You can get close to them for far less. Jev's mistakes cluster on a few answers, so we asked GLM-5.3 again only when Jev gave one of three answers it gets wrong most often. That setup reached 98.1% at $86 per million tasks. With Tev first, the best setup reached 97.1% and sent 39% of tasks to the larger model, because Tev's mistakes spread over many answers. LLM cascade walks through the method. The trade between a decision model and a chat model has its own post: decision model vs LLM.
What are the limits of this test?
Our benchmark is one small test, run once, on one day. Read the numbers with these limits in mind:
- Claude wrote all 400 tasks, and the same kind of model checked the labels. We spot-checked them. A careful person could still dispute an item or two.
- 400 tasks is small. The headline gap is solid. Family scores rest on 50 tasks each and could move a lot.
- Jev's timing includes our gateway hop; Tev's does not. Timing also depends on region and load.
- We tested only pick-one questions. Jev's yes/no and score types have no Tev match here.
- The two APIs count tokens differently, which is why the same list price gives different costs.
- GLM-5.3 and Opus 5.5 ran through AI Space with their default reasoning settings.
Jev and Tev also missed the same 7 tasks. We reviewed each; three could be argued. We kept every label as written. Full details are in the methodology.
FAQ
Is Tev as good as Jev?
Not on our test. Jev got 97.2% of 400 new tasks right and Tev 90.0%. On hard tasks the gap grew to 95.9% against 84.1%. Tev matched Jev on simple support requests. It fell furthest behind on tasks that apply a rule, such as approving an AI agent's proposed action.
Why is Tev cheaper if both cost $0.042 per million tokens?
Jev bills about twice as many input tokens for the same text: 468 on average against Tev's 246. TypeSafe probably adds its own prompt around yours. Neither charges for output. So a million tasks cost $19.64 on Jev and $10.33 on Tev in our run.
Can I run Tev on my own servers?
Yes. Together published Tev's weights on Hugging Face and the training recipe on GitHub. Jev is a hosted service only. If your data cannot leave your own servers, Tev is the one of the two you can use.
Did Together publish accuracy numbers for Tev?
No. Together's 23 September 2026 launch post explained how Tev was trained and how to call it, with no accuracy scores. Our test on 25 September is the first independent one we know of. The data and code are public so anyone can rerun it.
Should I use Tev or a frontier model?
Choose by how costly a wrong answer is. GLM-5.3 and Claude Opus 5.5 scored 99.0% and 99.2% on the same tasks. They cost 78 to 165 times more per task than Tev and took over two seconds per reply. One setup runs Jev first and sends only its riskiest answers to a larger model; in our test that reached 98.1% at $86 per million tasks.
Sources
- Together AI, How to train your own Jev for $17, 23 September 2026
- Together AI on X, Tev serverless pricing, 23 September 2026
- Together AI, Tev1-4B-experimental model card, accessed 25 September 2026
- Together AI, tev1 training recipe, accessed 25 September 2026
- TypeSafe, Introducing System One Models and Jev, 15 September 2026
- Cloudflare, GLM-5.3 on Workers AI, accessed 25 September 2026
- XY Space, Jev vs Tev benchmark: results and methodology, 25 September 2026
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project