AI ComparisonsSep 25, 202610 min read

Jev vs Tev: We Tested Together's $17 Jev Copy on 400 Tasks

Jev vs Tev on 400 new tasks: Jev got 97.2% right, Together's $17 Tev got 90.0%. Tev answered in 173 ms against 459 ms and cost about half as much per task.

On this page
  1. Which should you pick, Jev or Tev?
  2. At a glance: Jev and Tev side by side
  3. How did Jev and Tev do on our tests?
  4. Is Tev faster and cheaper than Jev?
  5. Can you trust their confidence scores?
  6. What is Tev, and how was it made?
  7. How do frontier models compare on the same tasks?
  8. What are the limits of this test?
  9. FAQ
  10. Sources

Jev is more accurate than Tev; Tev is faster and cheaper. Together AI released Tev1-4B-experimental on 23 September 2026 with no accuracy numbers. We tested it on 25 September against TypeSafe's Jev on 400 new tasks. Jev got 97.2% right and Tev 90.0%. Tev answered faster, in 173 milliseconds against 459. It also cost less: $10.33 per million tasks against $19.64.

Need the right answer on close reading? Pick Jev. Need to run it yourself, or need speed on easy tasks? Pick Tev.

  • Tev is Together AI's open-weight model trained to act like Jev. Together says it cost about $17 and 25 minutes to train. Together's launch post gave the recipe but no scores.
  • Jev is TypeSafe's decision model, launched on 15 September 2026 as a hosted service. It answers questions about text and writes no text.
  • Accuracy: Jev 97.2%, Tev 90.0%, on 400 tasks Tev had never seen.
  • Pairs: each task comes with a twin that differs by one small edit. Jev got both twins right 95.0% of the time, Tev 80.0%.
  • Speed: Tev's typical answer took 173 milliseconds, Jev's 459.
  • Cost: both list $0.042 per million tokens read, with output free. Tev still costs about half as much per task, because Jev bills about twice as many tokens for the same text.
  • The whole test cost $1.10. The data, code and results are public on GitHub.

Which should you pick, Jev or Tev?

Pick Jev when a wrong answer costs you something. Pick Tev when you need to host the model, need the fastest reply, or your tasks are easy.

Pick Jev if your task turns on one word: a "not", a date, a number, who is speaking. On our hard tasks Jev got 95.9% right and Tev 84.1%. Pick Tev if your data cannot leave your servers, you want to fine-tune your own version, or you sort plain requests where both models score near 100%. Keep a frontier model in reserve for the cases neither handles; see decision model vs LLM for that trade.

New to Jev? Start with what is Jev.

At a glance: Jev and Tev side by side

Jev and Tev each pick one answer from a list you give them, at the same list price. They differ in who runs them, how fast they answer and how often they are right. A token is a small chunk of text, about three quarters of a word.

TevJev
Who makes itTogether AITypeSafe
Run it yourself?Yes: open weights on Hugging FaceNo: hosted API
List price per million tokens$0.042 to read, output free$0.042 to read, output free
Cost per million tasks, our test$10.33$19.64
Typical reply time, our test173 ms459 ms
Accuracy, our test90.0%97.2%
How you call itA normal chat call; it replies with one option letterA choice question; it returns the option plus a probability for each

Prices come from Together on X and TypeSafe's launch post; the rest comes from our results.

How did Jev and Tev do on our tests?

Jev got 389 of 400 tasks right and Tev got 360. The gap is too large to be chance.

Right answers on 400 new tasks
Right answers on 400 new tasks
AccuracyBoth halves of a pair right
Tev90%80%
Jev97.2%95%

Source: Our benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev

We wrote 400 new tasks in 8 families: support requests, return policies, agent routing, content moderation, review sentiment, fact checking, approving agent actions and bug triage. They come in 200 pairs. The two halves of a pair differ by one small edit that changes the right answer. Here is one pair. A jacket was delivered on 2 March, and the policy allows returns within 30 days. In one half the customer asks on 20 March, so the answer is a full refund. In the other they ask on 15 April, so the answer is to decline. A model that skims for topic words gets one half right and the other wrong.

Pair accuracy, the second bar, catches that skimming: Tev got both halves right on 80.0% of pairs, Jev on 95.0%.

Head to head

Jev and Tev were both right on 356 tasks and both wrong on 7. Jev alone was right on 33. Tev alone was right on 4. A standard test for paired results (McNemar's) puts the chance of a gap this lopsided at about one in a million (p = 1.08e-06).

Hard tasks widen the gap

About 40% of pairs are marked hard: negations, sarcasm, rule edge cases, misleading words.

Accuracy by task difficulty
Accuracy by task difficulty
Easy (230 tasks)Hard (170 tasks)
Tev94.3%84.1%
Jev98.3%95.9%

Source: Our benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev

On easy tasks the two are 4 points apart. On hard ones, nearly 12.

Where Tev slipped most

Tev's three weakest families all ask it to apply a rule, not spot a topic:

  • Approving agent actions: 82% (Jev 100%). An AI agent proposes a shell command, a database query or an email. The model must approve it, send it to a person or block it. The difference is often one flag, such as a test run marker removed from a command, or an outside email address.
  • Content moderation: 84% (Jev 98%). Quoting abuse to report it is allowed; posting it is not. The words are the same.
  • Review sentiment: 84% (Jev 98%). Five levels, with "but" clauses and sarcasm.

Tev scored 100% on support requests, and so did Jev. Each family has only 50 tasks, so treat these family scores as a direction, not a precise number. Tev's weights are public, so a team could train it further on its own weak families.

Where each model's mistakes land

Tev leans toward caution and misses the rarer answers. Jev's few misses sit in one place: counting days on return policies.

  • Tev rarely asks a clarifying question when it should. Four tasks called for one; Tev asked once.
  • Tev sends safe actions to a person. It approved only 62% of the safe agent actions and passed the rest to human review. That errs toward caution, which is the cheaper kind of mistake.
  • Tev misses negative reviews. It picked "negative" on only 40% of the reviews that were negative.
  • Jev's misses cluster in return-policy day counting. It declined only 71% of the returns that fell outside the window. When it offered a full refund it was right 88% of the time, and store credit 75%.

One balanced score sums this up. It rates each possible answer on its own, then averages them, so a model that ignores rare answers scores lower. Out of 100, Tev scored 88.5 and Jev 95.5. Some answers appear only a handful of times, so read those per-answer rates as a direction.

Is Tev faster and cheaper than Jev?

Yes on both. Tev's typical reply came back 2.7 times faster, and it cost about half as much per task.

Seconds per reply, as the caller sees it
Seconds per reply, as the caller sees it
Typical (p50)Slow end (p95)
Tev0.173s0.215s
Jev0.459s0.954s

Source: Our benchmark, 25 September 2026, github.com/choyiny/jev-vs-tev

Latency is the wait between sending a task and getting the answer. Half of Tev's replies took 173 milliseconds or less; for Jev it was 459. The slowest 1 in 20 replies show the spread. Tev's took up to 215 milliseconds, Jev's up to 954. Tev is also steadier. One caveat favours Tev: we reached Jev through our AI Space gateway, which adds a hop. We called Tev on Together directly.

Why does Jev cost more at the same price? Both list $0.042 per million tokens read, and neither charges for output. But for the same task, Jev billed 468 input tokens on average and Tev 246. TypeSafe likely wraps its own prompt around your text. So a million tasks cost $19.64 on Jev and $10.33 on Tev. At that scale, both bills are small. Jev pricing works through bigger jobs.

Can you trust their confidence scores?

Yes, for both, and Jev's are closer to the truth. Calibration means a model's stated confidence matches how often it is right: when it says 90% sure, it should be right about 9 times in 10.

Jev returns a probability for every option. For Tev we read the probability from its first reply letter. Both scored well on the usual calibration measure, expected calibration error. Stated confidence missed the real hit rate by 2.0 points on average for Jev. For Tev the miss was 3.1 points. A second measure, the Brier score, also rewards being right more often, so the accuracy gap shows up there too: Jev 0.041, Tev 0.158, where lower is better. In practice, both let you send low-confidence answers to a person. Jev's threshold will hold more reliably. How to use Jev shows that setup.

What is Tev, and how was it made?

Tev is a 4 billion parameter open model that Together AI trained to answer like Jev. Together built it as a small add-on (a LoRA fine-tune) on top of Alibaba's Qwen3.5 4B.

Together trained it on 37,840 examples. The run took about 25 minutes and cost about $17. The examples come from well-known public datasets of yes/no questions, bank support requests, news headlines and film reviews. Together added generated tasks on policies and routing. The recipe is on GitHub. You call it with temperature 0 and a limit of 8 output tokens, and it replies with one option letter. It still writes that letter the way a chat model does. It does not use Jev's method of answering many questions in one pass.

Because Tev learned from those public datasets, testing it on them would reward memory. So we wrote new tasks. Claude drafted all 400, and none come from those datasets. How to test a classifier explains the method.

How do frontier models compare on the same tasks?

Two large chat models scored higher still, at far greater cost. GLM-5.3 is an open model from Z.ai; it got 99.0%. Claude Opus 5.5 from Anthropic got 99.2%. They cost $809.08 and $1,707.74 per million tasks, and took about 2.5 to 2.9 seconds per reply. We ran them only as a reference point.

You can get close to them for far less. Jev's mistakes cluster on a few answers, so we asked GLM-5.3 again only when Jev gave one of three answers it gets wrong most often. That setup reached 98.1% at $86 per million tasks. With Tev first, the best setup reached 97.1% and sent 39% of tasks to the larger model, because Tev's mistakes spread over many answers. LLM cascade walks through the method. The trade between a decision model and a chat model has its own post: decision model vs LLM.

What are the limits of this test?

Our benchmark is one small test, run once, on one day. Read the numbers with these limits in mind:

  • Claude wrote all 400 tasks, and the same kind of model checked the labels. We spot-checked them. A careful person could still dispute an item or two.
  • 400 tasks is small. The headline gap is solid. Family scores rest on 50 tasks each and could move a lot.
  • Jev's timing includes our gateway hop; Tev's does not. Timing also depends on region and load.
  • We tested only pick-one questions. Jev's yes/no and score types have no Tev match here.
  • The two APIs count tokens differently, which is why the same list price gives different costs.
  • GLM-5.3 and Opus 5.5 ran through AI Space with their default reasoning settings.

Jev and Tev also missed the same 7 tasks. We reviewed each; three could be argued. We kept every label as written. Full details are in the methodology.

FAQ

Is Tev as good as Jev?

Not on our test. Jev got 97.2% of 400 new tasks right and Tev 90.0%. On hard tasks the gap grew to 95.9% against 84.1%. Tev matched Jev on simple support requests. It fell furthest behind on tasks that apply a rule, such as approving an AI agent's proposed action.

Why is Tev cheaper if both cost $0.042 per million tokens?

Jev bills about twice as many input tokens for the same text: 468 on average against Tev's 246. TypeSafe probably adds its own prompt around yours. Neither charges for output. So a million tasks cost $19.64 on Jev and $10.33 on Tev in our run.

Can I run Tev on my own servers?

Yes. Together published Tev's weights on Hugging Face and the training recipe on GitHub. Jev is a hosted service only. If your data cannot leave your own servers, Tev is the one of the two you can use.

Did Together publish accuracy numbers for Tev?

No. Together's 23 September 2026 launch post explained how Tev was trained and how to call it, with no accuracy scores. Our test on 25 September is the first independent one we know of. The data and code are public so anyone can rerun it.

Should I use Tev or a frontier model?

Choose by how costly a wrong answer is. GLM-5.3 and Claude Opus 5.5 scored 99.0% and 99.2% on the same tasks. They cost 78 to 165 times more per task than Tev and took over two seconds per reply. One setup runs Jev first and sends only its riskiest answers to a larger model; in our test that reached 98.1% at $86 per million tasks.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…