AI ComparisonsOct 5, 20269 min read

Yev: An Open Source Jev Alternative at 94.2% on 400 Tasks

Yev is an open source Jev alternative you run on your own GPU. It got 94.2% of DecideBench's 400 tasks right vs Jev's 98.0%, with its training code public.

On this page
  1. What is Yev?
  2. How does Yev compare with Jev, Tev and other open models?
  3. Where does Yev fall short?
  4. How was Yev trained?
  5. How do you run Yev yourself?
  6. When should you pick Yev over Jev?
  7. FAQ
  8. Sources

Yev is an open source Jev alternative from XY Space, released on 4 October 2026. Jev is TypeSafe's hosted decision model: you send it text, a question and a short list of answers, such as whether to approve a refund, and it picks one. Yev does that job on your own GPU. On DecideBench, our public test of 400 such decisions, Yev got 94.2% right and Jev 98.0%.

Is your data not allowed to leave your servers? Yev is a model you can run there.

Do you need the highest accuracy, and can you call an outside API? Jev is still ahead.

Do you feed it long cases, such as a full invoice packet? Wait for the next Yev model.

  • It is a decision model with 4 billion parameters. It picks one answer from a list you give it and writes no other text.
  • The weights and code are public and free for commercial use.
  • It beats Together AI's open model Tev (92.8%). It sits three tasks behind imajev-4b (95.0%), another open model, a gap too small to count.
  • Running it costs more than calling Jev: $58.10 per million decisions on one rented graphics card, against $32.26 for Jev's API. You pick it for control, not price.
  • Long, many-step cases are its weak spot. It scored 51.5% on WorkflowEvals, a test of long business workflows, where Jev scored 67.8%.

What is Yev?

Yev is a System One model you can download and run yourself. TypeSafe coined "System One" for Jev when it launched on 15 September 2026: a model that makes a fast, structured decision your software can use directly, instead of writing prose. New to the idea? Start with what is Jev.

You give it three things: a piece of text (the state), a question, and between two and six options. It answers one of three question types:

  • Choice: pick one of the options.
  • Noul: yes or no.
  • Score: a point on an ordered scale, such as ticket severity from 1 to 4.

The model answers in a single pass. It reads how likely each option's letter is as the next token, so nothing has to be generated or parsed.

Each question type has its own calibration setting, fitted on 2,000 rows the model never trained on. Calibrated means a 90% answer should be right about 90% of the time.

The name has a version in it. The "0" in yev0-4b is the training stage. This is the Stage 0 model, and a Stage 1 model trained on longer text is in progress.

How does Yev compare with Jev, Tev and other open models?

Yev lands ahead of Tev, level with imajev-4b, and 3.8 points behind Jev. imajev-4b also costs less to run.

Right answers on DecideBench's 400 tasks
Right answers on DecideBench's 400 tasks
AccuracyBoth halves of a pair right
Tev (self-hosted)92.8%86%
yev0-4b94.2%88.5%
Clef 27B94.8%89.5%
imajev-4b95%90.5%
Jev98%96%

Source: DecideBench v1.1 leaderboard, results measured 28 September to 4 October 2026

DecideBench is our public benchmark of 400 decisions in 200 pairs. The two halves of a pair differ by one small edit that changes the right answer, such as a return request one day past the window. A model that matches on topic words gets one half wrong. How to test a classifier explains why we built it that way.

ModelWho runs itAccuracyPair accuracyCost per million tasksMedian wait
JevTypeSafe, hosted98.0%96.0%$32.26639 ms
imajev-4bYou95.0%90.5%$28.57499 ms
Clef 27BCloudflare, hosted94.8%89.5%$130.85811 ms
yev0-4bYou94.2%88.5%$58.101,058 ms
TevYou92.8%86.0%$46.82823 ms

Self-hosted rows ran on one NVIDIA L4, a data-centre graphics card, priced at $0.81 an hour. Their wait times include no network.

The benchmark sent four requests at once. Our server answers one at a time, using the plain transformers library, so each request waited behind the others. That is most of why Yev costs more than imajev-4b.

Without worked examples, Yev does slightly better. The model card reports 95.0% with no examples, against Tev's 90.0%. That gap is too large to be chance. The model card's own run with examples gave 94.5%, one task more than the leaderboard run on a different graphics card.

By task family, it got every agent-action review and every agent-routing task right. Its weakest families were content moderation (86%) and review sentiment (88%). Each family has only 50 tasks, so treat those two numbers as directional. Jev vs Tev covers the hosted pair in more depth.

Where does Yev fall short?

Yev struggles with long cases, many-step workflows and arithmetic on dates. These are the cases where we would still route work to Jev or a larger model.

TestYevComparison
WorkflowEvals, overall51.5%Jev 67.8%
WorkflowEvals, invoice processing (exact action set)46.7%Jev 61.78%
JevBench public items, hard tier55.9%imajev-4b 72.07%
JevBench hard tier, time and number questions1 of 15

WorkflowEvals gives the model one long case, such as an invoice packet several times longer than anything Stage 0 trained on, and asks many questions about it. Stage 0 trained on short cases with at most six options. Most WorkflowEvals questions fall outside both limits. On invoices, it picked the main action correctly 75.8% of the time but got the full set of actions right only 46.7% of the time.

Other limits from the model card:

  • No more than six options. The server accepts more, but the model never trained on them.
  • English only.
  • Not a safety guard. It scored no better than chance on a public test of judging whether an agent's actions are safe.
  • Not for decisions about a person without a human reviewing them: medical, legal, credit or hiring.
  • The synthetic labels had no human check. About 9,600 training rows were written by Claude Opus 5.5, and separate runs of that model checked them blind. Rule-generated rows, whose labels are computed, partly offset that.

How was Yev trained?

Yev starts from the open Qwen3.5-4B-Base model. We trained a small set of extra weights on top (a LoRA fine-tune) using 44,459 decision examples. Training took 4.71 hours on one NVIDIA L40S graphics card. All six training runs, including the ones we did not ship, took 25.3 hours of graphics-card time.

Training dataRows
Public datasets (21 sources, permissive licences only)20,013
Rule-generated cases, labels computed from written rules9,884
Cases written by Claude Opus 5.5, checked blind9,562
Decisions about public agent skills5,000

The rule-generated rows test the edges a policy cares about. A 30-day return window gets cases on the last day, and one day either side. The Opus rows come in small families: one base case, then variants that each change about 15 tokens to flip the answer. A fresh Opus context then had to name the edit that flipped it, or the pair was dropped.

DecideBench never touched training. Every row was checked against all 697 DecideBench items, for shared eight-word runs and for near-identical meaning. A final check found none of the 400 test items in the training file. We read the test once, at the end.

We picked the shipped run by a rule fixed before the experiments. One variant scored better on our own held-out data but worse on DecideBench's practice set, so it lost. The repository has the configs for all six runs.

How do you run Yev yourself?

Clone the repository, download the calibration file, and start yev serve. It downloads the weights on first run. For useful speed you need an NVIDIA graphics card.

git clone https://github.com/xyspacedev/yev && cd yev
uv sync --extra train --extra serve
uv run python -c "from huggingface_hub import hf_hub_download; hf_hub_download('choyiny/yev0-4b', 'calibration.json', local_dir='.')"
uv run --extra train --extra serve yev serve \
  --model choyiny/yev0-4b --calibration calibration.json --port 8000

The server offers two endpoints. The /v1/systemone endpoint takes the same request shape as Jev, so existing code needs little change. The /v1/chat/completions endpoint works with any OpenAI-style client and replies with one letter.

curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": "I was charged twice. Please help.",
  "questions": {"billing": {"instructions": "Is this message about billing?", "type": "noul"}}}'

Three details matter in practice:

  • Ask several questions about one text in one request. The server encodes the shared text once and reuses it for each question.
  • Install `flash-linear-attention` and `causal-conv1d`. Without them, the model falls back to a much slower kernel.
  • The server listens on 127.0.0.1 and checks no authentication. Put it behind your own gateway before exposing it.

When should you pick Yev over Jev?

Pick Yev when the text cannot leave your network, when you want to fine-tune your own copy, or when you need to inspect how the model was built. Pick Jev when accuracy matters most and a hosted API is acceptable.

The probabilities are useful in their own right. On our development set, answers with confidence of 0.9 or higher were right 99.3% of the time, and they covered 80.8% of items. So you can act on the confident answers and send the rest to a person or a larger model. LLM cascade walks through that setup with Jev, and the same method works here.

If you are weighing a decision model against a general chat model, decision model vs LLM puts numbers on that trade. Other open options launched recently too. Cloudflare released Clef and Clef-flash as open weights on 1 October 2026.

XY Space builds agentic workflows that make decisions like these inside real operations. Yev gives those workflows a decision model that runs inside a client's own network.

FAQ

Is Yev free to use?

Yes. The weights and code are under the Apache-2.0 licence, which allows commercial use. You pay only for the hardware. On one rented graphics card at $0.81 an hour, our benchmark run came to $58.10 per million decisions. The public datasets in the training mix keep their own licences, listed on the model page.

Is Yev better than Tev?

On DecideBench, yes, by a small margin. With worked examples it got 94.2% against Tev's 92.8%, a gap that could be chance. Without examples it got 95.0% against 90.0%, which is not. Tev is cheaper to self-host, at $46.82 per million tasks against $58.10.

Can Yev replace Jev?

For short texts with up to six options, it comes within four points of Jev on DecideBench. For long cases it does not. It scored 51.5% on WorkflowEvals, against Jev's 67.8%. A common setup keeps Yev for the confident answers and sends the rest to Jev or a larger model.

What hardware does Yev need?

One graphics card is enough. We trained it on an NVIDIA L40S with 48 gigabytes of memory and benchmarked it on a smaller NVIDIA L4 with 24 gigabytes. It also runs on an ordinary processor, slowly. With an NVIDIA card, install flash-linear-attention and causal-conv1d for the fast code path.

What is next for Yev?

A Stage 1 model is in training, with longer cases and questions with more options. That targets the long-workflow gap on WorkflowEvals. The pipeline will also extend its contamination filter to every benchmark we report, not only DecideBench.

Sources

Written by

Cho Yin Yong

Principal AI Solutions Engineer, XY Space

Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.

More from Cho Yin Yong

Share this article

Work with us

We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.

Start a project
Work with us

Book a call.We'll come back with specifics.

Start with the map of your organization, or with the one job that hurts. Measured in hours and money, and everything we build stays yours.

Loading form…