On this page
Yev is an open source Jev alternative from XY Space, released on 4 October 2026. Jev is TypeSafe's hosted decision model: you send it text, a question and a short list of answers, such as whether to approve a refund, and it picks one. Yev does that job on your own GPU. On DecideBench, our public test of 400 such decisions, Yev got 94.2% right and Jev 98.0%.
Is your data not allowed to leave your servers? Yev is a model you can run there.
Do you need the highest accuracy, and can you call an outside API? Jev is still ahead.
Do you feed it long cases, such as a full invoice packet? Wait for the next Yev model.
- It is a decision model with 4 billion parameters. It picks one answer from a list you give it and writes no other text.
- The weights and code are public and free for commercial use.
- It beats Together AI's open model Tev (92.8%). It sits three tasks behind imajev-4b (95.0%), another open model, a gap too small to count.
- Running it costs more than calling Jev: $58.10 per million decisions on one rented graphics card, against $32.26 for Jev's API. You pick it for control, not price.
- Long, many-step cases are its weak spot. It scored 51.5% on WorkflowEvals, a test of long business workflows, where Jev scored 67.8%.
What is Yev?
Yev is a System One model you can download and run yourself. TypeSafe coined "System One" for Jev when it launched on 15 September 2026: a model that makes a fast, structured decision your software can use directly, instead of writing prose. New to the idea? Start with what is Jev.
You give it three things: a piece of text (the state), a question, and between two and six options. It answers one of three question types:
- Choice: pick one of the options.
- Noul: yes or no.
- Score: a point on an ordered scale, such as ticket severity from 1 to 4.
The model answers in a single pass. It reads how likely each option's letter is as the next token, so nothing has to be generated or parsed.
Each question type has its own calibration setting, fitted on 2,000 rows the model never trained on. Calibrated means a 90% answer should be right about 90% of the time.
The name has a version in it. The "0" in yev0-4b is the training stage. This is the Stage 0 model, and a Stage 1 model trained on longer text is in progress.
How does Yev compare with Jev, Tev and other open models?
Yev lands ahead of Tev, level with imajev-4b, and 3.8 points behind Jev. imajev-4b also costs less to run.
| Accuracy | Both halves of a pair right | |
|---|---|---|
| Tev (self-hosted) | 92.8% | 86% |
| yev0-4b | 94.2% | 88.5% |
| Clef 27B | 94.8% | 89.5% |
| imajev-4b | 95% | 90.5% |
| Jev | 98% | 96% |
Source: DecideBench v1.1 leaderboard, results measured 28 September to 4 October 2026
DecideBench is our public benchmark of 400 decisions in 200 pairs. The two halves of a pair differ by one small edit that changes the right answer, such as a return request one day past the window. A model that matches on topic words gets one half wrong. How to test a classifier explains why we built it that way.
| Model | Who runs it | Accuracy | Pair accuracy | Cost per million tasks | Median wait |
|---|---|---|---|---|---|
| Jev | TypeSafe, hosted | 98.0% | 96.0% | $32.26 | 639 ms |
| imajev-4b | You | 95.0% | 90.5% | $28.57 | 499 ms |
| Clef 27B | Cloudflare, hosted | 94.8% | 89.5% | $130.85 | 811 ms |
| yev0-4b | You | 94.2% | 88.5% | $58.10 | 1,058 ms |
| Tev | You | 92.8% | 86.0% | $46.82 | 823 ms |
Self-hosted rows ran on one NVIDIA L4, a data-centre graphics card, priced at $0.81 an hour. Their wait times include no network.
The benchmark sent four requests at once. Our server answers one at a time, using the plain transformers library, so each request waited behind the others. That is most of why Yev costs more than imajev-4b.
Without worked examples, Yev does slightly better. The model card reports 95.0% with no examples, against Tev's 90.0%. That gap is too large to be chance. The model card's own run with examples gave 94.5%, one task more than the leaderboard run on a different graphics card.
By task family, it got every agent-action review and every agent-routing task right. Its weakest families were content moderation (86%) and review sentiment (88%). Each family has only 50 tasks, so treat those two numbers as directional. Jev vs Tev covers the hosted pair in more depth.
Where does Yev fall short?
Yev struggles with long cases, many-step workflows and arithmetic on dates. These are the cases where we would still route work to Jev or a larger model.
| Test | Yev | Comparison |
|---|---|---|
| WorkflowEvals, overall | 51.5% | Jev 67.8% |
| WorkflowEvals, invoice processing (exact action set) | 46.7% | Jev 61.78% |
| JevBench public items, hard tier | 55.9% | imajev-4b 72.07% |
| JevBench hard tier, time and number questions | 1 of 15 |
WorkflowEvals gives the model one long case, such as an invoice packet several times longer than anything Stage 0 trained on, and asks many questions about it. Stage 0 trained on short cases with at most six options. Most WorkflowEvals questions fall outside both limits. On invoices, it picked the main action correctly 75.8% of the time but got the full set of actions right only 46.7% of the time.
Other limits from the model card:
- No more than six options. The server accepts more, but the model never trained on them.
- English only.
- Not a safety guard. It scored no better than chance on a public test of judging whether an agent's actions are safe.
- Not for decisions about a person without a human reviewing them: medical, legal, credit or hiring.
- The synthetic labels had no human check. About 9,600 training rows were written by Claude Opus 5.5, and separate runs of that model checked them blind. Rule-generated rows, whose labels are computed, partly offset that.
How was Yev trained?
Yev starts from the open Qwen3.5-4B-Base model. We trained a small set of extra weights on top (a LoRA fine-tune) using 44,459 decision examples. Training took 4.71 hours on one NVIDIA L40S graphics card. All six training runs, including the ones we did not ship, took 25.3 hours of graphics-card time.
| Training data | Rows |
|---|---|
| Public datasets (21 sources, permissive licences only) | 20,013 |
| Rule-generated cases, labels computed from written rules | 9,884 |
| Cases written by Claude Opus 5.5, checked blind | 9,562 |
| Decisions about public agent skills | 5,000 |
The rule-generated rows test the edges a policy cares about. A 30-day return window gets cases on the last day, and one day either side. The Opus rows come in small families: one base case, then variants that each change about 15 tokens to flip the answer. A fresh Opus context then had to name the edit that flipped it, or the pair was dropped.
DecideBench never touched training. Every row was checked against all 697 DecideBench items, for shared eight-word runs and for near-identical meaning. A final check found none of the 400 test items in the training file. We read the test once, at the end.
We picked the shipped run by a rule fixed before the experiments. One variant scored better on our own held-out data but worse on DecideBench's practice set, so it lost. The repository has the configs for all six runs.
How do you run Yev yourself?
Clone the repository, download the calibration file, and start yev serve. It downloads the weights on first run. For useful speed you need an NVIDIA graphics card.
git clone https://github.com/xyspacedev/yev && cd yev
uv sync --extra train --extra serve
uv run python -c "from huggingface_hub import hf_hub_download; hf_hub_download('choyiny/yev0-4b', 'calibration.json', local_dir='.')"
uv run --extra train --extra serve yev serve \
--model choyiny/yev0-4b --calibration calibration.json --port 8000The server offers two endpoints. The /v1/systemone endpoint takes the same request shape as Jev, so existing code needs little change. The /v1/chat/completions endpoint works with any OpenAI-style client and replies with one letter.
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice. Please help.",
"questions": {"billing": {"instructions": "Is this message about billing?", "type": "noul"}}}'Three details matter in practice:
- Ask several questions about one text in one request. The server encodes the shared text once and reuses it for each question.
- Install `flash-linear-attention` and `causal-conv1d`. Without them, the model falls back to a much slower kernel.
- The server listens on 127.0.0.1 and checks no authentication. Put it behind your own gateway before exposing it.
When should you pick Yev over Jev?
Pick Yev when the text cannot leave your network, when you want to fine-tune your own copy, or when you need to inspect how the model was built. Pick Jev when accuracy matters most and a hosted API is acceptable.
The probabilities are useful in their own right. On our development set, answers with confidence of 0.9 or higher were right 99.3% of the time, and they covered 80.8% of items. So you can act on the confident answers and send the rest to a person or a larger model. LLM cascade walks through that setup with Jev, and the same method works here.
If you are weighing a decision model against a general chat model, decision model vs LLM puts numbers on that trade. Other open options launched recently too. Cloudflare released Clef and Clef-flash as open weights on 1 October 2026.
XY Space builds agentic workflows that make decisions like these inside real operations. Yev gives those workflows a decision model that runs inside a client's own network.
FAQ
Is Yev free to use?
Yes. The weights and code are under the Apache-2.0 licence, which allows commercial use. You pay only for the hardware. On one rented graphics card at $0.81 an hour, our benchmark run came to $58.10 per million decisions. The public datasets in the training mix keep their own licences, listed on the model page.
Is Yev better than Tev?
On DecideBench, yes, by a small margin. With worked examples it got 94.2% against Tev's 92.8%, a gap that could be chance. Without examples it got 95.0% against 90.0%, which is not. Tev is cheaper to self-host, at $46.82 per million tasks against $58.10.
Can Yev replace Jev?
For short texts with up to six options, it comes within four points of Jev on DecideBench. For long cases it does not. It scored 51.5% on WorkflowEvals, against Jev's 67.8%. A common setup keeps Yev for the confident answers and sends the rest to Jev or a larger model.
What hardware does Yev need?
One graphics card is enough. We trained it on an NVIDIA L40S with 48 gigabytes of memory and benchmarked it on a smaller NVIDIA L4 with 24 gigabytes. It also runs on an ordinary processor, slowly. With an NVIDIA card, install flash-linear-attention and causal-conv1d for the fast code path.
What is next for Yev?
A Stage 1 model is in training, with longer cases and questions with more options. That targets the long-workflow gap on WorkflowEvals. The pipeline will also extend its contamination filter to every benchmark we report, not only DecideBench.
Sources
- XY Space, yev0-4b model card, 4 October 2026
- XY Space, yev code repository, 4 October 2026
- XY Space, DecideBench dataset and results, v1.1, accessed 5 October 2026
- XY Space, DecideBench leaderboard, accessed 5 October 2026
- TypeSafe, Introducing System One Models and Jev, 15 September 2026
- TypeSafe, WorkflowEvals, accessed 5 October 2026
- Cloudflare, Introducing Clef: our open-source decision models, 1 October 2026
- Qwen, Qwen3.5-4B-Base, accessed 5 October 2026
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project