Jev and Laya explained: what a decision model is, and how it differs from an LLM
Two "System One" models launched within three days of each other this month: TypeSafe AI's Jev, a hosted API, and Convai's Laya, an open-weight copy of the same idea. Both are classifiers, not a new kind of LLM. That is a compliment. This explains what they do that an LLM cannot, what they cannot do that an LLM can, and where each belongs in a system.
Watch
Jev and Laya Explained: Decision Models vs LLMs
~11 min spoken. Keeps playing while you work in another tab.
On 15 September TypeSafe AI opened early access to Jev, which it calls its first "System One Model". Three days later Convai Innovations published Laya on Hugging Face under Apache 2.0: 421 million parameters, a compatible interface, and weights anyone can download. Within a week there were ONNX ports, a Node.js runtime, an MCP server and a crop of explainers calling both of them the end of hallucination.
The plainest accurate description is that they are classifiers: models that read an input and choose among answers you define in advance, with a probability for each. That is one of the oldest jobs in machine learning. What is new is how general the interface is, how much care has gone into the probabilities, and how much cheaper and faster the pair are than asking a chatbot to do the same job. Everything here is as it stood on 24 September 2026, from the vendors' own announcement and model card and the independent reviews listed at the foot of the page.
What an LLM does when you ask it to decide
A large language model writes. Given a prompt, it produces the next token, appends it, and runs again, one token at a time, until it decides to stop. Everything it can do (summarise, argue, code, translate) comes from that single loop.
When you use an LLM to decide something - is this comment abusive, which team should get this ticket, is this invoice a duplicate - you are borrowing that writing machine for a job with a fixed answer. You write instructions, list the allowed labels, ask for JSON, and parse what comes back. It works well enough that most teams do exactly this. It also has four costs that are easy to stop seeing:
- The answer's shape is a request, not a guarantee. The model can return malformed JSON, a label that is not on your list, or a paragraph of reasoning where you asked for one word. You write the validation and the retry.
- There is no real probability. You can ask the model how confident it is, but the number it gives you is more generated text, not a measurement.
- Every token of the answer is another pass through the model. A short JSON verdict is dozens of passes, and you pay for them as output tokens.
- You are renting a general-purpose model for a narrow job, and when its maker retires it, the job stops. We found this out in our own product this week, when a model our safety checks called was withdrawn and every check failed until we moved to its successor.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| 1. Ask | An LLM | Takes instructions, the text and the allowed labels as a prompt | Generation |
| 2. Generate | An LLM | Writes the reply one token at a time | Your parser |
| 3. Parse | Your code | Validates the reply; retries when it is malformed or off the list | The label, or a retry |
| 1. Ask | A decision model | Takes the state and the typed questions with their options | One forward pass |
| 2. Score | A decision model | Scores every option at once and returns a probability for each | Your threshold |
| 3. Decide | Your code | Acts on a confident answer, or passes the case on | The next step, or an escalation |
What a decision model does instead
Jev and Laya take two things: a state (the text or JSON you want judged) and a list of typed questions, each with its options fixed in advance. They read all of it at once and score every option together in a single pass, then return a probability for each. There is no second pass and nothing is written, so the answer is always one of the options you defined. There is nothing to parse, because there is no text.
TypeSafe defines three kinds of question, and Laya copies them:
- choice picks one option from a list, and returns the pick, a confidence, and a probability for every option.
- score places the state on an ordered rubric, such as urgency from 1 to 5, and returns the probability of each level.
- noul, TypeSafe's name for a yes-or-no question, returns the probability that a statement about the state is true.
That is where the "cannot hallucinate" line comes from, and it is worth reading carefully. It is true about the shape of the answer: the model cannot invent a sixth option when you gave it five, or return broken JSON. It is not true about the content. A decision model can still choose the wrong option, and do it with confidence. What it gives you that an LLM does not is a probability you can use: a number you can put a threshold on, so the uncertain cases go somewhere else.
Laya's model card describes the machinery. It is ModernBERT-large, a bidirectional encoder that reads the whole input at once rather than left to right, with a small decision head added on top: two transformer layers, a scorer for the options, and a head meant to say whether to act or escalate. It was trained with a method both companies call RLCD, reinforcement learning for calibrated decisions, which rewards the model for probabilities that match how often it turns out to be right. TypeSafe has published no architecture or parameter count for Jev.
Not a new idea, but a new package
Classifiers are how most production machine learning has always worked. Spam filters, fraud scores and sentiment models are classifiers, and since BERT in 2018 the standard recipe has been to take a pretrained encoder and fine-tune it on labelled examples of your task. Zero-shot classifiers, which choose among labels they were never trained on, have been in open-source libraries for years.
Three things are genuinely new here:
- The questions are set at call time. One model, and one request, can answer a routing question, a 1-to-5 urgency score and three yes-or-no policy checks about the same ticket, without a separately trained model for each.
- The probabilities are the product. A classifier's scores are often poorly calibrated, meaning a stated 90% is right far less than 90% of the time. RLCD is aimed squarely at that.
- It is sold like an LLM. An API priced per token, SDKs, and a pitch aimed at the developers who currently send these small decisions to a chatbot.
Jev and Laya side by side
| Jev | Laya | |
|---|---|---|
| Made by | TypeSafe AI | Convai Innovations |
| Released | 15 September 2026, early access | 18 September 2026 |
| How you get it | Hosted API only | Open weights, Apache 2.0, on Hugging Face |
| Size | Not disclosed | 421M (English, ModernBERT-large); 322M (multilingual, mmBERT-base) |
| Context | Not published | 512 tokens (English); 1,024 (multilingual and fine-tuned) |
| Price | $0.042 per million input tokens; output free | Free; you pay for the GPU |
| Speed, as the maker states it | 70-500 ms end to end, over the network | 39.5 ms (English) and 32.8 ms (multilingual) per question on one Tesla T4 |
The speed row compares two different measurements. Jev's figure includes the trip across the internet; Laya's is the model alone on your own card. Both are the makers' numbers.
The accuracy story is where the two separate, and Laya's own model card is candid about it:
- Out of the box, Laya is not a decision engine. On a 2,000-question typed-decisions benchmark the base checkpoints score 0.362 and 0.342 - above the 0.318 you would get by guessing, but below the 0.461 you would get by always giving the most common answer. The card calls them "near chance" zero-shot.
- Fine-tuned, it is very good. The typed-decisions checkpoint scores 0.766 on the same benchmark, ahead of the 0.727 published for Jev, but it was trained on that benchmark's own training split. It shows what fine-tuning buys, not what you get on day one. Convai also says its Jev figures are third-party numbers it did not measure itself.
- Many options are a weak spot. On Banking77, which asks for one of 77 banking intents, independent comparisons put Laya at 0.425 against Jev's 0.870. The model card itself warns about questions with more than about 20 options.
- Its confidence can fail silently. Laya's English base was trained on English. One review reports it scoring 0% on Khmer text while stating 95% confidence: it cannot read the script, and it does not know it cannot. A calibrated probability is only calibrated on the kind of data it was calibrated on.
- The confidence needs refitting on your data. The card reports a calibration error of 0.213 as shipped, falling to 0.081 after a temperature fit on the user's own examples. Plan on doing that fit before you trust a threshold.
In short, Jev is the one that works out of the box, and Laya is the one you can own, inspect, run anywhere and fine-tune - which, on narrow tasks with a few thousand labelled examples, is exactly what beats a general model. How that fine-tuning works, where the labelled examples come from, and the calibration step both models need are the subject of our follow-up on fine-tuning decision models.
What it costs
At Jev's price, a decision about a 1,000-token input costs $0.000042: a million such decisions cost $42, and the answers are free because nothing is generated. Laya costs whatever the GPU costs. On one Tesla T4, the card's figures work out to about 25 to 30 questions a second one at a time, and published figures for batched questions run from 103 to 332 a second. At 421 million parameters the weights are under a gigabyte, small enough for the smallest data-centre cards.
The cost that matters most is the one being replaced. A decision that currently goes to a general-purpose LLM pays for the prompt, the instructions, the list of labels and every token of the JSON verdict, and waits for the whole answer to be written. For high-volume, narrow judgments, a decision model is cheaper on every one of those lines. At low volume the difference is too small to justify a second model in the stack.
Where each belongs
A decision model is the right tool when the answer is a choice among things you can name in advance, the input is short, and there are many of them:
- routing tickets, emails and support chats;
- spam, phishing and moderation of short text;
- policy checks and guardrails in an agent loop, such as "is this tool call allowed";
- scoring or tagging large volumes of records, the "smart if-statement" TypeSafe pitches it as.
It is the wrong tool when the job is to write, explain or reason at length; when the input is long (Laya reads 512 to 1,024 tokens, so a long email thread or contract will not fit); when the input is an image, audio or video, which neither model takes; when the answer is one of hundreds of labels; or when the language is one the model was not trained on.
Both vendors, and most of the people who have tested them, land on the same design: a decision model first, and an LLM behind it for the cases it is unsure of.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| 1. Decide | A decision model (Jev or Laya) | Answers the typed questions with a probability for each option | Action if confident; the LLM if not |
| 2. Escalate | An LLM | Reads the item in full and reasons about it | Action, or a person for the hardest cases |
| 3. Review | A person | Settles what neither model should decide alone | Action |
The threshold is the whole design. Set it from your own labelled examples, after refitting the confidence on them, and measure what share of traffic each path takes. If most items clear the threshold, the expensive model only ever sees the hard cases. If few do, you have added a step without removing one.
Our own content moderation, for Journal Mosaic, is a fair test of that. It checks shared journal entries, photos, video and audio, and the comments under them. The comments are short text with a fixed set of verdicts - a decision model's home ground - and could move to one. The entries run longer than Laya's context, and the photos, video and audio need a model that can see and hear. So the realistic version for us is the diagram above: a decision model for the comments, and an LLM for everything it cannot read.
That is a fair summary of the whole category. Jev and Laya are not a replacement for LLMs, and they are not a new kind of intelligence. They are a fast, cheap, well-calibrated classifier with an unusually general interface, and much of what teams currently send to a chatbot was a classification problem all along.