Fine-tuning a decision model: what you can change in Jev and Laya, and the part that is actually hard
Jev cannot be fine-tuned at all, and Laya can be fine-tuned on a free GPU in an afternoon. Neither fact is the hard part. The hard part is the labelled examples: where they come from, how they are split, and what their probabilities are fitted on. Laya's own training notebook got that last step wrong until this month.
~10 min spoken. Keeps playing while you work in another tab.
Our explainer on Jev and Laya ended on a pattern: a decision model settles the confident cases, and an LLM handles the rest. That pattern only pays if the decision model is right often enough, on your data, and knows when it is not. Out of the box, neither is guaranteed. Laya's base checkpoints score below the most-common-answer baseline on the benchmark its makers chose, and Jev's calibration is being argued about in public.
So this is about the work between downloading a decision model and trusting it: what each one lets you change, what the examples look like, where they come from, and the mistakes that make a model look better than it is. Everything is as it stood on 24 September 2026.
Three levers, and which model has which
People say "training" for three different things, and it helps to keep them apart, because they cost different amounts and change different things.
| Lever | What it changes | What it needs | Jev | Laya |
|---|---|---|---|---|
| Question design | Which answer it picks | Your judgment, and a test set to check it | Yes, the main lever | Yes |
| Recalibration | How far to trust its confidence | A few hundred labelled examples | Yes, on your side | Yes, built into the tools |
| Fine-tuning | Which answer it picks, and how well | Thousands of labelled decisions and a GPU | No | Yes |
The first two leave the model alone. Only the third changes its weights.
Jev: you shape the questions, not the model
TypeSafe serves the same weights to every account, and there is no fine-tuning endpoint. Its advice, as a review of the options quotes it, is to "shape answers through the request": put your domain rules in each question's instructions and in the description of each option, break a broad judgment into small questions, and combine the answers in your own code.
That last point does most of the work. "Should this refund be approved?" is one hard question. "Is the order older than 30 days?", "Does the customer say the item arrived damaged?" and "Has this customer had a refund this month?" are three easy ones, and your code can apply the policy to the answers. A decision model is much better at narrow questions than broad ones, and your policy lives in code you control instead of a prompt you hope it follows.
For everything else, the same review says TypeSafe's documentation points at training a small model of your own downstream of Jev: feed Jev's probabilities in as features, and fit a model on your labelled outcomes. You are not changing Jev; you are learning how to read it.
Recalibration: the step nobody should skip
A model is calibrated when its confidence means what it says: of all the answers it gives at 0.8, about 80% are right. Both vendors train for this, with the method they call RLCD, which rewards probabilities that match how often the model turns out to be right.
The catch is in "turns out to be right" - on which data? A critique published this week argues that Jev cannot be calibrated for everyone at once: calibration holds on a particular mix of data, and a single model returning the same probabilities to every customer cannot match every customer's mix. Its author reports Jev giving 0.92 for heads on a coin the prompt said was fair. The same failure shows up in Laya, which one review reports scoring 0% on Khmer text at 95% confidence. A probability is only calibrated on data like the data it was calibrated on.
The fix is cheap and it is the same for both: fit the confidence to your own data. Take a few hundred examples you have labelled, run the model on them, and fit a simple correction from the model's scores to how often it was actually right. For a yes-or-no question that is usually Platt scaling, a small logistic curve. For multi-option questions it is temperature scaling, one number per question type that softens or sharpens every probability. Neither changes which answer the model picks. Both change how far you should trust it, and that is what your threshold is built on. The critique's author suggests a few hundred labelled examples can be enough.
Laya: fine-tuning on a free GPU
Laya is the one you can actually retrain, and Convai publishes the recipe as a notebook that runs on Kaggle's free pair of T4 GPUs. It fine-tunes the 421M English checkpoint on the training split of a public benchmark, typed-decisions: 1,200 cases holding 6,000 decisions across four workflows - agent-trace observability, customer service, invoice processing and security incidents. It then evaluates on the benchmark's separate 400-case test split. Convai's model card puts the training at four to five hours; the current notebook says the training step itself now takes minutes. Either way, compute is not the obstacle.
The result is the jump our explainer quoted: from 0.362 accuracy for the base model to 0.766 after fine-tuning, on 2,000 decisions it had not trained on. That is a big improvement from a modest amount of data, and it is the strongest argument for Laya.
Each training case has the same shape as a request at run time: a state (a JSON record of the ticket, invoice or trace), a set of questions, each with a type and a short description of every option, and the answer for each question. The notebook trains with RLCD: the model proposes a probability for each option, is rewarded by a proper scoring rule for how close that was to the right answer, and is also pulled towards the labeller's own probabilities. Then it fits one temperature per question type, and saves the model.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| 1. Collect and label | Your data, a teacher LLM, or people | Gathers real inputs and records the right answer to each typed question | The split |
| 2. Split | Your code | Holds back a calibration set and a test set before any training | Fine-tuning, with the training set only |
| 3. Fine-tune | Laya (Jev cannot be fine-tuned) | Retrains on the training set; with Jev, the questions are improved instead | Calibration |
| 4. Calibrate | Your code | Fits temperature or Platt scaling on the calibration set | Threshold choice |
| 5. Threshold and run | Your code | Picks the threshold on the test set and runs the two-stage pattern | New labelled examples from escalated cases |
Where the labels come from
This is the hard part, and it is worth being honest about what the benchmark's labels are. The model card says they came from a teacher: a reference model that answered every question, whose answers became the ground truth. The card also reports that the teacher agreed with itself only 73.5% of the time. Laya's 0.766 therefore means it matched a noisy teacher more often than the teacher matched itself - a sign it learned the task's pattern rather than the teacher's noise, but not a measure of how often it is right.
For your own task, labels come from three places, usually in this order:
- Decisions you have already made. Tickets that were routed and resolved, comments a moderator removed or kept, invoices that were paid or disputed. This is the best data you have, because the label is what actually happened. Check that past decisions were right, and that the mix of cases still looks like today's.
- A teacher model. Have a strong LLM answer your typed questions over a large sample of real inputs, and train on its answers. This is how the benchmark was built, and it is how most teams without years of history will start. The student inherits the teacher's mistakes, so the teacher should be the best model you can afford, and asked the same narrow questions the student will be.
- People. Slow and expensive, and irreplaceable for one set: the examples you test on. A test set labelled by the same teacher as the training set can only tell you how well the student copies the teacher.
And once the two-stage pattern is running, it makes labels on its own. Every case the decision model sends up to an LLM, and every case the LLM sends to a person, comes back with an answer. Those are exactly the hard cases, and they are the most valuable examples you can feed the next round of training.
The mistakes that flatter a model
Fitting the confidence on the training data. Until this month, Laya's own notebook fitted its temperatures on a slice of the training set. A model that has just been trained on those examples is already confident and right about them, so the fitted temperatures barely moved: [1.0148, 1.0374, 1.0575], all within 6% of doing nothing. The card for the published fine-tune now says to "treat its confidence as uncalibrated" until it is refitted on held-out data, and the notebook has been fixed to hold a slice back. If Convai's own team made this mistake in public, assume you will too, and split your data three ways before you start: train, calibrate, test.
Testing on the benchmark you trained on. The 0.766 is on the benchmark's own test split, drawn from the same four workflows as its training data. The card warns that the checkpoint is specialised to those synthetic workflows and should not be dropped in as a general default. The only number that matters is on your data.
A test set too small to tell you anything. At 80% accuracy on 300 test examples, the uncertainty is about plus or minus 4.5 percentage points. That is fine for deciding whether a model works at all, and too coarse to tell two close fine-tunes apart. Count your test examples per question, not per case.
Many options at once. Laya's card warns that accuracy falls off above about 20 options. If a question has 77 possible answers, fine-tuning helps less than splitting it: a question for the department, then one for the intent within it.
Choosing the threshold
Only after calibration does the threshold mean anything. Run the calibrated model over the test set, and for each candidate threshold look at two numbers: the share of cases that clear it, and how often those cases are right. Raising the threshold makes the decision model more accurate on fewer cases, and sends more of them to the LLM behind it. Pick the point where the accuracy on the cleared cases is what your business can live with, and read off how much LLM traffic that leaves. Then keep measuring, because the mix of your data will move and the fit will drift with it.
Is it worth it for you?
The work is front-loaded: the labelled examples, the three-way split, the calibration fit and the threshold. It pays when the decision runs thousands of times a day and each run to an LLM costs real money or real latency.
Our own case is a useful counterweight. Journal Mosaic's comment moderation is exactly the kind of narrow, short-text decision Laya is built for, and we already have a teacher: the LLM that checks every comment today returns a verdict a smaller model could learn from. What we do not have yet is the volume to justify a second model to run, calibrate and monitor. So for now the LLM stays. The first step we would take, well before training anything, is to keep each verdict alongside the comment and what a moderator later decided about it. That is probably the right way for most teams to start: record your decisions in a shape a decision model could learn from, long before you train one.