All videos

Video briefing

Jev and Laya Explained: Decision Models vs LLMs

Jev · Laya 5:28

TypeSafe AI's Jev and Convai's Laya bring dedicated decision models to production pipelines. Here is how they differ from general LLMs and where they belong.

Transcript

What the video says

Jev and Laya explained: what a decision model is, and how it differs from an LLM. From TerraNet Technologies. Quick thing first: if this is useful, like the video and subscribe. When you ask a large language model to route a ticket or verify policy, you borrow a generative engine for a closed question. Two models, TypeSafe AI's Jev and Convai's Laya, approach the problem differently: as pure classifiers delivering true probabilities without generating text. Instead of running an auto-regressive writing loop that predicts token after token, these architectures ingest the input and evaluate predefined options simultaneously. This distinction marks a shift toward purpose-built discriminative evaluation, sidestepping the overhead, unpredictable parsing errors, and synthetic confidence scores typical of general chatbot workflows.

Figure 1 maps the steps taken across both execution paths. The upper flow traces an LLM taking instructions, text, and labels as a prompt, generating output token by token, and handing off to custom code to validate the reply and retry malformed formatting. The lower flow maps a decision model ingesting the state alongside typed questions in a single forward pass, scoring all options concurrently to deliver genuine calibrated probabilities.

The critical contrast for system architects is where uncertainty gets handled: rather than repairing broken text structures after generation, decision models immediately hand off calibrated probabilities to simple threshold logic, removing validation loops entirely.

TypeSafe AI defines three primary question formats that Laya mirrors across its own runtime. Choice selects a single item from an explicit list, returning probabilities across every candidate. Score evaluates states along an ordered rubric, assigning probabilities across discrete levels.

Noul resolves binary true-or-false policy statements. Because the candidate answers are constrained at invocation time, the model cannot invent outside labels or produce broken JSON. While this guarantees the structural validity of every answer, the model can still confidently select the wrong outcome, making thresholding essential.

Classifiers have operated in production since the arrival of BERT, but Jev and Laya introduce three practical shifts. First, questions are defined at call time: a single model call can simultaneously evaluate an urgency score, a routing label, and three policy checks without maintaining separate dedicated models. Second, the probabilities represent the actual product, using reinforcement learning calibration to ensure stated confidence reflects empirical reliability. Third, they package traditional classification into accessible developer tooling, offered either as a per-token managed API or open-weight models.

Renting a general generative model forces you to pay for prompt tokens, instruction formatting, and every returned token of JSON. When the vendor updates or retires that model, the workflow breaks. TypeSafe AI's Jev charges forty-two thousandths of a dollar per million input tokens with free output, running end to end over the network in seventy to five hundred milliseconds.

Convai Innovations' Laya runs on your own hardware, processing single queries in thirty-two point eight to thirty-nine point five milliseconds on a Tesla T4 card. Batching questions pushes throughput between one hundred three and three hundred thirty-two queries per second.

Laya packs ModernBERT-large into four hundred twenty-one million parameters with a two-layer transformer decision head. Out of the box, Laya's base checkpoints score zero point three six two on a two-thousand question typed-decisions benchmark, which Convai notes is near chance zero-shot against a guessing baseline of zero point three one eight and a common-answer baseline of zero point four six one. Fine-tuned on the benchmark split, Laya reaches zero point seven six six, passing the zero point seven two seven third-party figure published for Jev. Reaching production-grade reliability therefore requires domain-specific fine-tuning.

Figure 2 details the two-stage architecture across its three tiers. Step one routes every item to a decision model like Jev or Laya to output typed probabilities. Step two routes cases that fall below the confidence threshold to an LLM, which reasons over the full text. Step three escalates the most ambiguous cases to a human reviewer. The takeaway is simple: decision models handle routine high-volume judgments, leaving generative models only for uncertain inputs.

Decision models excel on short texts with fixed outcomes: routing inbound tickets, screening comments, and evaluating agent guardrails. They fail when tasks require long-form reasoning, multimodal inputs, or inputs exceeding Laya's five hundred twelve to one thousand twenty-four token window.

The entire architecture hinges on your confidence threshold. Calibrate that threshold against verified production data. If the bulk of your traffic clears it, expensive model calls drop dramatically; if not, you have simply introduced extra latency.

Explore the full technical breakdown, benchmark scores, and deployment considerations for Jev and Laya on the TerraNet Technologies blog.

Produced by TerraNet Technologies from the cited evidence behind the written article. Facts can change after the recorded date.