Editorial illustration for CLM-8B Decouples State-Action Scoring to Cut Decision Model Latency
AI analysis / Latest briefings
Release analysis · TerraNet Intelligence

CLM-8B Decouples State-Action Scoring to Cut Decision Model Latency

Contrastive-LM has released CLM-8B, an open-source decision model that scores candidate actions against a state instead of generating them. Its authors claim parity with Jev at up to 9× lower latency; here is what the repository shows and what it does not yet.

By TerraNet Intelligence9 min read3 sources
Editorial illustration for CLM-8B Decouples State-Action Scoring to Cut Decision Model Latency
Contrastive-LM
CLM-8B
System One Model
Decision Models
Qwen3-8B
Agentic Benchmarks
Apache 2.0
Listen to this article

~9 min spoken. Keeps playing while you work in another tab.

On September 23, 2026, the Contrastive-LM open-source project published the initial release of CLM-8B, a specialized "System One" decision model designed for rapid state-action scoring Source 1 · GitHub. Accompanied by serving utilities and client libraries, the project's repository accumulated 2,163 GitHub stars in its first five days. The release introduces a dedicated contrastive architecture packaged behind an API compatible with TypeSafe schema definitions, targeting real-time agent verification, tool selection, and state classification tasks.

Disaggregated Embeddings and the Shift from Autoregressive Action Selection

A generative model chooses an action by writing it out, token by token. CLM scores the options it is given instead, which is what makes it a decision model rather than a language model in the usual sense. For how decision models differ from LLMs, see what a decision model is, and how it differs from an LLM.

How CLM-8B scores an actionHow CLM-8B scores an action: State, then Candidate actions, then Qwen3-8B encoder, then Embedding cache, then State and action heads, then Answer.How CLM-8B scores an actionCandidates are embedded once and cached; a new state costs one embedding and a dot productoncecachescoreStateUp to 2,048 tokensby defaultCandidateactionsOptions, tools oranswers to chooseamongQwen3-8B encodervLLM in poolingmodeEmbedding cacheCandidateembeddings keptand reusedState and actionheads20M trainableparameters (a 75MB download)AnswerSoftmax over thecandidate scoresSources: GitHub.TerraNet Technologies · terranettechnologies.com

The numbers behind this chart

Step Model Does Hands off to
State - Up to 2,048 tokens by default Qwen3-8B encoder
Candidate actions - Options, tools or answers to choose among Qwen3-8B encoder
Qwen3-8B encoder - vLLM in pooling mode Embedding cache, State and action heads
Embedding cache - Candidate embeddings kept and reused State and action heads
State and action heads - 20M trainable parameters (a 75 MB download) Answer
Answer - Softmax over the candidate scores -

CLM-8B approaches the problem through a disaggregated contrastive formulation Source 1 · GitHub. Instead of running complete sequence-to-sequence passes across every state-action pair, CLM treats states and actions as distinct semantic entities whose representations can be projected, cached, and scored independently.

According to the developers, the pipeline relies on an underlying pooling encoder—specifically Qwen3-8B served via vLLM—paired with a compact 75 MB reference head that computes contrastive alignment scores Source 1 · GitHub. Because candidate action embeddings, questions, and discrete option criteria can be computed once and preserved in memory, a new state does not re-embed the candidates. In the README's example call, three typed questions about one state used 106 input tokens on a cold cache and 38 once the option texts were cached, and returned in 58.1 milliseconds. That is one demonstration call, not a benchmark.

Training Pipeline and Reported Performance Metrics

The CLM team trains in three stages, pulling each state toward the action that was actually taken and away from the rest Source 1 · GitHub:

CLM-8B training recipeThree stages, each counting a different kind of example
StageDataSize
Pre-trainingNemotron question-and-answer pairs60 million pairs
Mid-trainingSynthetic hard negatives30 million
Post-trainingAgentic trajectories1 million

Source: GitHub

The 30 million hard negatives - plausible but wrong answers - were generated by Gemini 2.5 Flash-Lite, so what CLM learns to reject is shaped by what that model thought a near miss looks like Source 1 · GitHub. The README does report its own ablation of the ordering: on about 100,000 held-out questions with ten hard negatives each, pre-training followed by a short hard-negative stage reached 69.2% top-1, against a peak of 62.4% for training on hard negatives from the start Source 1 · GitHub.

The authors claim that, zero-shot, CLM-8B performs on par with Jev across computer-use, gaming and tool-calling tasks while running up to 9× faster Source 1 · GitHub. The README says the speed-up is largest when there are many candidate actions to score (its WikiRacing task) or when actions recur across states (its T-Rex game), which is where caching pays off. It is a best case, not a typical figure.

The headline benchmark numbers need more care. The 87.6% on Terminal-Bench 2.1 and 81.6% on DeepSWE are not CLM solving those tasks. For each task, candidate solutions were sampled from Claude Fable 5 (Terminal-Bench) or Claude Opus 5 (DeepSWE), and a lightly fine-tuned CLM acted as the verifier that picked one Source 1 · GitHub. They were measured on 30 and 38 held-out tasks - the DeepSWE figure is 31 of 38 - and in that setting the README reports CLM running 4.1-5.7× faster than Jev, which it says fails as a verifier on these long-horizon tasks. What the result supports is that CLM is a promising best-of-N selector over strong generators. It says nothing yet about CLM acting as an agent on its own.

What the Announcement Leaves Out

Jacky Kwok's launch post on X, viewed about a million times since 24 September, compresses the README's results into claims that do not survive the compression Source 3 · X.

"New SOTA" is not a leaderboard result. The post says CLM-8B sets a new state of the art on DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%) Source 3 · X. On the public DeepSWE leaderboard, which measures agents solving the tasks themselves, the top score is 74.1% (our analysis of that leaderboard). CLM's figure measures something else: how often it picks a passing solution from candidates written by Claude Opus 5, on 38 held-out tasks, with a head fine-tuned for the benchmark Source 1 · GitHub. The post names neither the generator nor the number of tasks.

The samples are small, and one figure does not divide. 31 of 38 is 81.6%, but on 38 tasks the 95% interval for that rate runs from roughly 67% to 91%. The Terminal-Bench figure is stranger: 87.6% is not a possible pass rate on 30 tasks - 26 is 86.7% and 27 is 90% - so it is presumably an average over repeated runs, which neither the post nor the README says.

The verifier's baselines are missing from the text. A verifier is worth the gap between the generator on its own (pass@1) and how often any of its candidates passes (pass@N). The README's text reports neither, and the only verifier it compares against is Jev, which it says scores below pass@1 Source 1 · GitHub. There is no comparison with the generator judging its own candidates, a trained reward model, or a reranker.

"8B" is mostly Qwen3-8B. Each encoder is a frozen LLM backbone with a 20-million-parameter trainable projection head, and the served backbone is Qwen3-8B Source 1 · GitHub. The part CLM trained is 20 million parameters. The ablation that would show how much the contrastive training adds - a simple classifier on the same frozen Qwen3-8B embeddings - is not reported.

"Internet-scale" describes the backbone. The post says CLM-8B is pre-trained on internet-scale data Source 3 · X; the README's pre-training stage is about 60 million question-and-answer pairs Source 1 · GitHub. The internet-scale training is Qwen3-8B's own.

The scaling laws are about the training loss. The README's "scaling laws for verification" are fits of contrastive loss on held-out Nemotron question-and-answer data, with the fits published in a Notion post rather than a paper Source 1 · GitHub. Lower loss on question answering is not the same as better verification of agent runs, and nothing published yet ties the two together.

None of this makes the release weak. It ships reproduction commands, open weights and its own ablations, and its authors - Jacky Kwok with co-authors including Marco Pavone, Christopher Ré and Azalia Mirhoseini, per the README's citation - are serious researchers. It does make the headline numbers narrower than the post suggests.

Developer Workflow and Immediate Implementation Changes

In a multi-step agent loop, a large generative model is often called just to make a small decision: did this command succeed, which tool comes next, is this answer better than that one. Each call adds latency, and over a long run the delays add up. Those small decisions are what CLM is built to take over.

The three question typesEach is answered by scoring its options against the state
TypeAsksReturns
NoulIs this statement true?A probability
ChoiceWhich of these labelled options?The chosen option and a probability for each
ScoreWhere on an ordered scale?An expected level, e.g. 0 to 2 on a three-point scale

Source: GitHub

The README says a request written for TypeSafe replays unchanged as client.system_one(state, questions), so a team already using that format can point the same questions at CLM and compare Source 1 · GitHub. The obvious trial is a generative verification step - "did this command succeed?", "which tool next?" - replaced by a CLM call.

For free-form action candidate spaces—such as reranking best-of-N responses or choosing next moves—the framework provides an in-process rank method that scores candidate arrays against prompt states through an embedding endpoint Source 1 · GitHub. This allows engineers building routing layers or coding verification loops to offload classification decisions to local infrastructure.

Unverified Claims and Technical Boundaries

The repository has drawn attention quickly, which says nothing about whether its numbers hold. Separate what the architecture does from what the authors report it achieving:

What CLM claims, and what it rests onEvery figure is the authors' own; none has been replicated outside the project yet
ClaimReportedConditions
Parity with JevComputer use, gaming, tool callingZero-shot; the authors' own evaluation
Speed vs JevUp to 9× (zero-shot); 4.1-5.7× (verifier)Largest with many or reused candidate actions
Terminal-Bench 2.187.6%Picking among Fable 5 solutions; 30 held-out tasks; not a whole number of tasks
DeepSWE81.6% (31 of 38; 95% range about 67-91%)Picking among Opus 5 solutions; 38 held-out tasks; fine-tuned head
Model size"CLM-8B"Frozen Qwen3-8B backbone; 20M trained parameters
Scaling lawsLoss falls as a power lawContrastive loss on Q&A data, not verification accuracy
State length2,048 tokens by defaultUp to 8,192 with more GPU memory

Source: GitHub

The claims are at least checkable. The README includes the command that reproduces the DeepSWE result and ships the T-Rex benchmark in the repository, so anyone with a GPU can test the verifier claim Source 1 · GitHub. What is not reported is how scoring holds up on states near the 8,192-token ceiling, or on candidate sets from weaker generators than Opus 5 and Fable 5.

Availability, Setup, and Licensing Terms

The repository is licensed Apache 2.0 Source 1 · GitHub, and so is the CLM-v0.1-8B model on Hugging Face Source 2 · Hugging Face. The fine-tuned DeepSWE verifier heads are published separately under MIT. The API reference and fine-tuning guide are linked from the repository.

To run the stack locally, the client package is installed directly via Python's package index Source 1 · GitHub:

pip install contrastive-lm

Serving the architecture requires a two-step infrastructure deployment Source 1 · GitHub:

  1. Encoder Service: A pooling instance of Qwen3-8B served over vLLM at an assigned port Source 1 · GitHub:
   vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --runner pooling --max-model-len 2048 --port 8090
  1. Scoring Head: The dedicated CLM serving engine (clm-serve), which automatically retrieves the 75 MB reference head on initial launch and exposes the API over port 8700.

The default server cuts states off at 2,048 tokens. That is short for agent work - a terminal session or a coding trace routinely runs past it - and the DeepSWE heads the project published are an 8k variant, so the benchmark results were likely produced with a longer window than the default download gives you. Raise both limits, as the README describes, before judging CLM on your own traces Source 1 · GitHub.

Alternatively, teams can run the in-process Engine class by directly providing the endpoint URL of an existing vLLM embeddings service, bypassing HTTP server overhead entirely for co-located agent runtimes Source 1 · GitHub.

AI Tools