CLM-8B Decouples State-Action Scoring to Cut Decision Model Latency
Contrastive-LM has released CLM-8B, an open-source decision model that scores candidate actions against a state instead of generating them. Its authors claim parity with Jev at up to 9× lower latency; here is what the repository shows and what it does not yet.
~9 min spoken. Keeps playing while you work in another tab.
On September 23, 2026, the Contrastive-LM open-source project published the initial release of CLM-8B, a specialized "System One" decision model designed for rapid state-action scoring Source 1 · GitHub. Accompanied by serving utilities and client libraries, the project's repository accumulated 2,163 GitHub stars in its first five days. The release introduces a dedicated contrastive architecture packaged behind an API compatible with TypeSafe schema definitions, targeting real-time agent verification, tool selection, and state classification tasks.
Disaggregated Embeddings and the Shift from Autoregressive Action Selection
A generative model chooses an action by writing it out, token by token. CLM scores the options it is given instead, which is what makes it a decision model rather than a language model in the usual sense. For how decision models differ from LLMs, see what a decision model is, and how it differs from an LLM.
The numbers behind this chart
| Step | Model | Does | Hands off to |
|---|---|---|---|
| State | - | Up to 2,048 tokens by default | Qwen3-8B encoder |
| Candidate actions | - | Options, tools or answers to choose among | Qwen3-8B encoder |
| Qwen3-8B encoder | - | vLLM in pooling mode | Embedding cache, State and action heads |
| Embedding cache | - | Candidate embeddings kept and reused | State and action heads |
| State and action heads | - | 20M trainable parameters (a 75 MB download) | Answer |
| Answer | - | Softmax over the candidate scores | - |
CLM-8B approaches the problem through a disaggregated contrastive formulation Source 1 · GitHub. Instead of running complete sequence-to-sequence passes across every state-action pair, CLM treats states and actions as distinct semantic entities whose representations can be projected, cached, and scored independently.
According to the developers, the pipeline relies on an underlying pooling encoder—specifically Qwen3-8B served via vLLM—paired with a compact 75 MB reference head that computes contrastive alignment scores Source 1 · GitHub. Because candidate action embeddings, questions, and discrete option criteria can be computed once and preserved in memory, a new state does not re-embed the candidates. In the README's example call, three typed questions about one state used 106 input tokens on a cold cache and 38 once the option texts were cached, and returned in 58.1 milliseconds. That is one demonstration call, not a benchmark.
Training Pipeline and Reported Performance Metrics
The CLM team trains in three stages, pulling each state toward the action that was actually taken and away from the rest Source 1 · GitHub:
| Stage | Data | Size |
|---|---|---|
| Pre-training | Nemotron question-and-answer pairs | 60 million pairs |
| Mid-training | Synthetic hard negatives | 30 million |
| Post-training | Agentic trajectories | 1 million |
Source: GitHub
The 30 million hard negatives - plausible but wrong answers - were generated by Gemini 2.5 Flash-Lite, so what CLM learns to reject is shaped by what that model thought a near miss looks like Source 1 · GitHub. The README does report its own ablation of the ordering: on about 100,000 held-out questions with ten hard negatives each, pre-training followed by a short hard-negative stage reached 69.2% top-1, against a peak of 62.4% for training on hard negatives from the start Source 1 · GitHub.
The authors claim that, zero-shot, CLM-8B performs on par with Jev across computer-use, gaming and tool-calling tasks while running up to 9× faster Source 1 · GitHub. The README says the speed-up is largest when there are many candidate actions to score (its WikiRacing task) or when actions recur across states (its T-Rex game), which is where caching pays off. It is a best case, not a typical figure.
The headline benchmark numbers need more care. The 87.6% on Terminal-Bench 2.1 and 81.6% on DeepSWE are not CLM solving those tasks. For each task, candidate solutions were sampled from Claude Fable 5 (Terminal-Bench) or Claude Opus 5 (DeepSWE), and a lightly fine-tuned CLM acted as the verifier that picked one Source 1 · GitHub. They were measured on 30 and 38 held-out tasks - the DeepSWE figure is 31 of 38 - and in that setting the README reports CLM running 4.1-5.7× faster than Jev, which it says fails as a verifier on these long-horizon tasks. What the result supports is that CLM is a promising best-of-N selector over strong generators. It says nothing yet about CLM acting as an agent on its own.
What the Announcement Leaves Out
Jacky Kwok's launch post on X, viewed about a million times since 24 September, compresses the README's results into claims that do not survive the compression Source 3 · X.
"New SOTA" is not a leaderboard result. The post says CLM-8B sets a new state of the art on DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%) Source 3 · X. On the public DeepSWE leaderboard, which measures agents solving the tasks themselves, the top score is 74.1% (our analysis of that leaderboard). CLM's figure measures something else: how often it picks a passing solution from candidates written by Claude Opus 5, on 38 held-out tasks, with a head fine-tuned for the benchmark Source 1 · GitHub. The post names neither the generator nor the number of tasks.
The samples are small, and one figure does not divide. 31 of 38 is 81.6%, but on 38 tasks the 95% interval for that rate runs from roughly 67% to 91%. The Terminal-Bench figure is stranger: 87.6% is not a possible pass rate on 30 tasks - 26 is 86.7% and 27 is 90% - so it is presumably an average over repeated runs, which neither the post nor the README says.
The verifier's baselines are missing from the text. A verifier is worth the gap between the generator on its own (pass@1) and how often any of its candidates passes (pass@N). The README's text reports neither, and the only verifier it compares against is Jev, which it says scores below pass@1 Source 1 · GitHub. There is no comparison with the generator judging its own candidates, a trained reward model, or a reranker.
"8B" is mostly Qwen3-8B. Each encoder is a frozen LLM backbone with a 20-million-parameter trainable projection head, and the served backbone is Qwen3-8B Source 1 · GitHub. The part CLM trained is 20 million parameters. The ablation that would show how much the contrastive training adds - a simple classifier on the same frozen Qwen3-8B embeddings - is not reported.
"Internet-scale" describes the backbone. The post says CLM-8B is pre-trained on internet-scale data Source 3 · X; the README's pre-training stage is about 60 million question-and-answer pairs Source 1 · GitHub. The internet-scale training is Qwen3-8B's own.
The scaling laws are about the training loss. The README's "scaling laws for verification" are fits of contrastive loss on held-out Nemotron question-and-answer data, with the fits published in a Notion post rather than a paper Source 1 · GitHub. Lower loss on question answering is not the same as better verification of agent runs, and nothing published yet ties the two together.
None of this makes the release weak. It ships reproduction commands, open weights and its own ablations, and its authors - Jacky Kwok with co-authors including Marco Pavone, Christopher Ré and Azalia Mirhoseini, per the README's citation - are serious researchers. It does make the headline numbers narrower than the post suggests.
Developer Workflow and Immediate Implementation Changes
In a multi-step agent loop, a large generative model is often called just to make a small decision: did this command succeed, which tool comes next, is this answer better than that one. Each call adds latency, and over a long run the delays add up. Those small decisions are what CLM is built to take over.
| Type | Asks | Returns |
|---|---|---|
| Noul | Is this statement true? | A probability |
| Choice | Which of these labelled options? | The chosen option and a probability for each |
| Score | Where on an ordered scale? | An expected level, e.g. 0 to 2 on a three-point scale |
Source: GitHub
The README says a request written for TypeSafe replays unchanged as client.system_one(state, questions), so a team already using that format can point the same questions at CLM and compare Source 1 · GitHub. The obvious trial is a generative verification step - "did this command succeed?", "which tool next?" - replaced by a CLM call.
For free-form action candidate spaces—such as reranking best-of-N responses or choosing next moves—the framework provides an in-process rank method that scores candidate arrays against prompt states through an embedding endpoint Source 1 · GitHub. This allows engineers building routing layers or coding verification loops to offload classification decisions to local infrastructure.
Unverified Claims and Technical Boundaries
The repository has drawn attention quickly, which says nothing about whether its numbers hold. Separate what the architecture does from what the authors report it achieving:
| Claim | Reported | Conditions |
|---|---|---|
| Parity with Jev | Computer use, gaming, tool calling | Zero-shot; the authors' own evaluation |
| Speed vs Jev | Up to 9× (zero-shot); 4.1-5.7× (verifier) | Largest with many or reused candidate actions |
| Terminal-Bench 2.1 | 87.6% | Picking among Fable 5 solutions; 30 held-out tasks; not a whole number of tasks |
| DeepSWE | 81.6% (31 of 38; 95% range about 67-91%) | Picking among Opus 5 solutions; 38 held-out tasks; fine-tuned head |
| Model size | "CLM-8B" | Frozen Qwen3-8B backbone; 20M trained parameters |
| Scaling laws | Loss falls as a power law | Contrastive loss on Q&A data, not verification accuracy |
| State length | 2,048 tokens by default | Up to 8,192 with more GPU memory |
Source: GitHub
The claims are at least checkable. The README includes the command that reproduces the DeepSWE result and ships the T-Rex benchmark in the repository, so anyone with a GPU can test the verifier claim Source 1 · GitHub. What is not reported is how scoring holds up on states near the 8,192-token ceiling, or on candidate sets from weaker generators than Opus 5 and Fable 5.
Availability, Setup, and Licensing Terms
The repository is licensed Apache 2.0 Source 1 · GitHub, and so is the CLM-v0.1-8B model on Hugging Face Source 2 · Hugging Face. The fine-tuned DeepSWE verifier heads are published separately under MIT. The API reference and fine-tuning guide are linked from the repository.
To run the stack locally, the client package is installed directly via Python's package index Source 1 · GitHub:
pip install contrastive-lm
Serving the architecture requires a two-step infrastructure deployment Source 1 · GitHub:
- Encoder Service: A pooling instance of Qwen3-8B served over vLLM at an assigned port Source 1 · GitHub:
vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --runner pooling --max-model-len 2048 --port 8090
- Scoring Head: The dedicated CLM serving engine (
clm-serve), which automatically retrieves the 75 MB reference head on initial launch and exposes the API over port 8700.
The default server cuts states off at 2,048 tokens. That is short for agent work - a terminal session or a coding trace routinely runs past it - and the DeepSWE heads the project published are an 8k variant, so the benchmark results were likely produced with a longer window than the default download gives you. Raise both limits, as the README describes, before judging CLM on your own traces Source 1 · GitHub.
Alternatively, teams can run the in-process Engine class by directly providing the endpoint URL of an existing vLLM embeddings service, bypassing HTTP server overhead entirely for co-located agent runtimes Source 1 · GitHub.