# CLM vs Jev for agent decisions: open weights, latency and trade-offs
Direct answer
CLM-8B is the better fit when you need Apache-2.0 weights, self-hosting, candidate caching, and task-specific fine-tuning. Jev is the simpler managed option and showed much stronger zero-shot accuracy in one independent package-curation benchmark. Neither replaces a generative LLM: both choose among predefined actions. Benchmark your own decisions, latency, confidence thresholds, and operating cost before choosing.
Why this comparison matters
Most “CLM vs Jev” searches are really asking four questions: can the model run under your control, will it make the right decision without training, how quickly will it answer, and what will the system cost to operate? A single leaderboard number cannot answer all four. CLM and Jev target the same narrow layer of an agent stack—choosing an action from a typed set—but package that capability very differently.
CLM is an open-weight contrastive model and server. Jev is a hosted System One model exposed through TypeSafe's API. Both avoid generating free-form prose for decisions that should have a finite answer. That is useful for routing, policy selection, tool choice, UI actions and other branches where an agent already knows its legal options. It is not a replacement for the language model that writes an explanation, creates code or discovers an action that was never supplied.
This guide separates official claims from independent evidence, explains why latency varies by workload, and provides a test protocol that can prevent a quick demo from becoming an unreliable production router.
The shared idea: score choices instead of generating text
A conventional LLM is often prompted to output a label such as `approve`, `review` or `reject`. The application then parses text and hopes the model followed the schema. A decision model starts with the allowable choices and scores them directly. The result can be represented as a typed value plus probabilities, making invalid strings much less likely and thresholds easier to implement.
That narrower interface brings an important limitation: the model cannot invent a missing choice. If an incident router receives only `database`, `network` and `application`, it cannot honestly return `identity provider` unless that option exists. Candidate-set design is therefore part of model quality, not merely plumbing.
Jev and CLM both use TypeSafe-compatible typed structures. In practical terms, this means a team can describe a state, ask a typed question, and receive one of the declared alternatives or a score. The similarity makes comparison possible, but it does not make the deployment models interchangeable.
What CLM-8B is
CLM-8B is published by the Contrastive-LM project with Apache-2.0 code and weights. Its reference architecture freezes a Qwen3-8B encoder and adds two roughly 20-million-parameter projection heads. One head embeds the current state; the other embeds candidate actions. Training uses a bidirectional contrastive objective so the correct state/action pairs move closer in embedding space and incorrect pairs move apart.
The project describes training on roughly 60 million Nemotron DQA question-answer examples, about 30 million synthetic hard negatives, and approximately one million agent trajectories. It also mixes pretraining replay with agent data. Those details matter because CLM is not simply a generic embedding model with a new label: it was deliberately trained to rank typed decisions.
The official server uses Qwen3-8B in vLLM pooling mode, normally on a GPU, and a separate CLM service for the projection heads. The documentation includes an RTX 4090 example and permits the small heads to run on CPU or CUDA. Do not translate that into “the whole system is a lightweight CPU model”: the reference encoder is still an 8B model, and memory, quantization, throughput and operational support must be measured on the hardware you intend to use.
CLM's most distinctive systems advantage is candidate caching. If candidate action embeddings are stable, the application can precompute them. Repeated states then become relatively cheap dot-product comparisons. The official repository reports RTX 4090 examples in which revisited-state latency drops from about 1.7 to 0.6 milliseconds or 2.0 to 0.7 milliseconds. A new state, however, remains around 28 milliseconds because it still needs encoder work. Caching is valuable, but only when the workload actually repeats reusable candidates or states.
What Jev is
Jev is TypeSafe's hosted System One decision model. Applications call the `/v1/systemone` API with an API key and typed inputs. There are no downloadable Jev weights in the official material, so self-hosting, offline execution and weight-level adaptation are not available in the same way as CLM.
TypeSafe advertises 70–500 millisecond responses and a price of $0.042 per million input tokens, with output free. Those are provider claims and commercial terms, not guarantees for every region or workload; verify the live terms before budgeting. The appeal is operational simplicity: the provider handles serving and the client sends typed requests. That can be compelling when a team wants to validate the decision-model pattern without owning an 8B inference service.
Jev is also intentionally narrow. TypeSafe and LangChain position it as a complement to generative models: use Jev for routing, classification and bounded choices, then use an LLM when the next step requires open-ended reasoning or generation. That division is more robust than forcing either product to do a job it was not designed for.
CLM vs Jev at a glance
| Question | CLM-8B | Jev |
|---|---|---|
| Access | Open code and weights under Apache-2.0 | Managed API; no open weights published |
| Hosting | Your infrastructure | TypeSafe-hosted |
| Core operation | Contrastive scoring of supplied candidates | Typed probabilistic decision over supplied choices |
| Adaptation | Fine-tuning supported and central to strongest claims | Provider-managed model; no weight-level self-training documented |
| Latency profile | Can be very low with cached candidates; new states still invoke an 8B encoder | Network plus managed inference; provider says 70–500 ms |
| Direct price | No per-call license fee, but compute and operations are yours | Provider lists $0.042 per million input tokens |
| Data boundary | Can remain inside your environment | Requests leave your environment for the hosted service |
| Best fit | Control, customization, high-volume repeated choices | Fast integration and stronger demonstrated zero-shot transfer in one independent suite |
The table is a deployment comparison, not a universal quality ranking. A locally hosted model with the wrong fine-tune can be worse than an API, while a well-adapted local model can be faster, cheaper at scale and easier to govern.
What the official CLM benchmark actually says
The CLM team says CLM-8B is on par with Jev across computer use, gaming and tool-calling, with latency up to nine times lower. Its verifier evaluation uses 38 held-out DeepSWE examples and 30 held-out Terminal-Bench 2.1 examples on an H100. The reported fine-tuned accuracy is 81.6% and 87.6%, with a 4.1–5.7× latency advantage over Jev.
These figures are useful evidence for the tested configuration, but three qualifiers are essential. First, they are published by the CLM team. Second, the strongest verifier numbers require task-specific fine-tuning; the model card explicitly warns that they are not zero-shot results. Third, small held-out sets can produce unstable percentages. A difference of one or two examples materially changes the total.
The responsible conclusion is that CLM can be highly competitive after adaptation on these agent tasks—not that an untouched checkpoint will match Jev on every classification problem.
The independent benchmark changes the decision
An independent `jev-bench` repository compares Jev and CLM with the same states, rubrics, 768-token cap, held-out tests and fresh-process single-stream latency. Its tasks concern package curation: quarantine decisions, curation, typosquat detection, reachability and license-family classification.
Zero-shot CLM scored 17%, 19%, 52%, 44% and 5% on those five tasks, while Jev scored 100%, 94%, 94%, 89% and 63%. The overall median latency row reports 116 ms for CLM and 136 ms for Jev, although individual cached task rows can show CLM in the low single-digit milliseconds. This result is dramatically less favorable to zero-shot CLM than the official comparison.
Fine-tuning changes the picture. The benchmark reports fine-tuned CLM scores of 100%, 88%, 100%, 76% and 58%. The repository excludes the typosquat result from its broad claims because of template leakage, a welcome caveat. Even with that exclusion, adaptation closes much of the quality gap.
This is not proof that Jev always wins zero-shot, nor that the official CLM results are wrong. The task distributions differ. It is evidence that transfer is fragile: a model trained for general agent decisions may still need examples that resemble your actual taxonomy, phrasing and error costs.
Latency: measure the whole request path
Latency comparisons often mix unlike measurements. A local benchmark may time only model work after warm-up, while a hosted request includes TLS, internet routing, queueing and serialization. Conversely, a self-hosted total must include the encoder, scheduler, process boundaries and cache misses—not only the final dot product.
Measure at least p50, p95 and p99 from the calling application. Separate cold starts from warm requests. Record candidate count, state length, batch size and cache-hit rate. Run a sequential test for interactive latency and a concurrent test for throughput. If the decision blocks a UI action, tail latency usually matters more than an attractive average.
CLM is especially sensitive to workload shape. Stable choice sets can reuse action embeddings, and repeated states may hit deeper caches. Highly dynamic states with constantly changing candidates pay more encoder cost. Jev's network overhead may be acceptable for a background workflow but noticeable inside a sub-100ms control loop.
Accuracy, probabilities and abstention
Both systems return probabilities relative to the supplied candidate set. A 90% result does not prove the action is objectively safe; it says the model preferred that choice among what it was shown. Add an explicit `unknown`, `escalate` or `none_of_the_above` choice when abstention is legitimate.
Calibrate thresholds on held-out production-like examples. For a low-cost content route, a 70% threshold may be acceptable. For deleting data or changing permissions, the correct policy may require a deterministic rule plus human approval, regardless of model confidence. Track confusion matrices and cost-weighted errors rather than only overall accuracy.
Typed output improves mechanical reliability, but it does not remove semantic ambiguity. Short option descriptions, overlapping labels and missing context can defeat either model. Write mutually exclusive choices, include the evidence needed for the decision, and test adversarial or incomplete states.
Cost and operational trade-offs
CLM has no per-request model fee under Apache-2.0, but “open” is not “free to operate.” Budget GPU acquisition or rental, idle capacity, deployment engineering, monitoring, updates, storage and incident response. At sustained high volume, those fixed costs may amortize well. At low or spiky volume, a managed service can be cheaper even when every API call has a price.
Jev converts much of that burden into a variable API charge. The listed input price is low, but production cost also includes retries, duplicated prompts, observability and any generative model used after the decision. Data residency, vendor availability and rate limits may dominate the financial calculation.
If your priority is experimenting without paid inference, consult FreeAI Tokens' free AI API guide and verified free tiers. For self-hosting experiments, the free GPU credits guide can help, but promotional credits should never be mistaken for a durable production budget.
When to choose CLM
Choose CLM when data must stay inside your boundary, Apache-2.0 weights are a requirement, your team can operate Qwen3-8B-class inference, or task-specific fine-tuning is part of the plan. It is particularly attractive when candidate sets repeat enough for caching and request volume can justify dedicated capacity.
CLM is also the clearer research platform. You can inspect code, freeze versions, reproduce evaluations and adapt the heads. That control is useful for regulated change management, although compliance still depends on your entire system and dataset.
Do not choose it solely because an official chart shows a multiple on latency. Confirm that your cache-hit rate, hardware and state lengths resemble the benchmark. Do not assume the base checkpoint is production-ready for an unfamiliar taxonomy.
When to choose Jev
Choose Jev when you want the smallest integration surface, cannot operate a local encoder, or need a strong zero-shot baseline quickly. The independent package-curation results make it a serious control candidate for a new domain: if Jev performs well before any training, it can establish the quality bar that a local model must beat.
The trade-offs are dependency on a hosted provider, paid usage, API-key management and less control over weights and serving. Confirm current pricing, retention, regional processing, rate limits and contractual terms directly with TypeSafe before production use.
A fair evaluation protocol
- Define 200–1,000 representative decisions with adjudicated labels and realistic class imbalance.
- Freeze the exact candidate descriptions and state template for both systems.
- Keep a truly held-out test set; remove template duplicates and near-duplicates.
- Test Jev and untouched CLM first to measure zero-shot transfer.
- Fine-tune CLM only on the training split, then rerun the frozen test.
- Measure p50/p95/p99 end-to-end latency from the application, including network and cache misses.
- Calibrate confidence and an abstention route; report per-class precision, recall and cost-weighted failures.
- Load-test realistic concurrency and calculate monthly total cost, including local operations.
- Run shadow traffic before allowing either system to take irreversible actions.
Publish the prompt schema, candidate list, hardware, software versions and sample counts with your result. Without those details, “five times faster” or “94% accurate” is not portable evidence.
Verdict
CLM-8B is the stronger option for ownership, adaptation and cache-friendly self-hosting. Jev is the stronger convenience option and, in the independent package-curation suite, the much stronger unadapted model. The apparent contradiction is the lesson: architecture and licensing determine what you can control, while domain evaluation determines what you can trust.
Use Jev as a rapid zero-shot baseline, CLM as a controllable candidate, or both in a shadow evaluation. Then select on error cost, tail latency, privacy and total operating cost—not the most flattering vendor number. More evidence-led comparisons are available in the FreeAI Tokens editorial hub.
Sources and methodology
- CLM official repository: architecture, training, server, caching and official evaluations.
- CLM-8B model card: license, limits and fine-tuning caveats.
- TypeSafe's Jev introduction: product design, provider claims and pricing.
- TypeSafe homepage and API documentation: current product positioning and API surface.
- LangChain's Jev harness guide: complementary decision-model pattern.
- Independent jev-bench repository: same-protocol zero-shot, fine-tuned and latency results.
Sources were reviewed on 29 September 2026. Official performance and price statements are labeled as provider claims; independent results are limited to their published tasks and protocol.