AI engineering
Routing AI decisions on a confidence score you can trust
A calibrated model that reports 0.8 should be right 8 times in 10. How that turns an AI step into a threshold your code can actually branch on.
Listen to the audio explanation
5 min
Its own script, written for listening.
Most AI developers spend an incredible amount of time fixing errors that shouldn't even exist.
Editor's note
Why this matters now
TypeSafe trains Jev so that its confidence score means what it says. When Jev reports 0.8, it should be right about 8 times in 10 across enough decisions to count. TypeSafe describes the training as reinforcement learning for calibrated decisions.
That matters because of what usually happens without it. A general model returns an answer, the code takes it, and nobody finds out how close a call it was. Ask one for a probability and it tends to be overconfident. A calibrated score turns that uncertainty into a threshold you can actually set: automate above it, and send everything below it to a person.
The source
What it says
Distilled from the original. The notes above and below are the editor's own.
Moving from Prose to Probabilities
In current AI architectures, making a decision—such as routing a support ticket or choosing a tool for an agent—usually involves a heavy-handed approach. Developers write a prompt, ask a generative Large Language Model (LLM) for a JSON response, parse that text, validate the schema, and then run a retry loop if the model hallucinated a format error. This process is slow, expensive, and structurally fragile.
Jev, a "System One" model from TypeSafe AI, is designed to solve this by inverting the traditional generative flow. Instead of asking a model to produce text that happens to contain a decision, the application declares the valid answers before the model even runs.
By moving from open-ended prose to constrained, typed decisions, Jev removes the need for parsing and retrying. It treats decision-making as a specialized software function rather than a creative writing task. This shift allows the model to act as a high-speed decision signal within an AI gateway, providing the structured data—like categories or scores—that application code can branch on directly. The goal is to provide "frontier-intelligence" that fits into the millisecond-sensitive request paths used by modern software.
Question Primitives and RLCD
Jev does not generate text token-by-token. Instead, it consumes the current application state and evaluates it against three specific "question primitives." These primitives ensure that every output is strictly constrained to a predefined schema.
The three primitives are:
- Choice: Selects exactly one option from a list of up to 255 predefined choices (e.g., choosing between
billing,support, orfraud). - Score: Places the current state onto an ordered rubric defined by the user (e.g., rating a customer's risk as
low,medium, orhigh). - Noul: Acts as a Boolean check, estimating the probability (between 0 and 1) that a specific statement about the state is true.
Because Jev evaluates these questions in parallel and in isolation, a single call containing multiple questions costs roughly the same as a call with just one.
Want the technical picture?
To ensure these probabilities are actually useful for automation, TypeSafe uses a method called Reinforcement Learning for Calibrated Decisions (RLCD). In a "calibrated" model, the probability score is a reliable reflection of reality. For example, if the model assigns a 0.8 probability to a decision, that decision should be correct approximately 80% of the time across a large number of instances. This calibration is what allows engineers to set mathematical thresholds for automation.
It is important to distinguish between "format hallucinations" and "factual errors." TypeSafe claims a zero-hallucination rate for Jev, but this refers strictly to schema adherence. Because the model is physically incapable of answering outside the declared options, it will never return a malformed response or a random string. However, it can still pick the wrong valid option. This is why the accompanying probability scores are the most critical part of the output; they flag when the model is guessing, even if the format is perfect.
The Efficiency Gains: Speed, Cost, and Benchmarks
The primary motivation for adopting a decision model like Jev is the significant gap in efficiency compared to traditional generative LLMs. Because Jev skips the expensive process of token-by-token generation, it can operate within the tight latency budgets required for real-time request routing.
TypeSafe reports end-to-end response times between 70 and 500 milliseconds. In their own workflow evaluations, they claim Jev is up to 193.6x faster and 444.6x cheaper than generative LLMs.
The cost structure is also simplified. While generative models charge for the total volume of text produced, Jev's pricing focuses on the input state:
| Metric | Jev (TypeSafe AI) | Generative LLMs (Typical) |
|---|---|---|
| Input Cost | $0.042 per 1M tokens | Variable (often significantly higher) |
| Output Cost | Free | Charged per token |
| Latency | 70–500 ms | Seconds (scales with output length) |
Note on Benchmarks: These performance figures are vendor-reported. TypeSafe has disclosed that these benchmarks were derived from workflows built by their own team, using reference answers from two frontier LLMs. While the headline gains are impressive, they represent the high end of real-world potential. Until independent third-party benchmarks are published, builders should treat these figures as a best-case target rather than a guaranteed baseline.
Deployment Constraints and Scaling Uncertainties
While the efficiency gains are compelling, Jev is not a "plug-and-play" replacement for an entire routing system. It is a component—a decision signal—rather than a standalone router or gateway.
To use Jev effectively, an architect must still deploy a system (often an AI gateway) that manages the request path. This system is responsible for:
- Deriving the application state to be evaluated.
- Calling the Jev model.
- Applying the business policy to the returned decision and probability.
- Forwarding the actual request to the final destination (a model, a tool, or a human).
There are also several technical unknowns regarding how Jev scales. It is currently unclear how the model's performance or latency might change as the density of question primitives increases within a single call. Furthermore, the specific operational requirements for integrating Jev into a custom-built, high-scale gateway architecture remain undefined. Builders should prepare for a deployment model where Jev provides the "intelligence" and a separate gateway layer provides the "infrastructure."
Using Jev as a Policy Dial
For engineering teams, Jev should be viewed as a "policy dial" that allows you to balance automation against human intervention. Because every answer comes with a calibrated probability, you can move away from binary "yes/no" logic and toward a spectrum of confidence.
Implementing the "Policy Dial"
Architects should implement Jev within a semantic gateway to manage the automation-to-human escalation ratio. You can set different thresholds depending on the "blast radius" of the action:
- High-Confidence Automation: For routine, low-risk tasks (like routing a common support query), set a permissive threshold. If Jev returns a confidence score above 0.9, the system automates the action immediately.
- Low-Confidence Escalation: For high-stakes actions (like approving a refund or triggering a financial transaction), set a conservative threshold. If the confidence is below 0.99, the system automatically routes the request to a human reviewer.
Architectural Integration
Do not attempt to use Jev as a standalone router. Instead, integrate it as the decision layer of a broader gateway. This allows you to separate the decision-making (is this request urgent?) from the execution (send this to GPT-4o). By using Jev to "decide what intelligence is needed," you can route routine traffic to cheaper, faster models and reserve expensive frontier models only for the complex cases that Jev flags as uncertain. This approach optimizes both the budget and the user experience by minimizing latency for the majority of requests.
Further Reading
- Vercel: Jev on AI Gateway
Editor's note
What to do with this
You can check whether your current model deserves that kind of trust without adopting anything new. Take a few hundred past decisions where you know the outcome, bucket them by the score the model gave, and see whether the 90% bucket really was right 9 times in 10.
If it fell short, that is worth knowing before you set any threshold at all. A threshold on an uncalibrated score looks like a control, and nobody can say what it controls.
The original
What is Jev: TypeSafe AI's decision model and where it fits
Paper Compute · 17 September 2026
Read next
AI engineering
What is Jev? The fast model that answers in JSON
Jev is a System 1 model built for classification. What it does, where it is far faster, and the 4 jobs it is wrong for.
7 min read7 min listen
AI engineering
Jev vs general-purpose LLMs: what 75x cheaper actually buys
Jev matches a general model that costs 75 times as much, and trails the strongest by 6 points. How to decide whether that trade fits your workflow.
6 min read5 min listen