Jev is an AI model that makes decisions instead of writing text. You give it a situation and a question with a fixed set of answers, and it returns a choice with probabilities attached, fast enough to sit inside software that decides thousands of times a minute.
What a decision model is
A large language model (LLM), like the one behind ChatGPT, produces its answer one token at a time. A token is a small chunk of text, often part of a word. Even when all you need is 'yes' or 'no', the model wraps it in a sentence, and every token costs time and money.
Jev, from the San Francisco lab TypeSafe AI, works differently. TypeSafe calls it a 'System One' model, after the psychologist Daniel Kahneman's split between fast, instinctive thinking (System 1) and slow, deliberate thinking (System 2). Jev does not generate text at all. It reads the input and weighs every allowed answer in a single parallel pass, so the verdict arrives in one step instead of word by word.
That design comes with a hard rule: the possible answers must be known before you ask. Jev can pick 'billing', 'technical' or 'sales' for a support ticket, but it can't write the reply to the customer.
Asking Jev a question
A call has two parts:
- State: the thing to judge, as text. A log line, an email, a support ticket, or a JSON blob describing a game or a web page.
- Questions: one or more typed questions about that state, each with its answers defined in advance.
There are three question types:
- Choice picks one option from a list you supply, such as
ignoreorflag. - Score places the state on an ordered scale, such as calm, frustrated or angry.
- Noul is a yes-or-no question, answered as a single probability between 0 and 1.
You can ask several questions about the same state in one request. They run in parallel and in isolation, so adding a question barely changes the response time. TypeSafe's docs advise keeping each question to one specific, well-scoped thing, rather than packing a whole judgement into one vague prompt.
A worked example: log triage
Imagine a service that writes thousands of log lines a minute. Asking a large LLM 'is this worth a human's attention?' about every line would be slow and costly, and you would still have to dig the answer out of a sentence. This is the kind of job a decision model is built for.
The code below follows the shape of TypeSafe's Python SDK as its quickstart shows it. Treat it as a sketch, and check the current docs before copying it.
from typesafe_sdk import Choice, TypeSafeClient
client = TypeSafeClient()
def triage(line):
response = client.system_one(
state=line,
questions={
"action": Choice(
instructions="What to do with this line",
criteria={
"ignore": "Routine, expected lines",
"flag": "Errors or anything odd",
},
),
},
)
answer = response.answers["action"]
if answer.confidence < 0.8:
return "review"
return answer.choiceEvery line gets one of two labels, and your code gets a plain string back, not prose to interpret. The obvious lines are settled in a fraction of a second. Only the ones the model is unsure about go down the slower path, here a review queue, which might end with a larger model or a person.
Probabilities and confidence
A Choice or a Score comes back with three things: the chosen answer, a probability for every option, and a confidence value between 0 and 1.
The probabilities show how the model split its belief. 'flag 0.97, ignore 0.03' is a clear call. 'flag 0.52, ignore 0.48' is close to a coin toss, yet both return 'flag' as the answer. Confidence sums up that spread in one number: all the probability on one option gives 1.0, and the more evenly it is spread, the lower the confidence. A Noul has no separate confidence, because its one probability already says how sure it is: near 1 is a strong yes, near 0 a strong no, and near 0.5 means unsure.
This is what makes the output safe to automate. A common pattern has three bands:
- High confidence: act automatically.
- In the middle: act, but ask for confirmation or log it for review.
- Low confidence: hand it to a person, a larger model or another system.
Where you draw the lines depends on the cost of being wrong. Muting a noisy log line can take a lower bar than approving a payment, even inside the same system.
TypeSafe trains Jev with a method it calls Reinforcement Learning for Calibrated Decisions, aimed at making confidence track accuracy: a model that says 0.9 should be right about nine times in ten. LLMs are often overconfident, so calibration is the property to test on your own data before you lean on it.
Using Jev inside an AI agent
An AI agent is a program that uses an LLM to plan and carry out a task, calling tools such as a shell, a browser or an API along the way. The MCP vs API video covers how agents reach those tools. Between the big steps sit many small checks: 'is this command safe to run?', 'is this the page I expected?', 'has this step finished?'.
If every one of those checks goes to the large model, the agent spends most of its time and budget on questions that need no real reasoning. A better split is to let Jev answer the small questions first and pass only the harder problems to the larger model, which then spends its effort on planning and writing, the work it is good at. LangChain's write-up on building an agent harness with Jev draws the same line: an LLM for open-ended reasoning and generation, Jev for fast, structured decisions along the way. One of its examples checks each tool call for risk and blocks it before the tool runs.
Here is that gate at work, with one confident answer and one unsure one:
An agent checks with Jev before acting
Step 1 of 7: The agent wants to clear a temp folder, so it asks Jev: safe or risky?
Why it is faster and cheaper
The speed comes from skipping generation. An LLM needs one step per token, so a longer answer takes more steps. Jev answers every question in one pass, and because it writes no text, you pay for the input only.
TypeSafe's own benchmarks claim Jev is up to about 190 times faster and 440 times cheaper than leading LLMs on decision tasks. Those are the vendor's figures, on tests it chose, so measure on your own workload before building a budget around them.
Where it fits
The speed suits systems that decide constantly and in real time:
- Games, where a character needs a choice every frame or two.
- Robots and simulations, reacting to a steady stream of readings.
- Web automation, such as checking a page is the one you expected before clicking.
- Coding tools and agents, for quick checks before a larger model is involved.
- High-volume triage, such as routing tickets, scoring content or flagging logs.
The common thread is repetition with a known set of answers. If the same kind of question comes up thousands of times, and each answer is one of a handful of options, a decision model is worth trying.
Limits and trade-offs
- No text, no code. Jev can't write a reply, summarise a document or explain a bug. Those stay with an LLM.
- Bounded answers only. A question like 'what should we do about this?' has no fixed set of answers, so it isn't a fit.
- A number, not a reason. Jev tells you what it picked and how sure it is, but not why. If you have to justify a decision to a customer or a regulator, you need more than its answer.
- Text in only. For now it reads text, not images, audio or video.
- 'No hallucinations' has a narrow meaning. The answer always fits your schema, so you never get an option that doesn't exist. It can still pick the wrong one.
Common mistakes
- Treating it as a faster ChatGPT. It is a different kind of model, not a quicker chatbot. Use it for choices and an LLM for words.
- Ignoring the confidence. The answer alone hides how close the call was. A low-confidence 'safe' deserves less trust than a high-confidence one.
- One vague question instead of several sharp ones. Split 'is this ticket a problem?' into urgency, team and sentiment. Extra parallel questions cost little.
- Skipping your own tests. Check accuracy and calibration on real examples from your system before handing it decisions that matter.
Key takeaways
- Jev is a decision model: it returns a typed answer with probabilities and confidence, not text.
- It answers in one parallel pass instead of token by token, which is where the speed and low cost come from.
- It fits small, repeated decisions with a fixed set of answers, not open-ended work.
- Use confidence to decide when to act and when to escalate.
- In an agent, let Jev answer the small checks and send only the bigger problems to the larger model.