Jev is a new AI model introduced by TypeSafe AI. This is an AI lab co-founded by Diogo Almeida, one of the co-authors of OpenAI's InstructGPT paper, which applied RLHF to language models. (RLHF is an algorithm used to align LLMs to be helpful and harmless assistants to humans.)
Jev is not an LLM, but a System One model. This terminology is based on Daniel Kahneman’s systems of thinking, in which:
System 1 refers to fast, pattern-based thinking (e.g., recognizing faces)
System 2 refers to slow, effortful reasoning (e.g., solving complex math problems)
Like LLMs, System One models understand natural-language inputs.
But unlike LLMs, they do not write replies, produce code, or generate explanations of their reasoning. Instead of generating text, these models return typed decisions and probabilities. (Classical ML enthusiasts would love this!)
What do Jev’s inputs and outputs look like?
Jev accepts the following two inputs:
State: the content that the model has to evaluate. This is text or JSON.
Questions: the judgments the model should make given the state.
Given these inputs, Jev returns one typed answer per question with probabilities and a confidence score.
For example, an input could be:
State: “My app keeps crashing when I click the ‘Submit’ button.”
Question: “Which team should handle this user request: Technical or the Marketing team?”
For this input, the model would return a response such as:
“Technical (98% probability of this option with confidence of 95%).”
Join the paid tier today to get access to all articles in this newsletter and level up as an AI engineer.
What is State?
State is the content passed to Jev around which you ask questions.
It is declared using the state field in the input/ API request, and each such request evaluates one state against one or more questions.
A state could be either:
a string
"I have chest tightness for the last 2 days that worsens when going upstairs"an array
[ "Hi, I've had chest tightness that's worse when going upstairs for the last 2 days", "I have diabetes and high blood pressure" ]an object
{ "age": 67, "medical_conditions": ["Diabetes", "Hypertension"], "message": "I've chest tightness for 2 days that worsens when going upstairs" }
What are Primitives?
Primitives are the building blocks that form the inputs (questions) and outputs (answers) to Jev. These are of three types, and each answers something different:
Choice: Which of these options? (e.g., “Which team should the request be redirected to?”)
Score: Which level/ score? (e.g., “How relevant is this candidate's experience to the job posting?”)
Noul: Is this true or false? (e.g., “Is the customer requesting a refund?”)
These primitives are used to create questions that are passed to Jev.
Each question has the following fields:
ID: This is used to identify the answer in the response.type: This is the type of primitive, either Choice, Score, or Noul.instructions: This is the question you are asking about the state.If a question contains the Choice and Score primitives, they also take
criteria(the possible answers), which define the options for a Choice question or the levels for a Score question.Questions with the Noul primitive can also have an optional
criteriawhich clarifies what yes and no mean.
The answer depends on the type of primitive used in a question.
Choice question returns:
choice(the option that Jev chose)probabilities(probability distribution over the options given in thecriteriafield), andconfidence(this is a single value derived from the probability distribution which tells how certain Jev is about its choice)
A Score question returns:
score(a value between the given levels)legend(each level and its description)probabilitiesconfidence
A Noul question returns just
noul. This is the probability that the answer is yes/ true.
Aconfidencevalue isn’t returned here, and anoulvalue near 1 means a strong yes, near 0 means a strong no, and 0.5 means uncertain.
Why State and Primitives?
Primitives make Jev type-safe with its inputs and outputs. You get the exact type of response that you intend.
Also, every answer that Jev chooses comes from the options given to it. This ensures that it never hallucinates new options. Note that this does not mean that Jev does not make mistakes. It can still choose the wrong option from the given ones, but it will never invent a new option.
Confidence is what lets the model say it is not sure, rather than agreeing with what the user intends to get out of the model. This makes Jev foundational to building reliable systems.
Constructing an input to Jev
Now that we understand Primitives, it’s time to look at what an actual input to Jev looks like.
Consider the following input for a patient triage app where we send the state as an array of patient messages, and we ask three questions regarding it:
a
Noulquestion (labeled asurgency) on whether the patient message needs an urgent responsea
Choicequestion (team) of which team should handle it, anda
Scorequestion (severity) on how severe the symptoms are.
{
"model": "jev-latest",
"state": [
"Hi, I've had chest tightness that's worse when going upstairs for the last 2 days",
"I have diabetes and high blood pressure"
],
"questions": {
"urgency": {
"type": "noul",
"instructions": "Does this patient message need urgent response?"
},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"pharmacy": "For medication based issues",
"emergency_doctor": "For medical issues that need urgent or emergency attention",
"billing": "For pricing related issues"
}
},
"severity": {
"type": "score",
"instructions": "How severe are the symptoms described by the patient?",
"criteria": [
"Mild, no immediate concern",
"Moderate, can be seen in a few days",
"Serious, needs same-day attention",
"Critical, needs emergency care now"
]
}
}
}Jev will answer all three questions independently in parallel and return their respective probabilities and confidence.
Note that the answer to one question does not affect another in any way. Questions are completely independent, and one question’s answer is not hidden context for another.
An example response looks as follows.
{
"model": "jev-1.13.0",
"answers": {
"urgency": {
"type": "noul",
"noul": 0.94
},
"team": {
"type": "choice",
"choice": "emergency_doctor",
"probabilities": {
"pharmacy": 0.02,
"emergency_doctor": 0.95,
"billing": 0.03
},
"confidence": 0.93
},
"severity": {
"type": "score",
"score": 2.2,
"legend": {
"0": "Mild, no immediate concern",
"1": "Moderate, can be seen in a few days",
"2": "Serious, needs same-day attention",
"3": "Critical, needs emergency care now"
},
"probabilities": {
"0": 0.01,
"1": 0.09,
"2": 0.58,
"3": 0.32
},
"confidence": 0.71
}
},
"usage": {
"input_tokens": 180,
"output_tokens": 24
}
}Notice the input and output tokens? You're billed only for input tokens, while the output tokens are completely free. We will discuss this in the next section.
What makes Jev better than LLMs?
LLMs are post-trained using RLHF. This teaches them to say things that people prefer. Now this works well for chatbots in most cases, but it also rewards sycophancy (models praising users insincerely), confident-sounding hallucinations (inventing things that do not exist), and a preference for a particular style of responses (“You’re absolutely right!”).
Instead of RLHF, Jev is trained using Reinforcement Learning for Calibrated Decisions (RLCD). This algorithm trains Jev to output decisions and calibrated probabilities rather than generated text aligned with human preferences.
Jev is also cheap, at only $0.042 per million input tokens with completely free output tokens. This is because most computation occurs on the input tokens, and its outputs are tiny, since it does not generate text or reasoning tokens.
Compare this to Opus 5.5, a reasoning model that can produce a large volume of reasoning-related output tokens for decision-making and costs $4 per million input tokens and $20 per million output tokens.
(But take this price comparison with a grain of salt, as Jev struggles when making complex decisions and is not a replacement for reasoning LLMs in such use cases.)
Another great feature of Jev is its self-consistency, which makes its outputs reliable across repeated runs with the same input. In an experiment run by LangChain, Jev had the lowest observed mean per-case variance of 0.0000149. In the same experiment, GPT‑5.6 Luna’s variance was 433× higher, GPT-5.6 Terra’s was 913× higher, and Claude Sonnet 4.6’s was 92× higher.
What can you use Jev for?
Jev can be used for tasks that require fast, consistent decisions, ideally when those decisions do not require complex, multi-hop reasoning.
Some use cases of Jev are shown below. You can check out this page for a detailed description of these use cases.
TL;DR
Jev is a System One model from TypeSafe AI, trained to generate decisions based on fast, pattern-based thinking rather than the complex reasoning used by System Two models like LLMs.
Instead of generating text, Jev returns typed decisions with probabilities and a confidence score.
Inputs to Jev are a state and one or more questions, based on three building blocks, or Primitives: Choice, Score, and Noul.
‘Choice’ based questions return a choice of answer from the given options, ‘Score’ based questions return a score from the given levels, and ‘Noul’ based questions return the probability that something is true.
Answers always come from the options that are given to the model to pick from and are returned with probabilities (and confidence for ‘Choice’ and ‘Score’ based questions).
Each question is answered independently, and one question’s answer is not used as hidden context for another.
Jev is trained using RLCD (Reinforcement Learning for Calibrated Decisions) instead of RLHF. This means that Jev is not post-trained to produce outputs that are aligned with human preferences.
It is built to be used for fast judgment tasks like model routing, content moderation, guardrails, risk assessment, and more.
It is highly self-consistent across repeated turns of the same input, making it ideal for building reliable systems.
It costs $0.042 per million input tokens with free output tokens.
Join the paid tier today to get access to all posts in this newsletter, including:
and so many more!








