Decision Models: Jev and the End of Prompting for Classification
Jev from TypeSafe AI understands language at a frontier level but never generates a token. It is a non-autoregressive decision model on DigitalOcean Serverless Inference that returns typed answers with calibrated probabilities. Here is how it differs from an LLM, with working API calls for all three question types.

The short version
- No autoregressive loop. LLMs write one token at a time, each waiting on the last. Jev reads the input (the "state") and outputs probabilities straight from its hidden layers in a single forward pass. That makes it much faster (reportedly 70 to 190 times) and cheaper.
- Parallel questions. Jev reads the shared context once, then evaluates every question (intent, frustration, PII) in parallel branches that don't wait on or influence each other. An LLM works through them sequentially or needs separate API calls.
- System 1, not System 2. Borrowing Kahneman's Thinking, Fast and Slow: LLMs are the slow, analytical System 2. Jev is the fast, automatic System 1, the reflexes that route a ticket, score risk or pick a tool while larger LLMs handle the heavy reasoning.
- Calibrated by RLCD. Where RLHF tunes chat models to sound helpful, Reinforcement Learning for Calibrated Decisions tunes Jev to be accurate about its confidence. When it says 90%, it is right about 9 times out of 10.
In short: a model that can't chat, but makes structured, parallel decisions software can act on immediately.
Most of the "AI" inside a production app is not writing essays. It is answering small questions: Which team gets this ticket? Is this message abusive? Does this text contain a phone number? We usually hand those questions to an LLM, ask it nicely to reply in JSON, and write a parser to catch the days it doesn't.
Jev, from TypeSafe AI, takes a different approach. It is a large model that reads language at a frontier level, but it is not an LLM, because it never generates text.
A "System One" model
TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, calls Jev a System One model. If you come from classical machine learning, think of a general-purpose classifier or decision model that understands free-form text. Per the reported training method, it is trained with Reinforcement Learning for Calibrated Decisions (RLCD). Its only job is to read unstructured text and return a structured decision with calibrated probabilities.
What makes it different from an LLM
| LLM (GPT, Claude, Gemini) | Decision model (Jev) | |
|---|---|---|
| Output | Free text, one token at a time | A structured answer plus probabilities |
| Generation | Autoregressive | Non-autoregressive, single forward pass |
| Format guarantees | Best effort, even with JSON mode prompting | Constrained to the schema you send |
| Failure modes | Hallucinated keys, chatty preambles, broken JSON | Can only pick from the options you defined |
| API | Chat Completions (messages) | /v1/systemone (state + questions) |
| Best at | Writing, summarizing, open-ended reasoning | Routing, moderation, classification, scoring |
Three properties matter in practice:
- Non-autoregressive. An LLM's answer depends on every token before it, which is why you can watch it type. Jev evaluates your input and your questions together and returns the final decision and probability scores at once. There is no sequential generation.
- Schema-constrained by design. Ask an LLM for JSON and you get text that looks like JSON. Jev's output space is the schema. Give it a choice question with three options and it can only return one of those three, with a probability for each.
- Calibrated. Every answer carries a calibrated probability, so you can set explicit thresholds for when your application acts automatically and when it escalates to a person, instead of branching on whatever the model happened to say.
If an LLM is the creative brain of an application, Jev is its reflexes: a fast, cheap if/else that happens to understand English.
The API: state plus questions
Jev is served as a serverless gateway on DigitalOcean Inference. It uses its own /v1/systemone endpoint rather than Chat Completions, so it is not OpenAI-compatible: instead of a messages array you send a state and a questions object.
model: the model ID, heretypesafe-jev-1.13.0.state: the context every question is evaluated against. It can be a string, a JSON object, or an array of text values.questions: a map of named questions. Each sets atype,instructionstelling Jev what to decide, and whatever fields that type requires.
Jev supports three question types: choice, score, and noul. Create a model access key from the Serverless Inference tab of the DigitalOcean Control Panel and export it as MODEL_ACCESS_KEY; the examples below read it from there.
choice: pick one option
Route a support ticket. Each option gets a key and a description.
curl -sS -X POST https://inference.do-ai.run/v1/systemone \
-H "Authorization: Bearer $MODEL_ACCESS_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe-jev-1.13.0",
"state": "I need help resetting my password, the link in the email expired.",
"questions": {
"department_routing": {
"type": "choice",
"instructions": "Which department should handle this ticket?",
"criteria": {
"billing": "Issues with payments, invoices, or subscriptions",
"technical_support": "Issues with login, bugs, or system errors",
"sales": "Inquiries about buying the product",
"general": "Anything else"
}
}
}
}'
{
"model": "typesafe-jev-1.13.0",
"answers": {
"department_routing": {
"type": "choice",
"choice": "technical_support",
"probabilities": { "billing": 0, "general": 0, "sales": 0, "technical_support": 1 },
"confidence": 1
}
},
"usage": { "input_tokens": 368, "output_tokens": 50 }
}
The choice value is always one of your criteria keys, so you can use it directly as a routing table index with no parsing or validation step. The confidence field lets you send low-confidence tickets to a human instead.
score: rate on a scale
Measure how frustrated a customer is. You supply the levels and a description for each. In the reference docs, criteria is an ordered array from lowest to highest, where each level is a descriptive string or an object with a label. The example below also passes a levels array, which the reference docs don't mention, so treat criteria as the field you rely on.
curl -sS -X POST https://inference.do-ai.run/v1/systemone \
-H "Authorization: Bearer $MODEL_ACCESS_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe-jev-1.13.0",
"state": "I have been waiting for 3 weeks for my refund! This is completely unacceptable and I will be reporting you to the BBB if this is not fixed today!!!",
"questions": {
"frustration_level": {
"type": "score",
"instructions": "Rate the customer frustration level on a scale of 1 to 5.",
"levels": ["1", "2", "3", "4", "5"],
"criteria": [
"Completely calm and polite",
"Slightly annoyed",
"Annoyed but reasonable",
"Very angry",
"Extremely angry, threatening, or using aggressive language"
]
}
}
}'
{
"model": "typesafe-jev-1.13.0",
"answers": {
"frustration_level": {
"type": "score",
"score": 3.99,
"legend": {
"0": "Completely calm and polite",
"1": "Slightly annoyed",
"2": "Annoyed but reasonable",
"3": "Very angry",
"4": "Extremely angry, threatening, or using aggressive language"
},
"probabilities": { "0": 0, "1": 0, "2": 0, "3": 0.01, "4": 0.99 },
"confidence": 0.99
}
},
"usage": { "input_tokens": 372, "output_tokens": 20 }
}
Note that the response legend is zero-indexed, so your five levels come back as 0 to 4. The score is the probability-weighted average of the level positions, which is why 99% of the mass on level 4 gives 3.99. Because the score is continuous, you can set thresholds (escalate to a human above 3.5, for example) instead of matching on a label.
A smaller scale works the same way. Asking for overall satisfaction on ["low", "medium", "high"] about the sentence "The customer said the delivery was late but the product quality was excellent." returned:
{
"type": "score",
"score": 1.32,
"legend": { "0": "low", "1": "medium", "2": "high" },
"probabilities": { "0": 0, "1": 0.68, "2": 0.32 },
"confidence": 0.52
}
The confidence of 0.52 is much lower than the 0.99 above. The model is genuinely split between medium and high, and it tells you so. That is the signal to route this one to a person.
noul: a calibrated yes/no
noul is TypeSafe's question type for yes/no decisions. It does not return a literal yes or no. It returns a calibrated probability between 0 and 1 that the statement in your instructions is true, and you threshold it at whatever level your workflow needs. It needs no criteria.
curl -sS -X POST https://inference.do-ai.run/v1/systemone \
-H "Authorization: Bearer $MODEL_ACCESS_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe-jev-1.13.0",
"state": "My phone number is... nevermind I love running",
"questions": {
"contains_pii": {
"type": "noul",
"instructions": "Does this message contain Personally Identifiable Information (PII) such as phone numbers, physical addresses, or email addresses?"
}
}
}'
This is a good test case: the text mentions a phone number but never gives one. A keyword filter would flag it. A model that reads the sentence should not. I don't have the response for this request, so run it yourself and check the noul value.
Here is a real noul response, from asking whether the feedback "The customer said the delivery was late but the product quality was excellent." needs a follow-up from the support team:
{ "type": "noul", "noul": 0.68 }
A 0.68 is not a verdict. Whether it triggers a follow-up depends on the threshold you pick, which is the point.
Several questions in one request
The questions map takes any number of named entries. Jev evaluates them all against the same state in a single call, answers each independently, and returns a result per key. Types can be mixed.
curl -sS -X POST https://inference.do-ai.run/v1/systemone \
-H "Authorization: Bearer $MODEL_ACCESS_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe-jev-1.13.0",
"state": "The customer said the delivery was late but the product quality was excellent.",
"questions": {
"sentiment": {
"type": "choice",
"instructions": "What is the overall sentiment of this feedback?",
"criteria": {
"positive": "Customer is happy overall",
"negative": "Customer is unhappy overall",
"mixed": "Customer has both good and bad things to say"
}
},
"needs_followup": {
"type": "noul",
"instructions": "Does this feedback require a follow-up from the support team?"
}
}
}'
{
"model": "typesafe-jev-1.13.0",
"answers": {
"needs_followup": { "type": "noul", "noul": 0.68 },
"sentiment": {
"type": "choice",
"choice": "mixed",
"probabilities": { "mixed": 1, "negative": 0, "positive": 0 },
"confidence": 1
}
},
"usage": { "input_tokens": 368, "output_tokens": 59 }
}
One call gave a label and a gating probability. With an LLM that is two prompts or one prompt and a fragile parser.
Limits, pricing and access
- Context limits. Two separate limits apply to each request: 64K tokens for the whole request, and 32K tokens for the
stateplus the single longest question. Requests over either limit are rejected. - Billing. Jev bills per input token and output tokens are free, so cost scales with the size of the
stateand questions you send, not the length of the answer. Check the DigitalOcean docs for current rates. - No streaming, no multimodal. There is nothing to stream, since the answer arrives in one pass. Jev does not accept image, audio, or video input.
- Tier and terms. Jev is a third-party model, so your account must be on a qualifying tier. A request from a lower tier returns
this model is not available for your subscription tier. Use of Jev is subject to the TypeSafe AI Master Customer Agreement.
Where it fits
Use a decision model where the answer is a fixed set of outcomes and the cost of a malformed response is high:
- Routing: ticket triage, intent detection, picking which agent or tool handles a request.
- Moderation and safety gates: PII checks, policy checks, toxicity, run before text reaches an LLM or a user.
- Scoring: sentiment, urgency, lead quality, relevance filtering.
- Judging: a cheap first-pass grader in an eval pipeline. See Agent Model vs. Judge Model.
Keep the LLM for anything that needs to produce language or reason across many steps. The common pattern is Jev at the front door, deciding what kind of request this is and whether it is safe, and an LLM or agent behind it doing the work that needs generation. That is also the cleanest answer to the question in Agents vs. Structured Workflows: the branching points of a structured workflow are exactly the decisions a System One model is built for.
Caveats
- The classification is only as good as your question. The
instructionsandcriteriaare your prompt, so write them as carefully as you would any other. - Calibrated probabilities help, but validate on your own labeled examples before trusting thresholds in production. A
noulof 0.68 means different things in different workflows. - Text only, no streaming, and a 64K-token request ceiling. Long documents need chunking.
- It cannot explain itself. There is no rationale field, so log the probabilities if you need to audit decisions later.
- Keep API keys in environment variables or a secrets manager, never in scripts you commit.
The bottom line
An LLM brings a lot with it: token-by-token generation, free-form output, and a parser to catch the days the format breaks. Jev strips all of that away. It is incredibly fast, extremely cheap, and mathematically guarantees the output format.
So, if you are looking for a magical new chatbot, Jev is completely useless. But if you are a software engineer trying to build reliable if/else logic into an AI pipeline without it breaking in production, it's exactly the boring, predictable, 3-mode tool you actually want.
