← All articles

Article · 26 September 2026 · 6 min read

Jev hands you a probability. Caveat Emptor.

By Doryan Gowty, Principal, Anneal

So all the talk over the last week seems to be around TypeSafe AI's release of Jev, their "System One" model built to give structured output for classification problems. There is a lot of excitement about the cost and latency benefits it offers.

I don't plan to go into great detail about Jev in this post. Prosper Otemuyiwa (unicodeveloper) has written an excellent article covering it over at Medium. In short, it's much faster and cheaper than using an LLM for the same task, you define the categories you want considered, and enumerated, when you send the call to the model. That's very attractive when a classifier sits inside your workflow, because vagaries, hallucinations and inconsistent category labels no longer need exception handling.

LLM agent loop vs typed response model Two flows for classifying an inbound email. With an LLM, the free-text answer is validated, retried when the label is invalid, and sent to a person when retries run out. With Jev, the model returns one of the supplied categories with a probability, and a cut-off decides whether to act on the label or send it for review. LLM AGENT LOOP NO · RETRY YES RETRIES OUT Email infree text LLM callprompt + parse Valid label? Act on label Personexception queue TYPED RESPONSE MODEL ABOVE BELOW Email in+ categories Jevlabel + p p ≥ cut-off? Act on label Reviewperson or fallback
With an LLM, you validate its free-text answer, retry when it invents a label and hand the email to a person when retries run out. Jev can only return one of your categories, but the cut-off that decides what happens next is still yours to set.

Classifiers aren't new. Modelled output has been used to handle uncertainty in decision workflows for decades. A credit scorecard and the decision strategy built on top of it, follow exactly this pattern. My background in banking focused on building models and setting strategies for this exact purpose. Zero-shot classifiers, trained on large general datasets, already exist, and there are packages that put a structured layer over an LLM so it returns a typed response the way Jev does.

What interests me in the current Jev discussion is that it's where data science and AI engineering intersect. Working on the same problem from opposite ends. I've written recently about decision models and putting LLMs in your workflow, and Jev sits right where those two pieces overlap and I'm keen to see where developers put it to use. In the agentic workflows I've built, a consistent output with a measure of confidence attached is very attractive. I can also see the temptation to use it where consequential decisions are made, such as a fraud call that limits someone's access to their accounts, or an underwriting decision.

The model is the easy part

Getting a classifier to return a label was never the hard part of the job. In a risk setting, the work is choosing the cut-off: the score below which you decline. A well-calibrated model lets you put a number on the cost of making a wrong decision. Most of the effort in deploying the pipeline goes into scrutinising the model's performance and the strategy settings.

Benchmarking of Jev's performance is already happening. Here are some examples such as this one which compares Jev with some traditional ML classifiers, and another that compares its classification performance with an LLM. So far it looks like Jev does fine across a range of problems, but purpose-built models trained on labelled data still perform better in most cases.

More of this testing will need to happen, since TypeSafe have said they won't report performance against public benchmarks. Jev returns probabilities and a confidence score with its output, and that's where the work begins. Knowing a model is right 90% of the time might give you some comfort that you can rely on it, but what is the other 10% costing you?

The cost of being wrong isn't symmetric. In pricing, finding the optimal offer for a given segment meant weighing the direct profit hit of offering a discount against the chance of losing the deal without it. A lost deal never shows up as a line in the P&L, but it's a real cost. A classifier knows none of this. It knows which label is most likely, and whether "most likely" is good enough depends on what it costs you to be wrong either way.

Too eager to please

A real shortcoming with any AI model is that it can be too eager to please. LLMs will give a response with supreme confidence whether it's correct or not. In my own work for the DTC beverage brand, I've seen LLMs guess at actions even when they have the data available to make the right call. Typed response models don't entirely remove this risk, because they still have to provide an answer. When bound to a set of defined responses, the model will simply favour the least worst option.

Jev has the same tendency in a tidier form. If none of the options you pass it fit the input, it still has to pick one, and you get the least worst label from your schema. The fix is to give it somewhere honest to land: an explicit none_of_the_above or unknown class, and a path in your workflow that knows what to do with it. Inputs that look nothing like what the model learned from are a harder version of the same problem. Its internal mapping of your categories breaks down, and it picks wrong labels with no sign that anything has changed.

The returned probabilities are the main defence. Your downstream code can reject any decision below a confidence threshold, say 85%, and send it to a person or a fallback. That threshold is a cut-off like any other, and it should be set against what a wrong call costs you. It also only catches the errors the model is unsure about. An input from well outside its training dataset can come back wrong with high confidence, so you still need to check calibration on your own data.

Traceability matters

Imagine a classifier making fraud decisions that flags a transaction and blocks a card. What if this card belongs to someone in senior leadership. You can guarantee somebody will ask what happened, and "the model said so" won't cut it.

The questions are predictable. What information did the model rely on? What score did it return? What was the cut-off at the time, who set it, and when did it last change? Which version of the model made the call, and is that the version running today? In a bank decisioning environment every one of those has an answer on file, because regulators and model risk teams make sure of it. Explainability is one of the bars modelling teams have to clear when choosing which technique to train a classifier with.

When the classifier is a hosted API call, some of those answers get harder to produce. You can log the input, the label, the probability and the cut-off you applied. You can't see how it transforms your input or which parts of it the decision relied on. Transparency about training data, model versions and retraining is entirely up to the vendor.

Where Jev fits

As I said at the start, Jev isn't a new idea. DSPy can already make an LLM return one of a fixed set of labels through a typed signature, and the major model APIs can constrain a response to an enum. Zero-shot classifiers, like the NLI-based models on Hugging Face, have handled "here are my categories, pick one" for years. What Jev claims is that it does the job far faster and more cheaply: 70 to 500 milliseconds end to end and $0.042 per million input tokens. If that holds up on your data, it matters anywhere you classify at volume, such as routing inbound requests, triaging tickets or tagging transactions before a rule fires.

Speed of deployment matters too. Getting a classifier into production used to mean labelled data, training, validation and a release process measured in months. Declaring your categories in a request and having something working that afternoon is a real change. A System One model like Jev makes perfect sense where it's replacing tasks LLMs are already doing poorly. There may even be cases where no training data exists and a zero-shot classifier can get an MVP up and running.

What doesn't change is the work around it. Someone still has to set thresholds with the cost of being wrong in each direction in mind, check the probabilities are calibrated on their own data, notice when the world drifts away from the model, and keep a record good enough to answer "why" when someone important asks. That's the decision science half of the job and AI engineering is increasingly getting involved in it.

If you're putting Jev into a workflow and working through the cut-off question, I'd like to hear how you're approaching it.