JEV AND THE CASE FOR AI THAT JUST MAKES THE DECISION.

On 15 September TypeSafe AI released a model that can't write a sentence, and it puts into perspective how much the industry is trying to fit LLMs where their downsides make them challenging to work with. Instead of traditional ‘next most likely word’ prediction, Jev is a more rigid form of AI model that might just work with our systems instead of against them.

Jev (named after the Victorian economist William Stanley Jevons - more on that below) is still in early access, but its launch video has passed 40 million views on X and TypeSafe's philosophy of "prod, not god" could be the most practically useful idea in AI this year - with a couple of caveats, because the launch marketing is running a little ahead of the evidence.

Jev answers questions in a very different way

You define a question and the possible answers in advance - which intent category does this search query fit, how urgent is this email on a scale of one to five, how likely is this ad copy to break the brand guidelines - and Jev returns one of your answers with a probability attached. Because the answer always comes from your list, it can go straight into your code without any cleaning up.

That constraint is also why it's fast and cheap. TypeSafe quotes 70 to 500 milliseconds per response against 3 to 329 seconds for the frontier models they tested, and The Batch reported that on TypeSafe's internal datasets Jev matched GPT-5.6 Terra and Claude Sonnet 5 on accuracy (67%) at $0.0007 per example against $0.06 and $0.12 - roughly 85 and 170 times cheaper (though those are TypeSafe's own datasets, and they haven't published external benchmarks yet).

"Prod, not god" puts the big models back in their lane

TypeSafe's argument is that we've started treating frontier models as god instead of production tools, sending them every task with fuzzy inputs and paying for all that generality even when the job only needed a yes or no. Their alternative is software where code owns the workflow, and AI makes narrow, structured decisions that are ready for ‘prod’ (CEO Diogo Almeida goes into it on the Latent Space podcast).

The big models are still the right tool for plenty - Jev can't write copy, hold a conversation, or act as an agent. But look across a typical SEO, paid media, or data workload and a lot of the AI jobs are decisions in disguise: classifying search queries by intent, splitting brand from non-brand, tagging reviews by emotion, routing leads. Plenty of those don't get done at all right now, because running them across every row costs too much.

That's where the name comes in. Jevons noticed in 1865 that more efficient steam engines increased total coal use, because cheaper energy made new uses worthwhile. At The Batch's per-example rates, classifying 100,000 search queries would cost around $70 with Jev against $6,000 to $12,000 with the frontier models - cheap enough to run on everything, every week, which is where I think the biggest opportunity sits.

As AI moves from passive answer engines to active project managers, a trend our Q1 2026 SearchPulse report terms the rise of 'Delegated Choice', models like Jev prove that specialised, structured decision-making will power the background workflows that get systems 'agent-ready'.

The psychology behind System One

TypeSafe took the "System One" label from Daniel Kahneman's *Thinking, Fast and Slow*, and it fits. System 1 is our fast, automatic, always-on thinking, while System 2 is slow, effortful and used sparingly - it only steps in when something feels off. Marketers have designed for their audiences' System 1 for years (it's the basis of nudge theory), and my colleague Lucy Todd has written about how the same split shapes search behaviour.

Applied to AI it's a sensible blueprint: a fast, cheap model makes the reflex calls, its confidence score flags when something's off, and only the uncertain cases go to a bigger model or a person. Most businesses have been running System 2 on everything (the organisational equivalent of consciously deciding which foot to lead with on the way to the kettle).

System 1 is also where the mistakes live

Biases are System 1 answers that nobody checked, and Jev's version is being confidently wrong. TypeSafe's "can't hallucinate" claim holds in the narrow sense that every answer comes from your list - it can still pick the wrong one from it. An independent Towards Data Science test on 3,080 banking support messages found Jev beat the Qwen model on accuracy (81.1% against 76.4%), but when it reported confidence between 0.7 and 0.9 it was right only 53% of the time, and passing those uncertain cases to an LLM often made things worse.

That matters because people tend to over-trust automated outputs (Parasuraman and Riley's 1997 work on automation misuse is the classic reference), and a precise-looking confidence score only encourages it. The categories matter too - if yours are wrong, or there's no "other" option, Jev will misfile things at thousands a second. Deciding what those categories should be is a question about human behaviour, and it's the same starting point behind HumanLens.

Start with one decision and test it properly

Jev is still very new, and TypeSafe themselves say their headline figures (193.6x faster, 444.6x cheaper) are likely to sit at the high end of real-world results. My suggestion is to pick one high-volume decision your team makes today, label a few hundred real examples, and test Jev against your current approach, checking accuracy at each confidence level as well as overall. Good places to start might be:

- intent classification for search queries, feeding into SEO prioritisation

- tagging brand mentions in AI answers for LLM tracking

- in-session decisions on which content or offer to show a visitor (CRO)

- guardrail checks on AI-generated copy against your brand guidelines

Earlier this year I wrote that 2026 would be the year boring wins, and the pragmatic approach behind Jev is exactly the kind of practical development that I was talking about.

Navigating the shift from generative tools to structured AI decision-making requires a balance of technical precision and behavioural insight. Whether you're looking to optimise your site for answer engines, audit your search workflows, or leverage behavioural science in your marketing, our team is here to help. Get in touch with our team or explore our latest SearchPulse report to stay ahead of shifting consumer search habits.

Contact Us
matt-g-image-01

MEET THE
AUTHOR.

MATT GREENWOOD-WILKINS

Matt is a data and spreadsheet nerd. Having worked in data pipeline engineering, business intelligence and data analysis - he helps us manage and understand data to generate interesting and actionable insights. He helps to drive efficiencies both internally and for clients, creating innovative solutions using automation, machine learning and AI.

More about Matt
hellofresh-logo2
brakes_logo.svg
sunsail
uktv
nidostudent
rspca_logo

Have a project you would like to discuss?