TLDR
Most of the AI calls running inside a business aren’t hard questions. They’re tiny decisions made thousands of times: is this message urgent, which folder does this go in, did that job actually finish. For the last few years we’ve been answering those with big language models that write paragraphs we then have to parse. In mid-September a company called TypeSafe released a different kind of model called Jev. It doesn’t write. It picks, and tells you how sure it is. I spent a week wiring it into my system. 3,342 calls cost a total of about twenty cents, it beat my keyword matching on a real test, and the most important thing I learned had nothing to do with the model.
The Genius at the Mail Slot
Here’s a line from Layton Gott’s post that made the rounds when Jev launched, and it’s the cleanest description of the problem I’ve seen:
“AI agents waste an unbelievable amount of intelligence on tiny decisions. Should I retry? Which tool should I use? Did this task actually finish? Is this action risky?”
That was my system. Every time a new prompt came in, something had to decide which of my 100-plus skills might help. Every time a memory search ran, something had to decide which results actually mattered. The honest options were a keyword match, which is fast and dumb, or a full language model call, which is smart, slow, costs real money, and answers in a paragraph that code then has to pick apart and hope it followed the format.
That’s like hiring a genius to stand at the mail slot and read every envelope out loud to you. The genius will do it. It’s just a strange use of a genius.
What Jev Actually Is
TypeSafe calls Jev a “System One” model, borrowing Daniel Kahneman’s split between fast, intuitive judgment and slow, deliberate reasoning. Reasoning models are System Two. Jev is deliberately not.
Their own description: “System One models… return typed answers and probabilities rather than generating text or reasoning explanations. Code owns the workflow; the model supplies programmable common sense where ordinary code needs semantic understanding.”
In plain terms: you send it the facts and one or more named questions, and every question comes back as a typed answer with a probability attached. There are three kinds of question:
- Choice: pick one option from a list you define, like a switch statement in code
- Score: place something on a scale, like sorting or a threshold
- Noul: the probability that a yes-or-no statement is true, like an if statement
No prose. Nothing to parse. Nate B. Jones put it bluntly: “You literally can’t chat with Jev.”
It’s cheap. At launch the price was $0.042 per million input tokens, with output free. One call that ranked 111 skills against a prompt, about 6,300 tokens of input, cost $0.00027.
Two properties matter a lot in practice:
- It’s calibrated, not deterministic. I asked the same question three times and got 0.65, 0.69, and 0.66. Same answer, slightly different confidence. So in my setup, a threshold within about 0.05 of a result counts as a coin flip.
- It never refuses. There’s no safety filter at the model layer. TypeSafe’s CEO argued in a long podcast interview that a refusal inside a software dependency behaves like a random error that breaks the calling code, so safety belongs in the questions you write. That’s my paraphrase, not his words. The practical upshot: if a decision is safety-related, you write that question yourself, and you own the threshold.
Where It Runs in My System
I didn’t rip anything out. Jev went in as an extra layer on top of what already worked, and it earns its place or it doesn’t.
Picking the right skill. Every prompt I send gets one Jev call that ranks my full skill roster against it, plus a yes-or-no on whether a skill is needed at all. It sits alongside the old keyword layer, and its picks lead when it’s confident.
Finding the right memory. A fast keyword pass narrows about 90 memory candidates to a shortlist, and Jev reorders the shortlist by actual relevance to the prompt.
Research. My research skill makes five small typed judgments along the way: which past runs matter, whether a question even needs a research run, which angles to prune, whether sources agree claim by claim, and how risky each recommendation is. The claim-by-claim check replaced a full language model agent that used to do the same job. A whole run’s worth of these costs about a tenth of a cent.
A coach that decides where Jev goes. I built a tool that splits a task into its decision steps and grades each one on whether it’s a good fit for a model like this. More on that below.
And one I’ve built and measured but haven’t switched on: a screen for my tarot app that reads each user message before the reading model does, and flags crisis language or attempts to hijack the prompt. On a 60-message test set it caught 19 of 20 crisis messages and 19 of 20 injection attempts, with zero false alarms on the tricky-but-harmless messages, at about two-thousandths of a cent per message and a median of 203 milliseconds.
The Four Questions Before You Use It
The best filter I found for “should Jev do this” comes from Nate B. Jones’s guide to what he calls Jev-shaped problems. All four have to be yes:
- Can you list the possible answers before you see the input?
- Does it happen a lot, or hold up a person waiting on it?
- Does code do something different depending on the answer? Route it, flag it, pause it, store it.
- Does a low-confidence answer have somewhere to go? A person, a review pile, a bigger model.
And two things kill it no matter what: the step needs something newly written that a human will read, or the decision is really arithmetic, counting, or dates. Code does those better than any model.
I wanted help applying that, so I built the coach. It splits a task into steps, grades each step on those checks, and gives a verdict: JEV, NOT JEV, or BORDERLINE, where borderline means “not enough information yet,” never “bad fit.” The verdict is computed by code from the answers, not eyeballed.
By the end of its first day, the coach’s ledger had 99 graded steps: 13 JEV, 15 BORDERLINE, and 71 NOT JEV.
Read that again. The most common answer from a tool built to find Jev work was “no.” That’s the tool working. Most steps in most workflows need writing, or math, or judgment that shouldn’t be automated at all. The win is finding the handful that are pure, repeated, listable decisions, and taking the genius off those.
The Numbers After One Week
Between September 16 and September 22, my system made 3,342 Jev calls. Every one was answered by the same engine version. Total spend: $0.1993.
Here’s the number that surprised me more than the price. Only 338 of those calls, about 10%, were the live work: the skill picker and the memory search running in front of real prompts. The other 90% was me and my AI proving it worked: accuracy rechecks, replaying real cases, a 300-case screening test, bake-offs. Getting a cheap, fast classifier trustworthy enough to run live took roughly nine calls of proof for every call of real use. Budget for that.
The one real head-to-head test: 60 real prompts pulled from 30 days of actual use, judged blind by a separate AI, with a control to measure the judge’s own noise.
- Precision in the top three results: Jev 0.807, keyword matching 0.727
- Found a memory that was actually needed: Jev 44 of 60 prompts, keywords 37 of 60
- Came back with nothing at all: Jev 0 of 60, keywords 6 of 60
- Where it missed the goal: I wanted it to cut irrelevant results in half. It cut them by 29%.
And a robustness check I ran after hearing the CEO describe it: I stuffed random IDs, hex strings, and junk timestamps into three real requests, fifteen calls total. Zero answers flipped. The biggest probability shift was 0.025.
One number went the wrong way. On September 16, calls took 131 to 392 milliseconds. Six days later, the same kind of call took 871 to 1,372 milliseconds, whether the input was 405 tokens or 6,000. Nobody on my side knows why yet. If you’re designing anything real-time on top of a young model, measure speed the day you ship, not the week before.
What Went Wrong
The wording mattered more than the threshold. This was the biggest practical lesson, and I found it three separate ways:
- A vaguely worded question (“would this change how someone answers the question?”) scored an exact duplicate 0.62, near the bottom of a twelve-item ranking. Asking “is this about the same subject?” scored the same duplicate 0.96.
- Asking only “does this source support the claim?” made two off-topic sources look like they disagreed with it.
- One question scored 0.75 when it got only the facts it needed, and 0.38 when the same facts came with extra fields attached. More context made it worse.
Every broken version still returned confident, well-formed numbers. A bad question doesn’t look bad. You only find it by testing on real cases before you argue about thresholds.
One job took five versions and still isn’t right. The first real job I gave the coach was triaging my open decision queue: is this item still live, who owns it, is it finished. The first version called a known-dead item “live” at 0.80 to 0.92. Later versions fixed that and got too eager to close things, then too timid. On the only three items with a confirmed real-world outcome so far, the fifth version matched one of three. It’s still in shadow mode, which is exactly where it belongs.
The code around the model needed real engineering. When I had Codex review my wrapper scripts, it found 13 blockers and 5 risks across three rounds: the model call happening before the code checked whether that call was even allowed, private text able to leak into a context field, a budget of zero that didn’t actually block anything. All fixed. But anyone pitching “just add a cheap classifier” as a weekend job is skipping this part.
Two AIs graded the same tasks very differently. Same five tasks, same coaching method. Claude split them into 9 steps and called 3 of them Jev-shaped. Codex split them into 18 steps and called zero of them Jev-shaped outright, because it refused to guess at facts it hadn’t been shown, like volume and where low-confidence answers would go. Charitably, that’s the system working: unknown stays unknown. Plainly, the same method gives different answers depending on who’s applying it.
The obvious win that wasn’t. One of my products has a public classification endpoint running on a regular language model. Perfect swap candidate. Then I checked the traffic: two lifetime requests. Before you replace an expensive call with a cheap one, make sure anybody’s making the call.
The Biggest Lesson Wasn’t About Jev
After a full day of watching the live skill picker, the Jev layer suggested a skill on about 89% of prompts. The old keyword layer suggested one on about 7%.
Neither suggestion got used. Zero adoptions, either layer, across 84 turns.
That’s not a model problem. It’s a plumbing problem. A more precise classifier doesn’t matter at all if nothing downstream acts on what it says. Check number three on the list, “does code do something different depending on the answer,” turned out to be the one that decides whether any of this is worth doing.
And there’s a quieter lesson in how I use it now. After a few days I noticed I was relying on myself to remember a list of best practices for writing these questions. So I said it out loud: I want to “rely less on my ability to remember any of this.” Now a linter checks every Jev job against six rules before it can run: small repeated choices only, one decision per question, only the facts a question needs, three real test cases before any threshold talk, the model judges and code acts, and every miss becomes a saved test. The rules live in code, not in my head.
Where a Small Model Fits in Your Business
You don’t need my setup to use this idea. You need to find the tiny decisions you’re currently paying a genius for, or worse, paying a person for.
Good fits:
- Routing incoming leads or support messages to the right person or pipeline
- Flagging a message as urgent, angry, or needing a human
- Screening form submissions for spam or prompt-injection attempts before an AI replies
- Checking whether an automated job actually finished, instead of searching a log for the word “BLOCKED”
- Sorting reviews or feedback into a handful of buckets you define
Bad fits:
- Anything that needs a written reply, a summary, or an explanation
- Anything that’s really math, counting, or dates
- Anything nothing will act on
The pattern that worked for me: keep the rules you have, add the small model as a layer, let low-confidence answers go to a person, and prove it on real cases before you trust it with anything that sends, deletes, or charges.
The future of AI inside a business probably isn’t one giant brain doing everything. It’s a lot of small, cheap, honest judgments, each doing one job, with code in charge and a person catching the maybes.
I appreciate you reading through this one. If you’ve got a pile of tiny decisions eating your team’s time, or a lead pipeline that needs a smarter router, book a call from the home page. And if you want the bigger picture of how this fits into the rest of my system, the harness post is where to start.
Build first, learn fast.
Keep reading
- Rules Didn't Work, So I Built a Harness (What It Actually Takes to Give AI Agents Structure)Seven months ago my AI setup was one identity file and a journal. Today it's 35 hooks across 12 lifecycle events, 62 situational rules, about 700 memories that maintain themselves, and a queue for the decisions that are actually mine. Every piece exists because something went wrong without it.
- How I Built an AI Agent That Dreams, Remembers, and Inherited My Worst HabitFrom the first planning conversation to first boot: about 24 hours. The actual coding took 50 minutes. She runs 24/7, has 7,877 memory vectors, screens Upwork jobs while I sleep, and has a nightly dream mode. The most interesting thing about her isn't what she can do - it's what she inherited from me.
- What Your AI Actually Needs From You (And Won't Ask For)I asked my Claude what would make our working relationship better. Instead of requesting better prompts or more context, it pointed to three gaps most AI users aren't thinking about - and the infrastructure to close them is simpler than you'd expect.