TL;DR :
A new kind of small model (Jev, from TypeSafe AI) does one thing: pick an answer from a fixed set, fast and almost free. It doesn't write. Right now engineers use it mostly as AI checking AI, like screening 17,000 agent commands and flagging 42 for a human, with about 7 in 8 flags correct.
Why it matters: checking used to mean calling another big model every time, so most teams checked a sample or nothing. At roughly a two-hundredth of the cost, you can check every step.
Three limits: it only knows what you show it; its mistakes repeat identically at scale; and someone still has to check its flags and what it waved through.
What to do: pick the one action that hurts most if it goes wrong, put the checker beside your existing approval rather than replacing it, write the risk rules in plain words, let it say "not sure," and count flags and misses from week one.Someone running an AI coding agent put a small, new kind of model in front of it with one job: before each command runs, decide if it's safe. Out of roughly 17,000 commands, it flagged 42 as unsafe and held them back for a person to look at. About seven in eight of those flags were right.
That model was Jev, released this month by a startup called TypeSafe AI. Unlike other models we’ve all been using recently, it doesn't write. You give it some text and a question with a fixed set of answers. It picks one and tells you how sure it is, in a fraction of a second, for next to nothing.
Right now, the people using Jev are mostly engineers, and most of what they've built is AI checking AI. Vercel uses it to check the commands its agent wants to run. LangChain and Langfuse use it to grade answers from other AI systems. Others use it to test whether an AI's source actually says what the AI claims. We used it in a similar way—scoring a few thousand home pages, so our knowledge base would know what a good page looks like before suggesting changes to our own.
We think this is where models like Jev will land first, and the reason is cost. Companies are giving agents real work: sending emails, updating records, running code. Any of those steps can go wrong. Until now, checking a step meant calling another large model every time, so in our experience most teams checked a sample, or nothing. One test found Jev gave the same answer as a leading model about nine times in ten, at about a two-hundredth of the cost. At that price, you can check every step.
It won't stay with engineers for long. New tools usually start with technical people and reach everyone else later. This is how diffusion works. Think of the clerk who approves routine invoices at a glance, or the person who reads support emails to decide which ones are urgent. Those are the same kind of quick, yes-or-no calls. If that happens, their job changes the same way it has for the developer in the coding example. The model does the checking, and they look at what it flags.
So what gets better? The model can check everything instead of a sample: every action before it runs, every answer a customer gets against its source, every email before an agent reads it. People only see what it flags.
What doesn't? Three things.
The checker only knows what you show it. Show it a refund without the customer's history, and it will pass a refund that looks right but goes to the wrong person.
Its mistakes repeat. A person who misreads a rule gets it wrong a few times before someone corrects them. A checker gets it wrong on every call, thousands of times, until someone notices.
Its flags need checking, and so does what it lets through. In the coding example, about one flag in eight was wrong: the command was actually safe. Nobody knows how many of the other 17,000 should have been flagged. Today a person looks at these. Soon a bigger model could do most of that. But someone still has to decide what the checker got right and what it missed, and keep count. If nobody does, nobody is really checking anymore.
None of this is a reason to wait. It's a reason to start with one action. Pick the one that would hurt most if it went wrong: sending money, deleting records, writing to a customer. Put the checker next to the approval you already have, not in place of it. Write down in plain words what makes the action risky, with a few examples, and let the checker answer "not sure." Then keep that count from the first week: how many flags, how many were right, and what it let through.