Keep this guide

Download the System One guide

An editorial PDF and a concise, visual PowerPoint handbook.

Okay, this one is actually very interesting.

For the last few years pretty much all AI products were built around LLMs. You send something in, model generates something back, maybe it gives you JSON, then your software uses whatever came out of it. Which works perfectly fine for many things, but also means we sometimes call some massive model just to answer a ridiculously small question like whether something is suspicious or where exactly this request should go.

Now there is Jev, created by TypeSafe, and they call this new class System One Models. Jev basically doesn’t talk to you. You give it some state and ask it to make a certain decision, and it returns the choice together with probabilities/confidence. So imagine something closer to an extremely smart decision function rather than another chatbot. TypeSafe’s introduction ↗

My first reaction was honestly kinda “okay, isn’t this basically a classifier?”. And well, in some sense yeah. But then I started thinking about how many stupidly expensive LLM calls in real products are basically classifiers wearing a $20 suit.

Let’s say you have some SaaS where every incoming request currently goes to a frontier model. Some requests are trivial. Some should go to a cheaper model. Some require human review. Some might be suspicious. Jev can make this first decision very fast and then your actual expensive model only wakes up when you really need it.

request → Jev → decision → whatever actually needs to happen

TypeSafe's real System One workflow diagram: triage an alert, choose a disposition, evaluate containment, then select an action
TypeSafe’s example of decisions inside a larger software workflow. This is a vendor diagram, not an independent test. See the original and its explanation ↗

There is one Reddit example I liked exactly because it is so simple. They had a big context going into an LLM basically to find out whether there was something of concern. After moving this first check to Jev they reported going from 3 to 6 seconds to milliseconds, while the bigger LLM is only called if Jev finds something worth checking further. That’s one user’s report, not a controlled benchmark. Read the discussion ↗

And honestly, for product people this might be more interesting than some new model scoring 2% higher on another benchmark. Product can become faster. Your inference bill can become smaller. You can also have these small decisions everywhere without feeling like you’re sending every tiny thing to Godzilla.

You don’t necessarily need the smartest giant model for every tiny decision in the product. Sometimes what you need is basically a very intelligent if statement.

There is also Liquid d1 now, built for this general type of decision workload as well. So at least we’re starting to see more than one implementation of the idea, which makes me curious where this whole category goes. Liquid’s documentation shows the same basic shape: state in; a Choice, Score, or yes/no Noul answer out. Read Liquid’s documentation ↗

LLMs still make perfect sense when I need something generated, explained, transformed, summarized etc. But if somewhere in my architecture I have an LLM call whose final useful output is literally true, “sales” or 0.83, I’d definitely start questioning whether that call should stay there. A Hugging Face community article about Jev makes basically this distinction around bounded outputs and software decisions. Read the community article ↗

TypeSafe's chart comparing accuracy and cost of Jev with several other models across four company-selected workflows
TypeSafe’s own accuracy-versus-cost evaluation, across four workflows it selected. Useful to inspect, but not independent validation. Read the methodology and caveats ↗

Now obviously don’t rebuild the whole company around this tomorrow. Jev is new, a lot of the impressive performance numbers come from TypeSafe themselves, and the Reddit thread has people saying this is glorified classification and/or heavily overhyped. Fair enough. We need more independent testing before declaring some completely new era of AI architecture. TypeSafe’s results ↗ · Reddit discussion ↗

But I would absolutely play with it if you are building AI products.

Especially if you look at your LLM usage and suddenly realize half of it exists just to make tiny decisions.

Keep this guide

Download the System One guide

An editorial PDF and a concise, visual PowerPoint handbook.