Skip to main content
· Last updated

![Jev: The Working Man's Classifier](https://storage.googleapis.com/coolhand-public/og-images/jev-the-working-mans-classifier.png)

# Jev: The Working Man's Classifier

We keep seeing the same pattern across the teams we work with. Somewhere inside an
agent — or even inside a fairly simple AI feature — there's a step that's really a
traditional machine learning task wearing an LLM costume. Sort these records.
Classify this email. Decide which of ten repos is the likely home of a bug. And
instead of building a real classifier for it, the team reaches for Haiku, or Luna, or
whatever small model is already wired up.

That's not a mistake. It's a rational choice. Building an actual ML classifier means
labeling a dataset, running training jobs, babysitting a GPU, and maintaining a model
that only does one thing. Most fast-moving teams don't have the time or the headcount
for that, and they shouldn't have to. What they want is the working man's version of
a classifier: something that shows up, does the job well, and doesn't ask for a
research budget. LLMs have been a genuinely good fit for that job. They're just not a
cheap one.

**This is why we're paying attention to Jev**, the first System One model from
TypeSafe AI. Jev doesn't generate text at all — you hand it some state and a set of
typed questions, and it hands back structured, probability-backed answers instead of
prose. (We've written up the fuller technical difference between this kind of model
and an LLM in [LLMs and System One models](/help/llms-vs-system-one-models), if you want
the whole picture.)

The practical version is the price. At current catalog rates,
[Jev 1.13](https://coolhandlabs.com/inference-apis/typesafe-jev-1-13-0) runs **$0.042
per million input tokens with free output**, against **$1.00 in / $5.00 out** for
[Claude Haiku 4.5](https://coolhandlabs.com/inference-apis/anthropic-claude-haiku-4-5-20251001).
For a typical classification call — a couple thousand tokens in, a one-word answer
out — that works out to roughly **$0.0021 on Haiku against $0.000084 on Jev**, or
about **25× cheaper**. You can check both figures yourself in our
[inference API catalog](https://coolhandlabs.com/inference-apis), which is free and
needs no API key.

That is not an argument for swapping everything. It's an argument that there's a
genuinely new option on the table, and that figuring out where it fits is a job
Coolhand is built for.

## Working today: Jev traffic shows up in your logs

Coolhand ingests Jev calls now. If you're already sending us request logs, pointing
some traffic at TypeSafe means it lands in the same place as everything else:
captured, priced at the real rate, grouped into templates, and searchable next to
your LLM calls. Jev calls get their own `system_one` grouping so they don't get
tangled up with chat templates, and evaluations and correctness scoring work on them
normally. [Our TypeSafe guide](/help/typesafe-best-practices) covers what we capture, what
TypeSafe's responses leave out, and how to keep feedback matching to Jev outputs.

> **Why this matters:** you can start measuring this on your own traffic today rather
> than taking anyone's benchmark on faith — including ours.

## Working today: finding a swap and checking it holds up

### 1. Spotting where a swap saves real money

A lot of LLM calls in production are, structurally, classification or routing
decisions — an inbound email getting tagged by intent, for instance. Coolhand now
looks at what a call is actually doing, flags when it looks like a System One-shaped
decision, and estimates what you'd save on cost and latency by moving it.

> **Why this matters:** most teams don't have a clean inventory of which of their LLM
> calls are "real" reasoning and which are disguised classification. Surfacing that
> automatically turns "we should look into this sometime" into a prioritized list.

### 2. Checking that the swap doesn't cost you quality

Cheaper and faster only matters if the answer is still good. Before Coolhand
recommends a swap, it runs the same task both ways and scores both against your
existing correctness evals and the human feedback Coolhand already helps you collect —
so if Jev holds up, or wins, you see it in your own data. The result is reported with
the cost and latency difference in one line, like "same accuracy, 25x cheaper," and
when there isn't enough labeled history to judge quality, it says so instead of
guessing.

This is more than pointing the same request at a different model, because the two
don't take the same request shape — Coolhand translates the task into both shapes and
scores the answers. That's also why this isn't a bakeoff: bakeoffs swap one model id
for another, which can't work when the request itself is different.

> **Why this matters:** "TypeSafe says it's fast" isn't a migration plan. Your own
> evals and your own users' feedback are what tell you whether a swap is safe for
> your specific task.

## What we're building next

### Finding the classifier hiding inside a bigger call

Some of the best opportunities aren't standalone calls at all — they're one judgment
buried inside a much larger prompt to a thinking model. Take a "triage this
production bug" agent step: a big chunk of that prompt might really be answering one
narrow question, like which of several repos is the most likely place to start
investigating. That question can often be pulled out and answered by a fast, cheap
decision layer, leaving the LLM to spend its tokens on the part of the job that
actually needs judgment and prose.

> **Why this matters:** this is the least obvious of the three, and often the biggest
> win. Shrinking what you ask the expensive model to do is a different kind of
> savings than replacing it outright.

## The honest caveat

We're not going to pretend Jev is a drop-in replacement for an LLM everywhere a small
model is in use today. It doesn't generate text, it can't explain its reasoning, and
it takes 64k tokens per request — with a tighter 32k budget for the state plus the
longest single question — which is a fraction of what a modern LLM will swallow.
Plenty of steps genuinely need an LLM's depth, and part of what we're building is
exactly the analysis to tell you which is which, honestly, rather than assuming the
answer is always "use the new thing."

## Putting it together

We think System One models like Jev are a real step change for a specific and very
common problem: the classification and routing work teams have been quietly solving
with LLMs because building a real classifier felt like too much. Today you can send
that traffic through Coolhand, see what it actually costs you, and get told when a
swap would save real money and whether it holds up on quality. Next, we're building
the part that finds where a big LLM call has a smaller decision hiding inside it.

If you want the deeper technical rundown on how System One models differ from LLMs,
that's in [LLMs and System One models](/help/llms-vs-system-one-models) — and if you try
Jev on something we haven't thought of, we'd like to hear about it.