551 Things People Built With Jev in 10 Days. The Ones That Will Make Money Are Boring.

Doom bots and trading demos got the attention. Triage, model routers and shell guardrails are where $0.042 per million tokens pays.

11 min read

Jev, the classifier TypeSafe launched on September 15, is a new kind of LLM: it doesn't write, it decides. You give it a state and closed questions (yes or no, a pick between options, a score), and it returns probabilities for $0.042 per million input tokens, with output free. In other words, it replaces the if in your code that needed to understand text.

The best uses published since then are boring, and that's a good sign. A research pipeline sorted 1,018 papers into 24 themes for $0.08, while the LLM summaries in the same pipeline cost $3.99. Someone ran a resume against all 6,245 YC companies in 25 seconds for $0.37 and came out with 156 founders to contact. In a shell guardrail, an ambiguous rm -rf came out "irreversible" with a confidence of 0.33, low enough for the code to ask a human before running it. On the production side, Metaview put Jev in all of its agents over a weekend and reports candidate searches about 10x faster, at the same accuracy.

Office worker shows off flashy robots and gadgets; a caped hero points to a simple label sorter making money.
551 demos built. One actually makes money. Guess which one's boring?

551 Builds and a Single Production Report

As of September 25, shipwithjev, an independent catalogue of Jev projects, lists 551 builds, 10 days after launch. Tools and apps lead with 213, followed by agents and browsers (79), research and data (72), games and real-time demos (66), content and growth (54), triage and routing (50), trading (10) and robotics (7). The entries come from 183 posts on X, 26 Reddit posts and 278 GitHub repos.

In all that material, a single production report has a company name attached, and it's the Metaview one. Shahriar Tajbakhsh, who posted it on X, adds a detail that matters more than the speedup: his team now fixes bugs by adding a question to Jev instead of rewriting a system prompt. Flavio Copes, in his deep dive on Jev, presents it as the first production feedback on the model.

Before trusting any figure below, you should know that shipwithjev isn't affiliated with TypeSafe and repeats what builders announce without measuring anything. Every number that follows belongs to whoever posted it.

The demos got the spotlight. Which builds are still running once the novelty wears off is a different question, and the answer starts with what Jev actually is.

A Smart If Statement, Not a Chatbot

Jev takes 2 things: a state (a ticket, a paper, a shell command, a web page) and a list of closed questions. For each answer it returns a probability, and in some builds a confidence score as well. It never writes a sentence.

Take a hypothetical support ticket. A single call can ask which category it belongs to, how severe it is, and whether the customer wants a refund. You get probabilities back that your code can branch on directly, with no paragraph to parse and no JSON that forgot a field.

TypeSafe prices it at $0.042 per million input tokens with output free, and announces latencies between 70 and 500 ms per call (as of September 2026). The launch post also puts Jev at 193.6x faster and 444.6x cheaper in its workflow evals. Those are the vendor's own benchmarks, run on workflows its own team wrote, and TypeSafe says itself that the multipliers sit at the high end of what users will see and that the setup may be biased.

What Jev can't do matters as much: it doesn't write, count, do math, compare dates or read images. (So it will never tell you "I'm afraid I can't do that, Dave". It will just hand you a low probability.)

That shape gives a simple filter. A use case pays when a narrow decision, cheap to get wrong or easy to catch, repeats thousands of times on a volume someone already pays to process. A Doom bot makes about 10 decisions a second for an audience of 0 paying customers, while ticket triage makes a single decision per ticket for a support team that might handle 100,000 tickets a month (a made-up volume, for scale) and already pays people to sort them. Each wrong label lands in a queue where a human can fix it.

The Boring Filter: Three-Layer Content Screening System
The Boring Filter: Three-Layer Content Screening System

Run the 551 builds through that filter and the catalogue looks odd: 66 of them are games and real-time demos.

The Demo Pile

The official Doom demo sets the tone. It runs at about 10 requests per second, which costs about $7 an hour, and TypeSafe admits in its own launch post that a non-AI bot would play better. Around it, the pile grew fast: a typesafe-mario project, a simulated town of 79 Jev agents that shipwithjev lists at $0.009 in total, and a jev-trader bot that a roundup by @0x_rody clocks at about 81 ms on the testnet of Monad (a blockchain, and a testnet is its practice version with fake money).

These demos travel for a good reason. Jev's real selling point is speed, and speed is invisible in a ticket queue. A ticket getting the right label in a fraction of a second makes a dull screenshot, while a game character reacting 10 times a second shows the latency with no explanation needed.

They make a poor business, though. A Doom bot that loses to a script is a $7-an-hour screensaver. Trading has the opposite problem: the buyer exists, but a latency figure measured with fake money says nothing about how often the bot is right once real money sits on the other side of the trade. Fast on testnet is the trading version of "it works on my machine". Of the 551 builds, 10 are trading projects.

The Boring Stack That Pays

The builds that pass the filter share a trait. Each one replaces a question someone used to answer by hand, or used to pay an LLM to answer in prose.

  • Labeling research. The paper sorter ran at a median of 256 ms per paper. The decision it replaces is "which of these 24 themes fits this paper", asked once per paper. The expensive part of that pipeline was the LLM summary, and a summary still needs someone to read it before anything gets sorted.
  • Prospecting. The resume run asks a single question per YC company: does this profile fit? Its shortlist works out to a 2.5% hit rate, a list short enough for a human to go through by hand.
  • Ad analysis. According to Flavio Copes's roundup, a builder classified 724 ads from 37 brands in about 40 seconds for 9 cents. Each ad gets closed questions instead of a marketing essay, so the output drops straight into a spreadsheet.
  • Surveys. A survey product rebuilt on Jev, posted on Reddit and catalogued by shipwithjev, reports running 2x faster and 85% cheaper than on Gemini 3.5 Flash-Lite. Free-text survey answers are the textbook case, since each of thousands of responses needs a category.
  • Existing platforms. The catalogue also lists Jev integrations for platforms that already run, like a Magento 2 module and Spliit Cloud. They're the least spectacular entries and the most interesting ones for the filter, because the volume exists before Jev shows up.

All of these figures are the authors' own, and cost per token is the wrong unit anyway. The number that matters is cost per resolved task, human review and retries included. A classifier that's 50 times cheaper but wrong 1 time in 10 (a hypothetical rate) can end up losing to the LLM it replaced. Per-token pricing stops mattering the moment a human cleans up every tenth answer.

Every item on that list sorts data that sits still. An agent that deletes files, books flights and deploys code is a more expensive place to be wrong.

Inside the Agent Loop

If you build with coding agents, this is where Jev matters most. An agent asks itself small closed questions all day, and usually it asks them to a large model that answers in prose.

The most repeated one is "am I done?". Flavio Copes documents a judge that asks Jev whether a task is finished, plugged into an agent's loop, and @0x_rody's roundup lists Canny, a project that blocks an agent claiming it has finished. An agent asks that question many times per task, which makes it a narrow, repeated decision by construction.

The shell guardrail is the most instructive case. Jev put "irreversible" at 0.56 on that ambiguous rm -rf, which looks like a verdict if you read fast. The 0.33 confidence is what saves the files: it tells the code the answer is weak, so the command goes to a human instead of running. On a guardrail, the confidence matters more than the verdict it's attached to. Without it, a slightly loaded coin flip would stand between your agent and rm -rf.

Other builds use Jev to decide who does the work. A Reddit post catalogued by shipwithjev describes a router that picks, per request, between Qwen (an open model running locally) and Sonnet (Anthropic's hosted model). It's the arbitrage you make when you rebuild a $200/month agent setup for $15, except it happens on every request instead of once.

Browser agents push the idea further. Gregor Zunic, from Browser Use, posted a flight search run with Jev in 7 seconds for $0.0039, with a small LLM as fallback to type into the form fields. Hunch, another browser agent, reports a median of 153 ms per decision and 24 correct decisions out of 24 in its authors' test (a score that makes any QA lead ask to see the other test set). The Reddit side of the catalogue fills in the pattern with a deployment approval split into 18 questions (r/devops), a classifier for the emails agents are about to send (r/aiagents) and a context garbage collector called jev-gc (r/LLMDevs).

Railway switching system for automated task routing decisions
Railway switching system for automated task routing decisions

Across all of these builds the split holds: Jev judges, the code executes, and an LLM keeps the cases that are hard or that need writing.

Typed output fixes the form of the answer, not its substance. A choice question always returns an allowed option, and that option can be the wrong one. It still beats free-form output, where LLM calls silently wrote garbage to my database for days because a JSON field went missing and nothing threw an error.

The "It's Just BERT" Crowd Has a Point

The objection since launch is that Jev is just a classifier with good marketing. The people best placed to judge half agree, and don't see it as a knock.

Will Depue, formerly at OpenAI, accepts the word on X: yes, it's a classifier, but a zero-shot one (it handles your task without training on your data) with near-frontier intelligence, and he wonders which other old ML ideas deserve a comeback. Sebastian Raschka locates the breakthrough in the fact that Jev generalizes, and bets the secret lies more in the training data than in the algorithm. Merve, at Hugging Face, goes further: many problems people solved with LLMs could have been solved with zero-shot classifiers, which she files under "skill issue".

Then came the tests. A Japanese developer, @xjuntaro, ran Jev on a Kaggle task (translated and summarized from Japanese). Without any training, it landed just behind a fine-tuned BERT (an older, smaller model you train for a single task), level with the big LLMs and ahead of TF-IDF with logistic regression (a classic keyword-counting baseline). It dropped sharply with the default 0.5 threshold, though, so you calibrate the threshold on your own data before trusting it. Another Japanese developer, @hawkymisc, answered the "BERT can do it" debate by building a Jev-compatible API on top of a BERT-type model, which is the most developer way to win an argument on X: ship the counterexample.

TypeSafe's own list of limits points the same way: math, counting, dates and adversarial content. The last one is the catch for any guardrail, because the person who writes the input is sometimes the person trying to get past it.

The price drop has a darker side too. A calculation posted by @kenonews puts gross margins at 99% for solving captchas, with a market paying $0.01 per captcha and Jev solving 100 of them for $0.0068. It's worth reading as an economic signal only: every decision that just got cheaper for builders also got cheaper for whoever automates what builders are trying to block.

So a well-trained BERT gets close. You're paying Jev less for intelligence than for skipping the training run, and that edge only holds as long as the price does.

What Survives

The filter comes down to 3 criteria: a narrow decision you can write as options, a volume someone already pays to process, and an error you can catch. Triage, labeling, model routing and shell guardrails pass all 3. Doom has no buyer for its volume, and trading fails on the error you can't take back.

What's still unknown is specific. No third party has measured the gains, Metaview is still the only named production report in the catalogue, the sign-ups that opened on September 20 were paused on September 22, and TypeSafe itself says it can't prove that $0.042 per million tokens isn't a subsidized price. If that price moves, every cost comparison in the boring stack moves with it.

Sources

This post may contain affiliate links. If you click them, I might earn a small commission — costs you nothing, and helps me keep shipping quality articles every day for your reading pleasure.


Jev classifies instead of writing, and the boring use cases (triage, routing, research pipelines) are what actually ship in production. The Demo vs Product Checklist in the welcome kit lays out the 8 criteria that separate the 66 games getting attention from the handful actually earning.

→ Get the welcome kit