- Monday Momentum
- Posts
- We Have Been Using Writing Machines to Make Decisions
We Have Been Using Writing Machines to Make Decisions
How TypeSafe's Jev skips text generation entirely to make agent decisions in milliseconds, and why a model that cannot chat may be the missing layer in every AI workflow
Happy Monday!

TypeSafe describes Jev as 'a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.' (Source: TypeSafe AI)
Diogo Almeida helped build the instruction-following research at OpenAI that became the foundation for ChatGPT. Then he left, spent two years in stealth, and returned on September 15 with a question that should bother anyone building agents: "Models have been superhuman at chat for years, so where is all the automation?"
His answer is a model that cannot chat at all. TypeSafe AI's Jev does not write sentences, code, or explanations. It takes your application's state and a set of questions you define in advance, then returns typed answers with calibrated probabilities in 70 to 500 milliseconds. TypeSafe prices it at $0.042 per million input tokens, and output is effectively free because there is almost nothing to meter.
The idea underneath it is one of the more useful framings of the year. Most of what an AI agent does is not writing, it is deciding. Should this ticket go to billing or support? Is this output safe to show? Which of these 40 links gets closer to the goal? We have been answering those questions by asking a text generator to write out its answer, then parsing the prose and hoping it did not invent a tool name. Jev's bet is that decisions deserve their own kind of model.
TypeSafe AI, founded by a former OpenAI researcher behind ChatGPT's instruction-following work, released Jev, a "System One" decision model that returns typed answers with calibrated probabilities instead of generated text. It runs in 70 to 500 milliseconds, costs $0.042 per million input tokens, and cannot return an off-schema answer by construction. It also cannot write, summarize, or count. Early builders are pairing it with frontier LLMs, letting the LLM write and plan while Jev handles the fast, repetitive decisions in between.
How Jev Actually Works
Start with what a normal LLM does when you ask it to classify something: it generates its answer one token at a time, each conditioned on the last. To return a small JSON object with a category and an urgency flag, it writes every character in sequence. Your code then parses and validates that output, and generation can still fail, stop early, or drift outside your schema.
Jev never generates a string. You declare the possible answers up front using three primitives: a yes-or-no question, a choice among defined options (up to 255), or a score on a scale. You pass in the current state of your application, such as a support ticket, a web page, or a game screen. Jev evaluates every question against that state in a single parallel forward pass and returns a probability distribution over the options you allowed.
Two consequences naturally follow from this structure. First, a type error is impossible, because an answer outside your schema is not among the possible outputs. TypeSafe is careful to say this guarantee covers the shape of an answer, not whether it is correct. Second, every answer carries a probability, and Jev was trained with a method TypeSafe calls Reinforcement Learning for Calibrated Decisions. The goal is that answers given 90% confidence turn out right roughly 90% of the time.
TypeSafe's launch post states the argument well: if a model can do a task 95% of the time but cannot tell you when it is in the other 5%, you cannot automate that task. A model that knows how confident it is lets your code decide when to act and when to escalate to a human or a bigger model.
Where Each Kind of Model Belongs
Jev gives up a great deal on purpose. It cannot summarize a document, draft a reply, or explain its reasoning. An independent deep dive by developer Flavio Copes found it is not a calculator either: it does not count reliably, and its scores work for thresholds and rankings rather than precise measurement.
Generative Models vs. Decision Models
Frontier LLM | Decision model (Jev) | |
|---|---|---|
Output | Generated text, parsed afterward | Typed values defined in advance |
Speed | 3 to 329 seconds end to end | 70 to 500 milliseconds |
Confidence | Stated, often overconfident | Calibrated probability on every answer |
Failure mode | Hallucination, malformed output | Wrong answer inside a valid schema |
Best at | Writing, planning, exceptions | Routing, scoring, classifying, verifying |
The interesting architecture uses both. One open-source browser agent, Hunch, has Jev choose the next action on a page, ordinary code check the result, and uncertain or irreversible steps go to a frontier LLM. Its author reports a 153-millisecond median per decision. Another pattern has an LLM write a summary while Jev checks each claim against the source material and flags low-probability claims for review. The LLM writes, Jev decides, and code ultimately verifies.
Parallel evaluation also changes workflow design. With LLMs, developers ask one question and then decide what to ask next, because every call is slow and costly. With Jev, you can ask every question you might need in a single call and let your code choose which answers to use. TypeSafe calls this speculative fan-out. In one cookbook, 13 questions in one call ran 12.2 times cheaper and 10 times faster than 13 sequential calls.
The Receipts, and What They Do Not Prove
TypeSafe claims Jev is 193.6 times faster and 444.6 times cheaper than frontier LLMs on its own workflow evaluations. The company deserves credit for how openly it qualifies that number, so the qualifications are worth quoting directly.
It calls those multiples "on the higher end of real world gains." Its evaluations use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which it acknowledges biases results toward those models. The workflows were built by its own team, and it concedes "some bias could exist." On pricing it writes plainly: "We can't prove it isn't subsidized." Its latency numbers are measured from the West Coast, where the service runs.
Early adoption is real but self-reported. An independent catalog, shipwithjev.com, listed 551 Jev builds as of September 23. One builder reported classifying 1,018 research papers across 24 topics for eight cents at a 256-millisecond median. Another ran 98,000 listing classifications in ten minutes. Those are builders' own figures rather than audited benchmarks, and a weekend demo is not a production system. The fair summary is that the architecture is plainly useful, and the size of the advantage will depend on your task.
Why the Name Matters
TypeSafe named its model class after Daniel Kahneman's distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning. Frontier LLMs with extended reasoning are System 2 machines. Jev is built for the fast, reflexive judgments that make up most of the work.
It named the model itself after William Stanley Jevons, the economist who observed that more efficient steam engines increased coal consumption rather than reducing it. TypeSafe's bet is that each order-of-magnitude drop in the cost of a decision unlocks orders of magnitude more places to use one. A person opens ChatGPT a few times a day while software could make thousands of small decisions in the background, every second, if each one costs a fraction of a cent and returns before a user notices.
This newsletter has tracked the price of intelligence falling for months, with DeepSeek, OpenAI's Luna cuts, and this week's flagship price war. Jev suggests the next drop may note come from a cheaper chat but from a different kind of model built for a different job.
What This Means for Practitioners
For developers building agents, audit your pipelines for steps that are really decisions wearing a text costume: routing, classification, guardrails, relevance scoring, pass or fail checks. Those are candidates for a decision model, and moving them can cut latency from seconds to milliseconds without touching the parts that actually need generation.
For product teams, speed changes what is possible. At 100 milliseconds, AI judgment can sit inside a user interface, a voice agent deciding whether to interrupt, or a search box re-ranking results as you type; at five seconds, it cannot.
For enterprise buyers, calibrated confidence is the feature to test first. Run any decision model against your own labeled data and check whether its 90% answers are actually right 90% of the time. If they are, you have a principled way to decide what to automate and what to route to a person. If they are not, the speed does not matter.
The Bottom Line
For three years the industry has treated the chat model as the universal interface to intelligence, and then bolted parsers, validators, and retry loops onto it to make it behave like software; Jev inverts that. It gives up the ability to write in exchange for answers software can use directly, delivered fast enough and cheaply enough to be called thousands of times.
Whether TypeSafe specifically wins is an open question, and its own numbers deserve independent testing. The larger idea looks durable: agents need something that writes and something that decides, and they probably should not be the same model.
In motion,
Justin Wright
If most of the work inside an AI agent is making small decisions rather than writing, how much of today's agent cost, latency, and unreliability comes from using the wrong kind of model for the job, and will the frontier labs build decision models of their own once they notice?

Introducing System One Models & Jev - TypeSafe AI
A deep dive into Jev, TypeSafe's System One model - Flavio Copes
Jev Explained: Typesafe AI's Non-Autoregressive System-1 Model - MindStudio
What Is Jev? A Guide to TypeSafe AI's System One Model - LangChain
TypeSafe Workflow Evals - TypeSafe AI
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - Simon Willison
The Closed Quorum: Inside the first reported autonomous AI C2 implant - Cisco Talos
Trump calls AI oversight a 'globalist scheme' as Amodei and Altman head to the UN - Fortune
Builder’s Note
I have been playing around with Jev for a number of use cases, and have been using it within my own quantitative investment model to benchmark decisions against the trained ML model that makes investment decisions now. While my initial tests were high level, routing decisions to Jev during down-market periods led to some years with better returns and smaller individual losses in backtesting. I will continue playing around with it and see if I can find more ways to integrate it into existing workflows for direct comparison.
Quick Hits
Anthropic released Claude Opus 5.5 at $4/$20 per million tokens, down from $5/$25. About 90 minutes later, OpenAI launched GPT-6 Sol at exactly half that, $2/$10, and GPT-6 Luna at $0.10/$0.50. (Simon Willison)
Cisco Talos disclosed CLOSEDQUORUM, the first reported autonomous AI malware, and released CAIRN, an open-source toolkit to help defenders track AI-integrated malware. Talos says it has no confirmation of in-the-wild deployment. (Cisco Talos)
Sam Altman briefed the UN Security Council on AI and international security on September 23, with Anthropic, DeepSeek, and Moonshot also participating, the same day President Trump called international AI oversight a "globalist scheme." (Fortune)

If you haven’t listened to my podcast Mostly Humans: An AI and business podcast for everyone yet, new episodes drop every week!
Episodes can be found below - please like, subscribe, and comment!
