A large share of the AI running inside enterprise software today is not writing anything. It's not chatting with an AI chatbot, and it's not even reasoning about how to give you the correct answer. It is deciding: which queue a support ticket belongs in, whether a RAG document is relevant, whether something looks weird with a transaction, or which model should handle a request. In many enterprise AI stacks, those decisions are being made by asking a language model to generate text, then parsing that text back into a label.

TypeSafe AI's Jev, released in early access on September 15, takes a different approach to those decisions. Jev doesn't chat, doesn't reason through a long chain of thought, and it doesn't even generate text. It takes application state and a set of typed questions and returns choices, scores and yes/no answers with probabilities attached, typically in 70 to 500 milliseconds, and at $0.042 per million input tokens, according to TypeSafe. And output tokens? Well, actually, they're free.

Within days, Jev had integrations in everything from observability platforms to agent frameworks like Pydantic AI and LangChain. TypeSafe had to pause their signup queue, and many people started using the model via OpenRouter. Meanwhile, a wave of open-source alternatives began appearing on GitHub and Hugging Face. While the attention is on Jev, the underlying approach is older: classification.

Why generate text when all you need is a label?

Chat models are optimized for a different job. Pretraining teaches them to predict the next token, and post-training with reinforcement learning from human feedback teaches them to produce responses people prefer. Both stages produce strings. Automation needs typed values, predictable latency, stable cost per call, and outputs that can be easily scored against ground truth.

Using a generative model for a decision (even when accurate) means paying for token-by-token decoding, then parsing and validating the output, then handling the cases where the model drifts from the format, adds commentary or refuses. Generative models also tend to be poorly calibrated when asked how confident they are. That matters more than it sounds: if a system can't tell when it's likely to be wrong, it can't reliably decide when to hand a case to a human. TypeSafe's launch post makes this argument directly, contrasting sequential generation with Jev's parallel sampler and self-reported confidence with calibrated probabilities.

While small language models have become fast and cheap, they still generate tokens to produce a decision. A classifier can return the label or score the calling software needs directly, without generating text.

What changed: Pre-training finally caught up with classification

So if classification was always the better fit for decisions, why is it suddenly returning now? The main answer is the quality of pre-training.

By 2019, researchers had already shown that zero-shot classification was workable. Yin et al. repurposed natural language inference models to test candidate labels against an input, and facebook/bart-large-mnli became a common default for classifying text against labels supplied at request time. But those models were brittle outside their training domains, so many deployments still needed labeled data and task-specific training.

BERT was trained on roughly 3.3 billion words. ModernBERT, released in December 2024, was trained on 2 trillion tokens with an 8,192-token context window. Stronger pretrained models can now handle a wider range of classification questions defined at request time, reducing the need to train a separate model for every task.

The open reproductions of Jev make this visible, because their internals are public. SemIf trains no model at all. It reads option probabilities directly from a frozen Qwen3.5-4B, and SemIf's developers measured about one second for 21 questions, against more than five seconds for generated JSON. The two methods agreed on 18 of 21 choices. Laya builds a decision model on a 421M-parameter ModernBERT-large backbone and answers in roughly 33 to 40 milliseconds per question on an Nvidia T4. Other projects apply LoRA adapters to Qwen3.5 or run on DiffusionGemma. The implementations differ, but they share the same underlying idea: using stronger pretrained models to make decisions without generating text.

The reproductions also show the limits. Laya's developers report that its base checkpoints score near chance on their typed-decisions benchmark, with substantially better results after fine-tuning. TypeSafe itself has disclosed little about Jev's internals beyond a new architecture, a parallel sampler and a post-training method it calls Reinforcement Learning for Calibrated Decisions (RLCD), which optimizes the model to return probabilities that match how often it is right. Making calibrated decisions its primary training objective is TypeSafe's big bet.

A new era: Modern pre-training, old deep learning tactics

Jev is the most visible example of a broader pattern: builders taking frontier-quality pretraining and applying the pre-ChatGPT playbook on top of it.

A cheap classifier decides whether a request goes to a small local model, a specialist, or a frontier API, and research such as FrugalGPT and RouteLLM has already formalized cascades for language models. Pydantic AI's Jev integration is built around the same split: Jev answers the typed fields, and anything that requires written text escalates to a language model.

Distillation and small specialists never really left the companies with large, long-standing ML teams, and VentureBeat's Beyond the Pilot podcast has documented how they use it. At LinkedIn, Erran Berger, who led the work as VP of product engineering (now as CTO), described building next-generation recommendation systems by fine-tuning a 7-billion-parameter teacher on a detailed product policy document, then distilling it through a 1.7B intermediate model down to a 0.6B student that serves production traffic. "There was just no way we were gonna be able to do that through prompting," Berger said. The process is now a repeatable recipe reused across LinkedIn's AI products. Shopify has taken the same idea further with an internal platform that lets R&D teams distill a frontier model into a fine-tuned open model for a single subtask in roughly a day, with evaluations built in. Farhan Thawar, Shopify's VP and head of engineering, said the distilled models run anywhere from 2x to 30x cheaper and faster, and in some narrow tasks outperform the frontier model they learned from.

The common thread is that generation gets reserved for the tasks that actually need it: drafting, summarizing, writing code. Narrower decisions, such as routing and tagging, can go to less expensive models that meet the application's accuracy and reliability requirements.

What enterprise teams should do before moving decisions off LLMs

The first step is an audit of existing LLM traffic, sorting calls into generation and decisions such as routing, triage, tagging, approval and scoring. A worked example by developer Flavio Copes illustrates the math: a company spending $10,000 a month on AI, with $6,000 of it on decisions, would save $5,700 a month if those decisions moved to a path costing 5% as much. Actual savings depend on how much of that spending can move and what the replacement actually costs.

The caveats are real. Jev is in early access, accepts text only, and currently serves from the West Coast. TypeSafe says it will take time to establish whether Jev's pricing is sustainable, and says its headline figures of roughly 194x faster and 445x cheaper come from its own workflow evaluations and likely sit at the high end of real-world gains. Pydantic's documentation also notes that text in the input written to steer Jev's answer can succeed, so Jev-based guardrails should sit alongside deterministic checks rather than replace them. The open alternatives remove the vendor dependency but hand teams the work of hosting, fine-tuning, calibration and monitoring.

The longer-term implication is about people as much as models. Engineers who built classifiers, cascades and evaluation pipelines before 2022 hold skills that have become relevant again, and teams that pair that experience with current pretrained models are positioned to build automation that is cheaper, faster and easier to audit.