The race to make AI agents cheaper and faster is moving beyond the models that write their answers. In the two weeks since TypeSafe AI introduced Jev, a model that chooses among predefined options instead of generating prose, developers have released a rush of competing decision models, many with downloadable weights.
Now, Amazon is joining them with an open model designed to sit inside an agent's workflow and decide whether a proposed action should proceed: Strands Decider 2B, a roughly 2-billion-parameter model fine-tuned from an Alibaba Qwen base that answers bounded questions with choices and probabilities in a single pass.
The open source model is free to download from Hugging Face and use under an enterprise-friendly, permissive Apache 2.0 license, though developers must supply the hardware or cloud computing resources to run it.
Amazon says they can also download its training materials and use the model to route requests, select tools, evaluate outputs or review an agent's actions.
The release is part of Strands Labs, AWS's experimental agent-development project. AWS provided VentureBeat the final announcement and three technical charts ahead of its noon ET publication.
The central idea is simple: if a program needs an answer such as “Should this tool run?” or “Which of these three routes fits this request?”, asking a large language model to write an explanation can add time and expense. A decision model scores the allowed answers directly. That can make it useful as a frequent, narrow checkpoint around a more capable agent, while leaving writing, coding and complex reasoning to a generative model.
What AWS built
Strands Decider begins with the pretrained Qwen3.5-2B base model. According to AWS, its developers removed the component that predicts the next word and replaced it with a small “pointer” component that scores supplied answer options.
They adapted the base model with a rank-16 LoRA update and trained the new scoring component, which AWS says adds just over a million parameters. The architecture diagram accompanying the announcement labels this design “Hobson.” Unlike a chatbot, it makes one forward pass and returns a distribution over the available choices.
AWS says its team iterated rapidly and will publish version 20 alongside the code, training data and scripts. That openness is a meaningful part of the offer: developers could inspect the recipe, fine-tune the model, and run it on their own infrastructure rather than send each decision to an outside API.
The example in AWS's announcement places the decider immediately before an agent calls a weather tool. If a user asks for the weather without naming a location, the agent in the demo guesses a city. Strands Decider examines the conversation and proposed call, asking whether the city came from the user's request and whether it is too early to call the tool. Application code then uses its answers to send the agent back to ask for the city.
That example uses the Strands agent framework's existing intervention mechanism, which can proceed, deny a tool call, request confirmation or guide the agent with feedback.
AWS says the example's questions and thresholds were chosen by hand, and that dedicated decision-model integration libraries are still in development. The model supplies a judgment; the developer still defines what the application does with it.
How fast and accurate is it?
AWS says the model can make local decisions in tens of milliseconds on short tasks and under 100 milliseconds on some common hardware and inputs.

TypeSafe, meanwhile, reports end-to-end Jev responses of roughly 70 to 500 milliseconds through its hosted service.
Amazon's accompanying latency for Jev chart gives a more specific, qualified view: version 18 recorded a median of 106 milliseconds and a 95th-percentile latency of 296 milliseconds across 230 requests on a local Nvidia RTX 3090, including the HTTP round trip. Longer prompts generally took longer.
The chart excludes a 7.7-second first request during server warm-up; AWS also reports roughly 150 milliseconds median for small tasks on an M3 MacBook. These figures should not be presented as a uniform sub-100-millisecond guarantee.
For quality, AWS measures both accuracy and the trustworthiness of the probabilities using the public portion of JevBench. Its chart traces successive versions through version 19, which improved over the team's earlier checkpoints.

The plotted v19 point is roughly 72% accurate with a 0.35 Brier score, while the chart places the similarly sized Mapika decider-2b v11 around 76% and 0.32, respectively. Higher accuracy and a lower Brier score are preferable, so that comparison favors Mapika on both measures.
AWS says its model ranks second among public models around its size on this test, and first among those publishing a complete training recipe. The chart does not show a result for the planned version 20 release, and its plotted comparison does not include Jev itself. Neither the email nor those charts establish that Strands Decider beats Jev overall.
Network distance, request shape and operating costs differ. An Amazon spokesperson said on background that AWS is offering only the open source model for now, with no hosted API or per-call charge.
Developers can run it on a MacBook, the spokesperson said, but using their own machines or cloud infrastructure still consumes resources; Amazon did not provide a general per-token or per-request operating-cost estimate comparable to an API rate.
That leaves the relative cost of operating Strands Decider and paying for Jev's service unproven.
Jev sparked a fast-moving field
TypeSafe launched Jev on September 15, calling it a “System One” model. It takes application state and typed questions, then returns choices, scores or yes/no probabilities without generating a written answer. TypeSafe charges $0.042 per million input tokens, with no output-token fee, and says its post-training emphasizes probability calibration. Its model is accessed through an API; TypeSafe has not released Jev's weights or full training recipe.
The appeal and its limits have already drawn scrutiny. VentureBeat reported that the approach revives classification using stronger modern pretrained models: many enterprise systems need a label or routing decision, not a paragraph. Another VentureBeat investigation showed why a cheap probabilistic guard cannot serve as an infallible security boundary: adversarial text in an agent's input can influence its verdict. TypeSafe and integration partners have acknowledged that risk.
The category has expanded almost as quickly as the conversation about it. Laya uses a much smaller 421-million-parameter ModernBERT-based design for local decisions. Jared Palmer's Kev publishes a family of models and a research trail.
Bespoke Nimble 9B offers an Apache-licensed adapter and reference code. Mapika's decider family spans several sizes and releases training code and model cards.
FLock's this-that-model is another small, locally runnable entrant. They differ in architecture, training, hardware requirements and how well their confidence estimates hold up; sharing the “decision model” label does not make their benchmarks interchangeable.
Researchers at Stanford and Nvidia have taken a related route with open CLM-8B, which can cache representations of reusable actions. In the researchers' tests, VentureBeat reported, it ran up to nine times faster than Jev on some tasks, while trailing Jev on tool-calling accuracy. That result is specific to their workloads and testing setup, but it illustrates how quickly specialized alternatives are probing different speed-and-accuracy tradeoffs.
Category | Strands Decider 2B vs. Jev | What the evidence supports |
Up-front/API price | Strands advantage | AWS says the model is free to download and use, with no hosted API or per-call fee. Jev costs $0.042 per million input tokens . |
Total operating cost | Unclear | Strands requires the user to provide CPUs/GPUs or cloud compute. AWS gave no per-request or per-token operating estimate, so you cannot establish that it is actually cheaper to operate than Jev’s extremely inexpensive hosted API. |
Speed | Roughly competitive, no demonstrated win | AWS measured an earlier Strands checkpoint at 106 ms median / 296 ms p95 locally on an RTX 3090. TypeSafe reports ~70–500 ms end-to-end for hosted Jev. Different hardware and network conditions make a direct winner impossible to establish. |
Accuracy | No evidence Strands beats Jev | AWS’s chart puts Strands v19 at roughly 72% accuracy on public JevBench. It does not include Jev itself, so there is no head-to-head result supporting an accuracy advantage. |
Calibration | No demonstrated Jev win or Strands win | AWS reports roughly a 0.35 Brier score for v19. But again, Jev itself is absent from the plotted comparison. TypeSafe specifically emphasizes calibration in Jev’s post-training, but the materials here don't establish which is better. |
Openness | Clear Strands advantage | AWS plans to release the model, code, training data and scripts; Jev remains API-only and TypeSafe has not released its weights or full training recipe. |
Self-hosting / privacy / control | Clear Strands advantage | Strands can run locally, including on a MacBook, so companies can keep decisions inside their own environment rather than sending them to an external API. |
Customization | Likely Strands advantage | Because AWS is publishing the training recipe and materials, developers can inspect and fine-tune it. Jev does not offer the same degree of model-level control. |
Ease of consumption | Jev advantage | Jev is a hosted API: call it and pay a tiny usage charge. Strands makes you deploy and maintain inference yourself. That is more flexible, but also more operational work. |
AWS/Strands integration | Strands advantage for that ecosystem | AWS shows it working directly with the Strands intervention mechanism as a checkpoint before agent actions. |
On price, there's a nice paradox: Strands has a sticker price of zero, while Jev has an operating price that is already almost negligible. Jev's $0.042 per million input tokens means one billion input tokens costs only $42 at the quoted rate. So unless a company already has spare compute or values local execution for privacy/control reasons, AWS hasn't yet demonstrated that self-hosting Strands saves money once hardware, cloud compute and operational overhead enter the equation.
The place where Amazon unquestionably pushes the category forward is open-source reproducibility. TypeSafe asks developers to trust and consume Jev as a service; Amazon is effectively saying: here is the small model, the recipe, the data and the machinery — run it yourself and modify it. That may be a bigger strategic differentiator for enterprises than squeezing another few percentage points out of a benchmark.
AWS's distinguishing proposition is therefore less a demonstrated Jev-killer than an open, reproducible decision layer tied to an existing agent framework.
A Strands developer can put a small local model before a tool call, decide when uncertainty warrants asking a person or handing control back to the agent, and retain an LLM for tasks that require generation. Its published recipe and ability to run on a laptop could make experimentation easier for teams that want control of their infrastructure.
The decisive test will be outside the launch examples: whether the released checkpoint stays accurate and well calibrated on enterprises' own policies, documents and unusual cases, and whether local inference plus maintenance beats Jev's unusually low hosted price. For sensitive tool calls, the weather demo is a useful illustration of a review point, not proof that a model judgment alone makes an agent safe.
In short: Amazon has produced a genuinely competitive Jev alternative whose strongest differentiation is not superior benchmark performance, but that it is open, self-hostable and reproducible.
