Amazon Web Services has open-sourced a decision model called Strands Decider 2B, the cloud giant's entry into one of the fastest-spreading niches in applied AI: small, fast models that don't write text at all, but simply pick the next move from a menu of options and say how confident they are.
The model came out of an internal experiment by Marc Brooker, a distinguished engineer at AWS, who started building his own version after studying TypeSafe's Jev — the model that kicked off the current "decision model" wave in September. The homebrew project was good enough to briefly top the Jevbench ranking for models of its size, and AWS engineers subsequently cleaned it up and shipped it under Strands Labs, the group that builds tools and protocols for deploying AI agents.
Architecturally, Strands Decider 2B keeps the "torso" of the Qwen3.5-2B language model for language understanding, then strips out the text-generation head and replaces it with a pointer head of roughly one million parameters that scores the developer's predefined options directly. The whole thing was fine-tuned with LoRA adapters rather than a full retrain, and the model, its training data, and its training scripts are all published — on Hugging Face and GitHub — small enough to run on ordinary local hardware.
Because the answer space is closed, the model can't hallucinate an option the developer never offered it. It instead returns a calibrated choice with a confidence score, which makes it useful for the plumbing inside agent workflows: routing requests, selecting tools, evaluating outputs, and sanity-checking an agent's plan before an action runs. In AWS's own demo, built on the open-source Strands agent framework, a Decider check catches a weather agent that guessed a city the user never mentioned and redirects it to ask a clarifying question first.
On JevBench, the emerging benchmark for this model class, AWS reports a perfect score on the easy tier, second place among public models of roughly 2 billion parameters overall, and first place among models that ship with a comprehensive training recipe. AWS also says the model decides in under 100 milliseconds on common hardware including an NVIDIA RTX 3090 — figures that are company-reported and not yet independently verified.
"What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — 'what is the next thing for me to do here, based on where I am?'" Brooker told TechCrunch. He noted the design challenge is pushing decision speed and calibration without degrading the multilingual understanding and general knowledge that make the base model broadly useful — and he doesn't expect frontier labs to own the category, since in smaller niches an interesting model costs hundreds or thousands of dollars to build, not billions.
The release lands the same week OpenAI opened a limited preview of its Decisions API, a hosted equivalent built on the Luna model — and barely a month after TypeSafe debuted Jev, named for the economist whose paradox holds that cheaper intelligence can increase demand. Dozens of lookalike models have appeared since. TypeSafe CEO Diogo Almeida told TechCrunch he isn't worried yet: "The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful."
The bigger picture is that agent stacks are splitting into "system one" and "system two": big, expensive frontier models for open-ended reasoning, and tiny calibrated classifiers watching every step in between. If most agent actions are routine routing decisions, most agent inference spend may soon belong to models like this one — a 2-billion-parameter component that costs almost nothing to run and never says a word.
Comments (0)
Log in to join the discussion
Log InNo comments yet