Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

AWS Open-Sources Strands Decider 2B, a 2B-Parameter Model That Answers With Choices Instead of Text — and Decides in Under 100ms on an RTX 3090

AWS Open-Sources Strands Decider 2B, a 2B-Parameter Model That Answers With Choices Instead of Text — and Decides in Under 100ms on an RTX 3090

AWS's Strands Labs released Strands Decider 2B, an open-source decision model built on a Qwen3.5-2B torso with a roughly 1M-parameter pointer head that picks from developer-defined options and returns a confidence score. AWS says it decides in under 100 milliseconds on an RTX 3090 (company-reported). It lands the same week OpenAI previewed its own Decisions API.

Amazon Web Services has open-sourced a decision model called Strands Decider 2B, the cloud giant's entry into one of the fastest-spreading niches in applied AI: small, fast models that don't write text at all, but simply pick the next move from a menu of options and say how confident they are.

The model came out of an internal experiment by Marc Brooker, a distinguished engineer at AWS, who started building his own version after studying TypeSafe's Jev — the model that kicked off the current "decision model" wave in September. The homebrew project was good enough to briefly top the Jevbench ranking for models of its size, and AWS engineers subsequently cleaned it up and shipped it under Strands Labs, the group that builds tools and protocols for deploying AI agents.

Architecturally, Strands Decider 2B keeps the "torso" of the Qwen3.5-2B language model for language understanding, then strips out the text-generation head and replaces it with a pointer head of roughly one million parameters that scores the developer's predefined options directly. The whole thing was fine-tuned with LoRA adapters rather than a full retrain, and the model, its training data, and its training scripts are all published — on Hugging Face and GitHub — small enough to run on ordinary local hardware.

Because the answer space is closed, the model can't hallucinate an option the developer never offered it. It instead returns a calibrated choice with a confidence score, which makes it useful for the plumbing inside agent workflows: routing requests, selecting tools, evaluating outputs, and sanity-checking an agent's plan before an action runs. In AWS's own demo, built on the open-source Strands agent framework, a Decider check catches a weather agent that guessed a city the user never mentioned and redirects it to ask a clarifying question first.

On JevBench, the emerging benchmark for this model class, AWS reports a perfect score on the easy tier, second place among public models of roughly 2 billion parameters overall, and first place among models that ship with a comprehensive training recipe. AWS also says the model decides in under 100 milliseconds on common hardware including an NVIDIA RTX 3090 — figures that are company-reported and not yet independently verified.

"What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — 'what is the next thing for me to do here, based on where I am?'" Brooker told TechCrunch. He noted the design challenge is pushing decision speed and calibration without degrading the multilingual understanding and general knowledge that make the base model broadly useful — and he doesn't expect frontier labs to own the category, since in smaller niches an interesting model costs hundreds or thousands of dollars to build, not billions.

The release lands the same week OpenAI opened a limited preview of its Decisions API, a hosted equivalent built on the Luna model — and barely a month after TypeSafe debuted Jev, named for the economist whose paradox holds that cheaper intelligence can increase demand. Dozens of lookalike models have appeared since. TypeSafe CEO Diogo Almeida told TechCrunch he isn't worried yet: "The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful."

The bigger picture is that agent stacks are splitting into "system one" and "system two": big, expensive frontier models for open-ended reasoning, and tiny calibrated classifiers watching every step in between. If most agent actions are routine routing decisions, most agent inference spend may soon belong to models like this one — a 2-billion-parameter component that costs almost nothing to run and never says a word.

Comments (0)

Log in to join the discussion

Log In

No comments yet