AWS Releases Strands Decider 2B

On October 1, 2026, AWS's Strands Labs released Strands Decider 2B, a small open decision model for agents. It takes a state and typed questions and returns a choice, a yes/no probability or a score, each with a confidence value. The code and training data are on GitHub and the weights are on Hugging Face under Apache 2.0. The post names Jev as the model that started this class.
Why it matters
An agent's routine choices, such as picking a tool or rating an output, are classification problems. At 2B parameters Strands Decider is small enough to run on a local GPU or CPU, so inference needs no per-call fee and no remote API. The outputs are bounded to the options you supply, and each answer comes with a confidence value you can log and review.
Because the weights are open and AWS published its training data and scripts, a team can fine-tune the model on its own labels. That's the work we do with hosted and open decision models, and the reason to measure on your own data first: AWS reports strong results on a public benchmark, and a client's task rarely looks like one. See fine-tuning Jev for the method and the decision models page for our measurements.
Strands Decider has no text-generation head. A small readout scores the options you supply, so the model can't write prose.
Key technical notes
- Architecture: Qwen3.5-2B with the language-model head removed and replaced by a pointer head of just over a million parameters that scores each option. The base model is fine-tuned with a rank-16 LoRA adapter. The released model is version 19.
- Median latency is about 115 ms on an Nvidia RTX 3090 and about 153 ms on an M3 MacBook, according to AWS.
- AWS reports that it measured accuracy on JevBench's public set and calibration with the Brier score on the same set, and ranks third of 33 models in the 2B class, or first of 30 if models just over 2B are excluded. The Hugging Face card self-reports 167 of 231 correct (72.3%) on JevBench public, with a Brier score of 0.348 and an ECE of 0.050.
- AWS says the model answers all of the easy JevBench tasks correctly. These are AWS's own measurements. AI Weekly reports that Mapika decider-2b v11 scores higher on the same benchmark, 76% against 72%.
- The blog post is by Marc Brooker, Mike Chambers and Fabio Nonato de Paula.
Sources
- Introducing Strands Decider 2B: a small, open source, decision model, Strands Agents blog, October 1, 2026
- strands-labs/strands-decider, GitHub
- StrandsAgents/strands-decider-2B-hobson-v19, Hugging Face
- Amazon Ships Strands Decider 2B, an Open-Source Jev Rival, AI Weekly
- Cover image: Strands Agents, architecture figure from Introducing Strands Decider 2B, padded to 1200x630
- TypeSafe AI's Jev now available on AI Gateway, Vercel changelog, September 16, 2026