Amazon's Strands Labs open-sources Strands Decider 2B, a local decision model for agents
Amazon's Strands Labs on Thursday released Strands Decider 2B, an open-source 2-billion-parameter decision model that chooses among supplied options and returns calibrated confidence instead of free-form text, small enough to run locally for agent workflow steps.
Most agent steps are not essays — they are yes/no, route-here, or tool-now decisions. A 2B open model that answers those in tens of milliseconds with a confidence score is the kind of cheap control plane that makes agent stacks affordable without calling Opus for every branch.
On Thursday, 1 October 2026, Strands Labs released Strands Decider 2B. The Strands Agents post is titled “Introducing Strands Decider 2B: a small, open source, decision model.” The page dates the post October 1, 2026, and says 8 min read. It does not print an hour. The byline is Marc Brooker, Mike Chambers, and Fabio Nonato de Paula. The post says the group announced Strands Labs earlier this year, as a place to try new approaches to agentic AI, and that Decider is the addition today. Agentic, here, means software that takes a next step, not only a chatbot that answers. The weights are on Hugging Face. The training data and the scripts are in the open repository strands-labs/strands-decider on GitHub. The model card names the checkpoint StrandsAgents/strands-decider-2B-hobson-v19. GitHub’s license record for that repository, and the model card, say Apache 2.0. Apache 2.0 is a license that lets someone use, change, and ship the work, including in a product, if the license and copyright notices stay with it. Those lines are Strands’.
A decision model, as the post describes it, does not write free-form text the way a large language model does. It picks among a closed set of options the caller supplies, or it scores an ordered scale the caller supplies, and it returns a reliability score: how sure it is that the answer is right. The post’s examples are a yes-or-no question about whether the sentence “turn on the lights” is about a coffee machine, a choice of English, Zulu, or Dutch for the phrase “sihamba ngokushesha,” and a 0-to-1 score for whether “this is the best doc I’ve ever read” is positive. The repository names the three question types noul, a yes or no, choice, one option from a list, and score, a level on a scale. In exchange for that narrower job, the post says, these models are faster at a given size, always answer from the options they were given, and can run at very low latency. Latency means how long the answer takes. The same post says the method, which produces every score in one pass, is significantly worse at complex problems than a reasoning model, and that the lack of text makes it a poor fit for coding, chatbots, and summarizing a document. Those limits are the post’s. A decision model is not a replacement for a frontier model on those jobs.
The architecture, as the post states it, starts from a pretrained language-model torso, Qwen3.5-2B, and removes the language-model head, the part that would have written the next word. A torso, in that sentence, is the body of the model that has already learned language. Two billion is the parameter count of that torso. A parameter is one of the numbers that shapes the model. In place of the old head is a pointer head that scores each offered option. The post says the head compares the model’s internal summary at each option with the internal summary at the end of the question. The head is small, just over a million parameters. The torso is then fine-tuned with a rank-16 LoRA adapter. Fine-tuned means trained a bit further. LoRA is a small set of extra weights, here of rank 16, so the team does not retrain every weight in the two billion. The model card names the base as Qwen/Qwen3.5-2B-Base and says the published files are that adapter plus the readout head. The post says the model is small enough to run on a local CPU or a GPU. A CPU is the computer’s main processor. A GPU is the graphics chip people also use for this math. Those lines are Strands’.
The post says this is the second major version of the architecture. The first used a slot head, and the team found it performed significantly worse. The release is v19, after many iterations under the covers. The repository, the post says, records what changed in each version. The results note in that repository says v19 adds answer-adequacy rows, a request, a response, and whether the response actually answers it, on top of the earlier recipe. Those lines are Strands’.
Strands says it cares about three things: accuracy, how often the answer is right, calibration, whether the confidence matches the hit rate, and latency. It measures the first two together, accuracy on JevBench’s public set and calibration with the Brier score on that same set. JevBench is the public test the post names for this class. A lower Brier score is a better match between confidence and outcome. The post says strands-decider-2b is 3rd of 33 in the 2 billion class on that accuracy and calibration, and 1st of 30 if models that sit just over 2 billion parameters are left out. It also says the model gets every easy task on JevBench right, and that those easy tasks look like some of the simpler steps people give agents. The model card prints 167 of 231 tasks right on the public set, an accuracy of 0.723, with a Brier score of 0.348 and an ECE of 0.050 at a window of 4,096 tokens. ECE is expected calibration error, another check on whether the confidence matches the hit rate. A token is a small piece of the input. The repository’s results note records the same 167 of 231, and on its reference table a Brier score of 0.342 and an ECE of 0.052. The two pages do not print the same Brier figure. Both figures are Strands’. The results note also says v19 is not on the public JevBench board. The 3rd-of-33 place is where Strands puts that score on the board snapshot from 25 September 2026: 50th of 90 systems on that full snapshot, first of 30 at 2 billion parameters or fewer, and third of 33 once three models just above 2 billion are counted. The same note lists a standard tier of 72 tasks at 0.875 and a hard tier of 111 tasks at 0.505, the same hard-tier score as v18. It also records a later v20 run at 169 of 231. The blog’s released checkpoint is v19. None of these scores is an outside lab’s rerun.
On short classification tasks the model had not trained on, the repository says a confidence of 0.9 or higher was right about 95 percent of the time. The results table for v19 puts that top band at 0.952, on 23 percent of those answers. The same note says that band is measured on short classification, and that this is the only place it holds. On long documents the model is under-confident: right about 0.83 of the time at a stated confidence around 0.64. Under-confident, here, means it is righter than its score admits, so a high bar sends more cases to a person than the hit rate alone would require. The model card’s own limits say questions are read less carefully than documents, that the hard tier of JevBench scores far below the easy tier, and that the confidence bands were fitted on short classification. Those lines are Strands’. A threshold copied from the short-task table is not a guarantee on a long document.
On speed, the post says strands-decider-2b can make a local decision in a median of about 115 milliseconds on widely available hardware. A millisecond is a thousandth of a second. The median is the middle result. The post says the time grows roughly in a straight line as the task gets longer. The graph on the page, the post’s Figure 3, is measured on a local Nvidia RTX 3090 against v18 of the model. On an M3 MacBook, the post says, the median for small tasks is about 153 milliseconds. The repository’s results note, for v19, says a median of 115 milliseconds on an RTX 3090, with a slow case, the 95th percentile, at 299 milliseconds, and a warm M3 Pro median of 234 milliseconds. The 115 millisecond line is the same middle number. The Mac figures are two different measurements in Strands’ own write-up, small tasks on the blog and the benchmark timing in the results note. Those times are Strands’. The release does not print a price. It is the open weights and a local tool.
The post says early uses of this class include model routing, tool selection, evaluations, guardrails, memory, context management, and policy classification. Routing means picking which model should take the next step. A guardrail is a check that can block a step. Context is the text the agent is holding. The post also describes hybrid agents: a large language model for the hardest decisions, and a decision model for the routine ones, to cut cost and wait. It says people are also trying fixed workflow languages, games, task automation, and mazes. Those are the post’s examples of experiments. They are not a customer list, and they are not a claim that the small model takes over the jobs the post already says it cannot do.
The easiest start, the post says, is the strands-decider command-line tool. A command line is the text window where you type. One example on the page is a support note, “Help! My payouts have been failing for 3 days!”, and a choice of billing, sales, or retail. The printed output picks billing, with a confidence of 0.768, and probabilities of 0.845 for billing, 0.091 for retail, and 0.064 for sales. The model card prints a nearby run, confidence 0.769 and probabilities 0.846, 0.090, and 0.064, and says a retrain can move the numbers while keeping the same answer. The same example can also ask whether the note is urgent, a yes-or-no, and how frustrated the writer is, on a scale of calm, frustrated, or depressed. Asking several questions about one note is the efficiency the post describes, because the model reads the note once.
The repository includes an example that puts the model inside a Strands agent, under examples/strands/. The agent runs locally, talks to Decider locally, and uses the default language model from Amazon Bedrock for the part that still writes text. Bedrock is Amazon’s service for calling models. The scenario is small on purpose. The agent has a weather tool and instructions that make it eager, so if someone asks “What’s the weather?” without a city, it guesses a city and tries to call the tool. Before that call, Decider reads the conversation and the proposed tool call and answers two yes-or-no questions: whether every argument traces back to something the user said, and whether it is too early to call the tool. In the example the city was guessed, and the agent goes back and asks which city, instead of reporting weather for a place nobody named. The post says this uses Strands’ intervention system. A handler runs before any tool executes and can proceed, deny the call, stop and ask a person, or hand the model feedback and another turn. The post says the questions, the threshold, and the policy in the example were picked by hand. It is an illustration, not a recommendation. The team says it is working on libraries for this kind of integration.
Tim Fernholz’s TechCrunch story is stamped 9:49 a.m. PDT on October 1, 2026, which is 12:49 p.m. Eastern. The headline calls the release Amazon’s own Jev clone, as decision models flood the web. The piece says Amazon Web Services released an open-source decision model inspired by TypeSafe’s Jev, small enough to run locally, that sorts between options decided in advance and returns a confidence. It says Marc Brooker, an Amazon distinguished engineer, started the project after seeing Jev, that the home-built version briefly reached the top of the JevBench ranking for models of its size, and that Amazon engineers then cleaned it up for Strands Labs. That “briefly reached the top” line is TechCrunch’s account of the earlier experiment. It is not the rank Strands prints for the v19 release. Brooker told TechCrunch the interest came from AWS customers whose agent steps did not always need a full language model. Quoted, as the piece prints him: the class is “a perfect decider for a workflow step,” the question of what to do next given where the agent already is, and it can be “more reliable, thanks to the confidence scores, thanks to the closed domain of answers,” with lower latency and, potentially, lower cost. The piece says TypeSafe named Jev for the economist William Stanley Jevons, and the hope that a cheaper kind of computer intelligence would be used more, not less. It says the release came the same week OpenAI announced a similar offering. It does not name that OpenAI model. Brooker told TechCrunch he does not necessarily expect the frontier labs to own the category, because in smaller markets the cost to build something interesting is in the hundreds or thousands of dollars. That sentence is his remark about the class. It is not a price for Decider. TypeSafe’s chief executive, Diogo Almeida, told TechCrunch he did not yet see real competition, and that the current batch “seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful.” That quotation is his.
The picture is the training chart from the post, Figure 2, with a burgundy bar along the left edge. The title reads “Accuracy and calibration across the training trajectory.” The left axis is JevBench accuracy, drawn as a blue line. The right axis is the Brier score, drawn as a red dashed line. The horizontal axis runs through the versions to v19. The blue line rises as the versions advance. The red line falls. A note on the chart says higher accuracy is better and a lower Brier score is better. The frame does not print a calendar date. It is the training chart. It is not a photograph of a chip or a data center.
In plain terms, Strands Labs said on Thursday that Strands Decider 2B is an open model of about 2 billion parameters that you can run on your own machine. It does not write an essay. It picks among the options you give it, or it scores a scale you give it, and it returns a confidence. Strands says the release is v19: a Qwen3.5-2B torso, a pointer head of about a million parameters, and a small LoRA adapter of rank 16. Strands says that score would sit 3rd of 33 in the 2 billion class on a late-September JevBench snapshot, and that a typical local decision takes about 115 milliseconds on an RTX 3090. The post points the model at routing, tool choice, checks, and guardrails, and it includes a command-line tool plus a Strands agent example that can stop a tool call when an argument was guessed. TechCrunch covered the launch the same day and traced the class to TypeSafe’s Jev. The post itself says the model is a poor fit for coding, chat, and summaries, and that the hard decisions can stay with a large language model.
RELATED
Sources
- Strands Agents — Introducing Strands Decider 2B, 1 Oct 2026
strandsagents.com
- Hugging Face — StrandsAgents/strands-decider-2B-hobson-v19
huggingface.co
- GitHub — strands-labs/strands-decider
github.com
- TechCrunch — Amazon releases its own Jev clone, 1 Oct 2026
techcrunch.com
