← News

Reflection's official Beam launch graphic: the Reflection logo mark and the word Beam beside concentric rings dissolving into dots on a dark green field

5 Oct 2026

Reflection AI

Reflection unveils Beam, a 501B open-weight model it says matches China's GLM 5.2 with 3–4× less compute

Reflection AI, the New York startup founded in 2024 by former Google DeepMind researchers, introduced Beam on Monday, Oct. 5, 2026, its first open-weight model. Beam is a mixture-of-experts model with 501 billion total parameters and 23 billion active per token, aimed at coding and agent work. Reflection says it scores close to Z.ai's GLM 5.2 on reasoning while using three to four times less inference compute. For now it is early access only; weights, a technical report, and a model card are promised later in October.

For about two years the best free-to-download AI models have mostly come from China, and that's been a real problem for US companies and governments that want to run AI on their own machines without trusting a Chinese model. Beam is the most serious American attempt yet to close that gap, and the efficiency claim is the interesting part: if it really matches GLM 5.2 at a fraction of the compute, it's cheap enough to actually use. But every number here is Reflection grading its own homework, and nobody outside early access can run it yet. Wait for the weights later this month before calling it a win.

On Monday, 5 October 2026, Reflection AI published “Introducing Beam: Reflection’s 501B open-weight model.” The page prints October 5, 2026. It does not print an hour. Reflection calls Beam its first open-weight model. Open-weight means the learned numbers, the weights, are meant to be released so someone else can run the model. Beam is a sparse mixture-of-experts model. A mixture of experts splits the model into many specialist pieces. Sparse means that for each small piece of text, a token, only some of those pieces run. Reflection says Beam has 501 billion parameters in total and 23 billion active per token. A parameter is one learned number. Active per token is the slice that does the work for each piece of text. The rest stays in memory. Reflection says it built Beam for coding, reasoning, and agent work. An agent, here, is software that can take steps on a task, not only reply in chat. Reflection says Beam is text-only. It reads and writes text. It does not take pictures or sound. Those lines are Reflection’s.

What a person can use today. Reflection says Beam is “undergoing final red-teaming and evaluations.” Red-teaming means testers try to make the model fail, or do harm, before a wider release. Developers can sign up now for early access. Reflection says that early version goes to a select group, through a waitlist. The company says it will release the weights, a technical report, a model card, and developer tools later this month. It also says those weights will ship under the Apache 2.0 license, with documentation and the pieces needed to run, test, and fine-tune the model. Apache 2.0 is a common license that lets others use, change, and share the work. Early access today is a sign-up. The weight release is the later-October promise.

How Reflection says it trained Beam. The post says the model was pretrained on 23.8 trillion tokens drawn from the web and from licensed data. Pretraining is the long first pass that sets the model’s numbers. A token is a chunk of text, often a word or part of a word, so 23.8 trillion tokens is an enormous amount of reading. Reflection then describes a reinforcement-learning run on 10,500 NVIDIA GB300 chips for four weeks. Reinforcement learning, shortened to RL, means the model practices tasks and gets a score on each attempt. A GB300 is an NVIDIA data-center graphics chip. Reflection says the run generated more than 100 million rollouts, the practice attempts, with context up to 256,000 tokens. Context is how much text the model can hold in view at once. The post says training and grading used about 1.3 billion sandboxes, temporary computers the model practices inside, and one million environments for coding, agent work, and science, technology, engineering, and math. Reflection calls the job one of the largest RL runs by any open lab. It says the model’s abilities kept improving as that RL compute grew, “with no sign of a plateau.” A plateau would be a stretch where more practice stops helping. Those lines are Reflection’s.

What Reflection says Beam can do. The scores below are the company’s, from the table on its own post. Reflection says Beam is competitive with Z.ai’s GLM 5.2, and approaching Alibaba’s Qwen 3.8-Max, on coding and agent tasks. It says frontier open models such as Kimi K3 remain ahead on raw capability. On reasoning tests, Reflection says Beam reaches scores comparable to GLM 5.2 while using three to four times less inference compute. Inference is the work of writing an answer, after training is finished. Compute, here, is the math the chips do for that answer. Reflection says the comparison is an estimate. It leaves out prompt prefill, the first pass over the question, and serving overhead, the extra cost of running the service around the math. Semafor’s interview is the same multiple in Misha Laskin’s words: he said Beam needs three to four times less computing power to reason through a problem than comparable open models. That sentence is his, as Semafor prints it. It is not a second lab’s measurement.

The example scores Reflection prints for Beam. SWE-bench Verified is 80.9. Terminal Bench 2.1 is 80.1. GPQA Diamond is 90.5. AIME 2026 is 97.8. Humanity’s Last Exam, with no tools, is 36.2. SWE-bench Verified is a test of fixing real software bugs that reviewers have already checked. Terminal Bench is a test of tasks in a terminal, the text window where a person types commands. GPQA Diamond is a set of hard graduate-level science questions. AIME is a high-school math contest. Humanity’s Last Exam is a broad, difficult exam. On the same table, where GLM 5.2 has a number, Reflection prints 81.0 on Terminal Bench 2.1, 91.2 on GPQA Diamond, 99.2 on AIME 2026, and 40.5 on Humanity’s Last Exam with no tools. The GLM 5.2 cell for SWE-bench Verified is empty. Reflection marks empty cells as not reported. “Comparable” is Reflection’s word for that table.

A control Reflection describes. A person can set a reasoning-effort level. A lower setting favors a shorter answer. A higher setting lets the model reason longer when the task is hard. Reflection says that is how a user trades response length for accuracy, and matches the effort to the task and the compute budget. Those lines are Reflection’s.

What the chief executive told Semafor. Misha Laskin, cofounder and chief executive, called Beam a “workhorse.” He said it is particularly strong at coding and agent tasks, and that it performs better than other Western open models. He said Reflection is already training a next model that would be “much more” powerful. He said the buyers in view are businesses and governments building sovereign AI: they want to run a model on machines they control, and they do not, or cannot, use Chinese models. His line, as Semafor prints it: “They don’t really have very good options today.” Those sentences are Semafor’s account of the interview.

Who the company is, as Semafor reports it. Reflection is a New York startup founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. Semafor says the investors include Nvidia, Sequoia, and Citigroup, and that in June Reflection confirmed it had closed a round at a $25 billion pre-money valuation. Pre-money is the price of the company before that new money is counted in. Those lines are Semafor’s.

The wider picture those two newsrooms draw. Semafor says Chinese labs have led open-weight AI, and that some large Western companies use Chinese open models to cut the cost of repetitive work, such as customer service and routine code. It names Mistral and Thinking Machines Lab as other Western startups that offer open-weight models. The day before Beam’s post, Axios reported that several Western open-weight models were set to arrive in October, Reflection’s first model among them. Axios said a Reflection spokesperson declined to comment for that story. The Axios piece is dated Oct. 4, 2026. It is a preview. The launch post is Reflection’s, on Oct. 5.

The picture is Reflection’s launch graphic for Beam. A dark green field fills the frame. On the left, a white mark of two circles sits beside the word Beam. One circle is a tight spiral. The other is a solid disc. To the right, thin white rings overlap. Some rings are solid lines and some are made of dots. The rings break apart into a field of white dots that runs off the right edge. The frame does not print a calendar date. It is the company’s graphic for this release. It is not a photograph of a lab, a chip, or a person.

In plain terms, Reflection on Monday introduced Beam, a text-only mixture-of-experts model it sizes at 501 billion parameters, with 23 billion of them active for each token. The company says the model is close to Z.ai’s GLM 5.2 on reasoning, at a fraction of the inference compute, on Reflection’s own table and Reflection’s own estimate. Early access is a sign-up. Reflection says the weights, a technical report, and a model card come later this month.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources