⏱️ Reading time: 8 min
Reflection AI unveiled the open Beam model, a mixture-of-experts system with 501 billion total parameters but only 23 billion active per token. The startup trained it for code, reasoning, and agentic tasks, and promises to release the weights this month.
📑 En este artículo
The most striking detail isn’t the size but the efficiency. According to Reflection, Beam matches or beats open rivals like GLM-5.2 while using 3 to 4 times less inference compute per generated token.
TL;DR
- Reflection launched Beam, an MoE model with 501 billion total parameters and 23 billion active.
- Training used 23.8 trillion tokens and more than 100 million RL rollouts on GB300 GPUs.
- Beam competes with GLM-5.2 in reasoning while using 3 to 4 times less inference compute.
- The weights, technical report, and developer tools are coming out this month.
- Final red-teaming is still pending before early access to the model opens up.
What the open Beam model is
The open Beam model is Reflection AI’s first open-weight family, a startup focused on language models for software agents. It’s a mixture-of-experts (MoE) model with 501 billion total parameters and 23 billion active per token, designed for coding, reasoning, and tool use.
The mixture-of-experts architecture (MoE) splits the network into specialized blocks, called experts, and activates only some of them on each token. That’s why Beam loads 501 billion parameters into memory, but the actual computation per token uses just 23 billion, about 22 times less. It’s the same logic used by GLM-5.2 and Kimi K3, though Reflection combines it with reinforcement learning at a different scale.
What happened with Reflection AI’s Beam
Reflection built Beam on two pillars. The first was pretraining on 23.8 trillion tokens curated from the web and licensed datasets, which the company says allows it to match or beat open base models of similar size.
The second pillar was a large-scale reinforcement learning (RL) run. Reflection developed its own algorithms, training environments, and infrastructure to sustain that compute, generating more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs over four weeks.
The company describes this run as one of the largest ever reported by an open lab. It’s a figure only Reflection can confirm with its own records, but the volume of rollouts and training sandboxes detailed in its report backs up the claim.
The model is still in the final red-teaming and safety evaluation stage. Reflection opened a waitlist for early access and announced that the weights, technical report, model card, and developer tools will be released later this month.
Context: the open model race
Reflection positions Beam as the model pushing the Western open frontier forward. In code and agent-use benchmarks, the company describes it as competitive with larger open models like GLM-5.2, and close to Qwen 3.8-Max. Models like Kimi K3 still lead in raw capability, but Reflection argues Beam’s advantage lies in inference efficiency, not in winning every benchmark.
The comparison also includes GLM-5.3, Nemotron 3 Ultra, DeepSeek V4.1 Flash, and Inkling, Thinking Machines’ mixture-of-experts model. All of them compete today in the same category: large, open-weight models built to code and operate tools autonomously.
Worth clarifying: most of the open models currently leading the benchmark tables (GLM, Kimi, Qwen, DeepSeek) come from Chinese labs. Reflection explicitly frames Beam as an open option trained outside that bloc, something that matters for companies whose internal policies prefer not to depend on Chinese models or closed APIs.
How it works under the hood: the mechanism behind Beam
The most expensive part of building Beam wasn’t pretraining, but the RL run. Reflection used 10,500 NVIDIA GB300 GPUs over four weeks to generate its more than 100 million rollouts, with a context window of up to 256,000 tokens during training. To grade and train on those rollouts, the company ran approximately 1.3 billion sandboxes, fed by a million coding, agent-use, and science environments that Reflection built specifically for this run.
Think of it as an emergency room with a million simulated cases: each sandbox is a patient with a real coding or tool-use problem, and the model gets a grade for how well it did. With enough cases and enough grades, the model learns which reasoning strategy works best for each type of problem, not just how to repeat text similar to its training data.
According to Reflection’s report, the model’s capabilities kept improving as RL compute scaled up, with no signs of plateauing. To give a sense of scale, the company compares its run to the 30 million rollouts used to train Inkling and the 753,000 used for MiMo: Beam trained on a rollout volume between 3 and 130 times larger than those two models.
flowchart TD
A["Pretraining: 23.8T tokens"] --> B["Large-scale RL: 10,500 GB300 GPUs, 4 weeks"]
B --> C["100M+ rollouts, 1,300M sandboxes"]
C --> D["Beam: 501B total, 23B active"]
D --> E["Efficient inference in production"]
Reflection measures Beam’s efficiency in estimated FLOPs, not in inference cost measured in production. The calculation multiplies the number of activated parameters by the tokens generated per attempt, leaving out prompt prefill, context-dependent attention, and serving overhead. There’s no direct way to verify this comparison from the outside, because Reflection didn’t publish the actual inference runs behind the chart, only the estimate’s methodology.
| Model | Terminal-Bench v2.1 | GPQA Diamond | HLE (no tools) |
|---|---|---|---|
| Beam | 80.1 | 90.5 | 36.2 |
| GLM-5.2 | 81.0 | 91.2 | 40.5 |
| GLM-5.3 | 88.2 | 91.7 | 42.3 |
| Kimi K3 | 88.3 | 93.5 | 46.9 |
| Qwen 3.8-Max | 86.6 | 92.6 | 43.6 |
| DeepSeek V4.1 Flash | 90.6 | 90.9 | 39.1 |
| Nemotron 3 Ultra | 56.4 | 87.0 | 26.7 |
| Inkling | 63.8 | 87.2 | 29.7 |
The numbers come from Reflection’s own report and from Artificial Analysis data, which the company uses as a reference for rival models. Beam falls below GLM-5.3, Kimi K3, and Qwen 3.8-Max on these three benchmarks, but Reflection argues those models need far more activated parameters per token to get there.
Impact and analysis
The open Beam model shows that the race is no longer won just with the highest benchmark score: it’s also won by how much each response costs. If a company in Latin America needs to deploy a code assistant at scale, paying for 23 billion activated parameters is much cheaper than paying for the more than 2 trillion that models like Qwen 3.8-Max activate on every token.
This doesn’t settle the debate over who leads the open model race. Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash still lead Beam in almost every benchmark in the table above. What’s changed is that there’s now an open option, trained outside China, competing in that same category with an explicit efficiency pitch.
💭 Key takeaway: Beam isn’t trying to be the most capable model on the chart. It’s trying to be the cheapest to run among those near the top.
What’s next for Beam
Reflection hasn’t released Beam’s weights yet. The company says it will complete final red-teaming and publish the weights, technical report, model card, and developer tools during the rest of October 2026. Until then, access is waitlist-only.
If the timeline holds, Beam would join GLM-5.2, GLM-5.3, Kimi K3, Qwen 3.8-Max, and DeepSeek V4.1 Flash as another option for running locally or on self-managed infrastructure, without depending on a closed API.
Try it yourself: sign up for the early access waitlist on Reflection AI’s official announcement to run Beam as soon as the weights are released this month.
Frequently Asked Questions
When will Reflection AI release Beam’s weights?
Reflection announced it plans to release the weights, technical report, model card, and developer tools during the rest of October 2026, once final red-teaming is complete.
How many parameters does Beam activate per token?
Beam activates 23 billion parameters per token, out of a total of 501 billion. It’s a mixture-of-experts architecture, so most of the parameters stay inactive at each computation step.
Is Beam better than GLM-5.2 or Qwen 3.8-Max?
On benchmarks like GPQA Diamond and Terminal-Bench v2.1, Beam falls just slightly below GLM-5.2 and Qwen 3.8-Max. Reflection argues the real difference lies in efficiency: Beam needs 3 to 4 times less compute per token than GLM-5.2 to reach a comparable result.
What hardware did Reflection AI use to train Beam?
The reinforcement learning run used 10,500 NVIDIA GB300 GPUs over four weeks, generating more than 100 million rollouts with up to 256,000 tokens of context.
Is Beam available to try today?
No. Beam is in the final red-teaming stage, and Reflection is only offering an early access waitlist; according to the company, public weights will follow later this month.
References
- Reflection AI: official Beam announcement, with benchmarks, RL methodology, and release timeline.
- Wikipedia: general explanation of the mixture-of-experts (MoE) architecture.
- Artificial Analysis: data source used by Reflection to compare Beam with other models.
- NVIDIA: maker of the GB300 GPUs used in Beam’s RL run.
📱 Like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de BoliviaInteligente en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments