TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Reflection has introduced Beam, a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active. The company reports strong coding and agentic benchmark results and says Beam uses less inference compute than some larger models, but the weights, technical report and model card have not yet been released.
Reflection has introduced Beam, a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, aimed at coding, reasoning and agentic workloads. The company says Beam’s weights, technical report, model card and developer materials will follow later this month, after final red-teaming and evaluations.
Reflection describes Beam as its first open-weight model. It says the model was pretrained on 23.8 trillion tokens from curated web material and proprietary licensed datasets. Reflection reports that Beam matches or outperforms available open base models of similar size, though it has not yet published the weights or technical report needed for outside assessment.
The company says it trained Beam with a major reinforcement-learning effort: more than 100 million rollouts generated over four weeks on 10,500 NVIDIA GB300 GPUs. Training and grading used about 1.3 billion sandboxes, according to Reflection. The company also says it assembled one million coding, agentic and STEM environments for the work.
Reflection reports that Beam is competitive with larger open models on coding and agentic tasks and approaches Qwen 3.8-Max on those measures. It says Kimi K3 remains ahead on raw capability, while Beam’s advantage is inference efficiency. On advanced reasoning benchmarks, Reflection says Beam scored comparably to GLM-5.2 while using three to four times less inference compute. These are company-reported comparisons; some benchmark entries for other models are marked unreported.
A Smaller Active Model for Coding
Beam’s architecture activates 23 billion parameters per token despite having 501 billion in total. That distinction matters for serving: Reflection argues that lower active parameter counts and shorter generated answers can reduce inference compute while retaining strong performance on coding and agentic tasks.
If independent evaluations support those results, Beam could offer organizations a model suited to repeated software tasks and tool-using workflows at a lower compute burden than some larger systems. The practical value remains unverified until users can access the model and compare it under consistent conditions. Reflection’s comparisons are based on estimates and selected benchmarks, not published measurements of actual deployment costs.
As an affiliate, we earn on qualifying purchases.
Reflection’s Reinforcement Learning Push
Reflection says its training strategy combined large-scale pretraining with reinforcement learning designed for multi-step tasks. The company used asynchronous policy gradients and says it developed methods to keep training stable when samples came from older model versions and when training and inference systems produced numerical differences.
According to the company, its largest RL campaign included about 80 million rollouts for a reasoning expert within the broader effort of more than 100 million. Reflection says performance continued to improve as rollout volume rose, with no sign of a plateau in its evaluation suite. That trend is based on the company’s own training and benchmark data.
“Beam is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks.”
— Reflection
As an affiliate, we earn on qualifying purchases.
Independent Results Still Pending
Beam is still undergoing red-teaming and evaluations, and the company has not yet released its weights or technical report. That means independent researchers cannot currently reproduce the reported results or assess the model’s behavior directly.
Reflection’s compute comparisons are estimates based on active parameters and generated-token counts. The company says these calculations exclude prompt prefill, context-dependent attention operations and serving overhead, so they are not measured deployment costs. The exact release date, access terms and final evaluation results have not been specified in the source announcement.
As an affiliate, we earn on qualifying purchases.
Weights and Evaluations Due Later
Reflection says it plans to publish Beam’s weights, technical report, model card and developer artifacts later this month. The company is also offering a sign-up for early access while it completes red-teaming and evaluations. Once those materials are available, outside testing can clarify how Beam performs across tasks, how its efficiency compares under consistent measurement, and what safeguards or limitations apply.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Beam?
Beam is Reflection’s first open-weight model, a sparse mixture-of-experts system designed for coding, reasoning and agentic workloads. It has 501 billion total parameters, with 23 billion active.
Are Beam’s weights available now?
No. Reflection says Beam is undergoing final red-teaming and evaluations. It plans to release the weights and related materials later this month, but has not given an exact date.
What performance has Reflection reported?
Reflection says Beam is competitive with larger open models on coding and agentic tasks and reports reasoning scores comparable to GLM-5.2 with three to four times less inference compute. Those comparisons remain company-reported pending independent testing.
How much compute did Beam’s reinforcement learning use?
Reflection says the effort generated more than 100 million rollouts over four weeks on 10,500 NVIDIA GB300 GPUs. It also reports using about 1.3 billion sandboxes for training and grading.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
