Reflection AI announced Beam, the company’s first open-weights artificial intelligence model, with a clear focus on programming, reasoning, and tasks carried out by AI agents. The model is based on a Sparse Mixture-of-Experts architecture and contains 501 billion parameters in total, but only 23 billion of them are active when processing each token.
The company says Beam was pretrained on 23.8 trillion tokens drawn from web content and licensed private datasets. It also provides a context window of one million tokens, a capacity intended for handling long documents and tasks. In addition to pretraining, Reflection AI used reinforcement learning on a large scale to improve the model’s performance.
Intensive Training and a Focus on Inference Costs
During the reinforcement-learning phase, the company ran 10,500 Nvidia GB300 graphics processing units for four weeks and produced more than 100 million rollouts. The number of sandbox operations used in training and evaluation reached approximately 1.3 billion operations.
Reflection AI places inference efficiency at the forefront of its messaging about Beam. According to the company’s tests, the model delivers performance close to Z.ai’s GLM-5.2 on advanced reasoning benchmarks, while using three to four times less computing capacity during inference. However, these results have not yet undergone independent verification, and the company acknowledges that larger open models, such as Kimi K3, outperform Beam in raw performance in some cases.
Programming Results and Control Over Inference Effort
Reflection AI reported that Beam scored 77.2 on SWE Bench Pro v2-Hard, 80.1 on Terminal Bench v2.1, and 80.9 on SWE Bench Verified. The company says the model competes with GLM-5.2 and approaches larger Qwen models on some tasks, but it does not claim to outperform competitors on all tests.
Beam includes a reasoning effort parameter that allows users to specify the amount of computing time allocated to reasoning. Lower levels produce shorter answers at a lower cost, while higher levels allow the model to use more tokens to process complex tasks.
What Changes in Practice?
This design means that Beam’s value depends not only on the model’s size or its highest benchmark score, but also on balancing answer quality against the resources used to produce it. This could benefit developers and organizations running programming tasks or agents that rely on repeatedly calling the model. However, the practical benefit will remain tied to independent results, infrastructure costs, and the terms governing access to the weights and tools.
Beam is a text-only model, not a multimodal one. When provided with tool use and web access, it can search different sources and make use of information supplied by external tools. Reflection AI says the model learned during training to use other language models and to turn to OCR interfaces to read documents when web access is available.
Availability and the Broader Strategy
The company is targeting enterprises, developers, and government entities, describing Beam as a practical model suitable for use in everyday workloads. The plan falls within Reflection AI’s push to build what it calls “AI factories,” enabling organizations to customize systems that operate on their data and run them locally.
Reflection AI was founded in 2024 by former Google DeepMind researchers and has raised approximately $4.7 billion in investments, with participation from Nvidia, Sequoia Capital, and Lightspeed Venture Partners. Its pre-money valuation in the latest funding round was approximately $25 billion. The company plans to release Beam’s weights, technical report, model card, and developer tools during October, after completing safety tests and final evaluations.