Reflection previews Beam before its planned October weight release
Reflection has introduced Beam, a text model for coding, reasoning and tasks involving tools, while accepting applications for early access. Downloadable weights are planned for later October as final safety evaluations continue. The company describes separate capability and safety training and intends to publish a technical report, model card and developer tools with the weights under Apache 2.0.
Artificial Intelligence··Morning
Early access precedes downloadable weights
Reflection, an AI developer founded by former Google DeepMind researchers, introduced Beam on October 5. Its first planned open-weight model handles text and is designed for coding, reasoning and multistep tasks using tools. The company opened applications for a selected early-access group while final safety evaluations continued. The announcement therefore offers an early version through a waitlist; the downloadable model weights are scheduled for later October. Reflection plans to publish those weights with a technical report, model card and developer materials.[1], [2]
Training used separate pretraining and reinforcement-learning runs
Beam uses a mixture-of-experts architecture, activating a subset of its parameters for each token, the units into which model inputs are divided. Reflection disclosed 501 billion total parameters, with 23 billion active at a time. Its account of pretraining lists 23.8 trillion tokens from web material, public sources and proprietary licensed datasets. That run used 6,144 NVIDIA GB300 processors and finished in under four weeks. A separate reinforcement-learning run used approximately 10,500 GB300 processors for four weeks and generated more than 100 million task attempts, according to the company.[1], [2]
Midtraining used structured code repositories, long-running tasks and lengthy documents to extend the context to one million tokens. Reflection’s technical description also covers recovery from failures, saved training checkpoints and balancing the work assigned to the model’s experts. The training figures are disclosures by the developer.[1]
A separate safety model feeds the planned release
For safety work, Reflection trained a second model from the pretrained checkpoint. It combined a capability teacher’s behavior with a safety teacher’s behavior through multi-teacher distillation. The stated principles include following instructions, acknowledging uncertainty and maintaining the model’s AI identity. The company intends to publish its internal safety tests and safety results with the technical report. Apache 2.0 is the planned license, and the release package is expected to include tools for running and fine-tuning Beam.[1], [2]