Beam: Reflection's 501B open-weight model
9 hours ago
- Beam is an open-weight sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active, optimized for coding, reasoning, and agentic tasks.
- It was pretrained on 23.8 trillion high-quality tokens and trained with a massive reinforcement learning run using 10.5K NVIDIA GB300 GPUs for four weeks, generating over 100 million rollouts.
- Beam achieves competitive performance with frontier open models like GLM-5.2 and approaches Qwen 3.8-Max on coding and agentic benchmarks, while using 3–4× less inference compute on reasoning tasks.
- The model features a controllable reasoning effort parameter, allowing users to balance reasoning depth and token usage for task-specific efficiency.
- RL training developed generalizable agentic capabilities, such as browsing and tool use, even without explicit training on those domains.
- Advanced infrastructure supported fully asynchronous execution, high resilience, efficient trainer packing, and robust reward integrity, enabling stable training at scale.
- Pretraining emphasized curated data quality, including proprietary licensed datasets, and achieved stable MoE optimization with near-perfect expert utilization.
- Midtraining extended context length to 1M tokens and built a strong foundation for RL by exposing the model to reasoning-rich, long-form documents.
- Safety alignment used multi-teacher distillation and deliberative alignment techniques to enforce rules, qualities, and a proactive interaction style.
- The model will be released under an Apache 2.0 license later this month, with full documentation and ecosystem integration.