Hasty Briefsbeta

Bilingual

Beam: Reflection's 501B open-weight model

9 hours ago
  • Beam is an open-weight sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active, optimized for coding, reasoning, and agentic tasks.
  • It was pretrained on 23.8 trillion high-quality tokens and trained with a massive reinforcement learning run using 10.5K NVIDIA GB300 GPUs for four weeks, generating over 100 million rollouts.
  • Beam achieves competitive performance with frontier open models like GLM-5.2 and approaches Qwen 3.8-Max on coding and agentic benchmarks, while using 3–4× less inference compute on reasoning tasks.
  • The model features a controllable reasoning effort parameter, allowing users to balance reasoning depth and token usage for task-specific efficiency.
  • RL training developed generalizable agentic capabilities, such as browsing and tool use, even without explicit training on those domains.
  • Advanced infrastructure supported fully asynchronous execution, high resilience, efficient trainer packing, and robust reward integrity, enabling stable training at scale.
  • Pretraining emphasized curated data quality, including proprietary licensed datasets, and achieved stable MoE optimization with near-perfect expert utilization.
  • Midtraining extended context length to 1M tokens and built a strong foundation for RL by exposing the model to reasoning-rich, long-form documents.
  • Safety alignment used multi-teacher distillation and deliberative alignment techniques to enforce rules, qualities, and a proactive interaction style.
  • The model will be released under an Apache 2.0 license later this month, with full documentation and ecosystem integration.