Hasty Briefsbeta

Bilingual

Instella-Moe: An Open Mixture-of-Experts Language Model

14 hours ago
  • AMD introduces Instella-MoE, a fully open Mixture-of-Experts language model with 16 billion total parameters and 2.8 billion active parameters per token.
  • The model is trained from scratch on AMD Instinct MI300X and MI325X GPUs using the AMD ROCm software stack, incorporating architectural innovations like Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective connectivity.
  • Instella-MoE undergoes a multi-stage training pipeline including pre-training, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning, with all artifacts openly released.
  • The model achieves competitive performance on benchmarks against dense and MoE models of comparable or larger active parameter counts, with the final base model averaging 76.7 and the post-trained Think model averaging 73.22.
  • AMD releases model weights, training configurations, data mixtures, and code for all stages to promote reproducibility and collaboration in the open AI community.