Instella-Moe: An Open Mixture-of-Experts Language Model
14 hours ago
- AMD introduces Instella-MoE, a fully open Mixture-of-Experts language model with 16 billion total parameters and 2.8 billion active parameters per token.
- The model is trained from scratch on AMD Instinct MI300X and MI325X GPUs using the AMD ROCm software stack, incorporating architectural innovations like Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective connectivity.
- Instella-MoE undergoes a multi-stage training pipeline including pre-training, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning, with all artifacts openly released.
- The model achieves competitive performance on benchmarks against dense and MoE models of comparable or larger active parameter counts, with the final base model averaging 76.7 and the post-trained Think model averaging 73.22.
- AMD releases model weights, training configurations, data mixtures, and code for all stages to promote reproducibility and collaboration in the open AI community.