Sparse Mixture-of-Experts Architectures for Scaling Large Language Models: Routing, Specialization, and Computational Efficiency

Authors

  • Maya Srinivasan Author

Keywords:

LLM, sparse neural networks, MoE, Mixture of Experts

Abstract

Scaling large language models by increasing the number of parameters improves representational capacity but also substantially increases training and inference cost when all model parameters are activated for every input token. Sparse Mixture-of-Experts architectures address this limitation by conditionally routing tokens through subsets of specialized neural-network components. This review examines the evolution of Mixture-of-Experts language modeling through 2024, beginning with classical adaptive and hierarchical expert models and progressing toward sparsely gated neural networks, GShard, Switch Transformers, expert-choice routing, GLaM, sparse upcycling, Mixtral, and DeepSeekMoE. Modern sparse models separate total parameter capacity from active computational cost by using routing mechanisms that select a limited number of experts for each token. Mixtral demonstrated the practical feasibility of sparse expert activation within competitive open-weight large language models, while DeepSeekMoE emphasized fine-grained expert segmentation and shared experts to encourage stronger specialization. The review analyzes token-choice and expert-choice routing, auxiliary load-balancing losses, expert capacity, communication overhead, expert collapse, distributed training, inference memory, routing instability, and specialization. Particular attention is given to the relationship between the nominal parameter count of an MoE model and the much smaller subset of parameters activated during individual forward passes. Challenges include load imbalance, cross-device communication, expert redundancy, routing interpretability, memory requirements, and increasingly complex deployment infrastructure. By 2024, Mixture-of-Experts modeling had emerged as an important route toward increasing model capacity without proportionally increasing per-token computation.

 

References

Jacobs RA, Jordan MI, Nowlan SJ, Hinton GE. Adaptive mixtures of local experts. Neural Comput. 1991;3(1):79-87.

Jordan MI, Jacobs RA. Hierarchical mixtures of experts and the EM algorithm. Neural Comput. 1994;6(2):181-214.

Eigen D, Ranzato M, Sutskever I. Learning factored representations in a deep mixture of experts. arXiv. 2013;1312.4314.

Shazeer N, Mirhoseini A, Maziarz K, Davis A, Le Q, Hinton G, Dean J. Outrageously large neural networks: the sparsely-gated mixture-of-experts layer. In: International Conference on Learning Representations. 2017.

Lepikhin D, Lee H, Xu Y, Chen D, Firat O, Huang Y, et al. GShard: scaling giant models with conditional computation and automatic sharding. In: International Conference on Learning Representations. 2021.

Lewis M, Bhosale S, Dettmers T, Goyal N, Zettlemoyer L. BASE layers: simplifying training of large, sparse models. Proc Mach Learn Res. 2021;139:6265-6274.

Fedus W, Zoph B, Shazeer N. Switch Transformers: scaling to trillion parameter models with simple and efficient sparsity. J Mach Learn Res. 2022;23(120):1-39.

Du N, Huang Y, Dai AM, Tong S, Lepikhin D, Xu Y, et al. GLaM: efficient scaling of language models with mixture-of-experts. Proc Mach Learn Res. 2022;162:5547-5569.

Riquelme C, Puigcerver J, Mustafa B, Neumann M, Jenatton R, Pinto AS, et al. Scaling Vision with Sparse Mixture of Experts. Adv Neural Inf Process Syst. 2021;34:8583-8595.

Zhou Y, Lei T, Liu H, Du N, Huang Y, Zhao V, et al. Mixture-of-Experts with expert choice routing. Adv Neural Inf Process Syst. 2022;35:7103-7114.

Zoph B, Bello I, Kumar S, Du N, Huang Y, Dean J, et al. ST-MoE: designing stable and transferable sparse expert models. arXiv. 2022;2202.08906.

Komatsuzaki A, Puigcerver J, Lee S, Mustafa B, Ainslie J, Tay Y, et al. Sparse upcycling: training mixture-of-experts from dense checkpoints. In: International Conference on Learning Representations. 2023.

Gale T, Narayanan D, Young C, Zaharia M. MegaBlocks: efficient sparse training with mixture-of-experts. Proc Mach Learn Syst. 2023;5:288-304.

Clark A, De Las Casas D, Guy A, Mensch A, Paganini M, Hoffmann J, et al. Unified scaling laws for routed language models. Proc Mach Learn Res. 2022;162:4057-4086.

Jiang AQ, Sablayrolles A, Roux A, Mensch A, Savary B, Bamford C, et al. Mixtral of Experts. arXiv. 2024;2401.04088.

Dai D, Deng C, Zhao C, Xu RX, Gao H, Chen D, et al. DeepSeekMoE: towards ultimate expert specialization in mixture-of-experts language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024. p. 1280-1297.

DeepSeek-AI. DeepSeek-V2: a strong, economical, and efficient mixture-of-experts language model. arXiv. 2024;2405.04434.

Roller S, Sukhbaatar S, Szlam A, Weston J. Hash layers for large sparse models. Adv Neural Inf Process Syst. 2021;34:17555-17566.

Rajbhandari S, Li C, Yao Z, Zhang M, Aminabadi RY, Awan AA, et al. DeepSpeed-MoE: advancing mixture-of-experts inference and training to power next-generation AI scale. Proc Mach Learn Syst. 2022;4:243-259.

Mustafa B, Riquelme C, Puigcerver J, Jenatton R, Houlsby N. Multimodal contrastive learning with LIMoE: the language-image mixture of experts. Adv Neural Inf Process Syst. 2022;35:9564-9576.

Published

2024-06-01