Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
Published in System Card, 2026
Recommended citation: Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty (2026). Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models. System Card.
Paper Link: https://arxiv.org/abs/2607.01844
Abstract
This work presents a memory-efficient distributed training stack for large Mixture-of-Experts models.
