Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

Published in System Card, 2026

Recommended citation: Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty (2026). Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models. System Card.
Paper Link: https://arxiv.org/abs/2607.01844

Abstract

This work presents a memory-efficient distributed training stack for large Mixture-of-Experts models.