Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
Published in Preprint, 2024
Recommended citation: Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, Shafiq Joty (2024). Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction. Preprint.
Abstract
This work studies how early-layer representations can reduce input-token computation for long-context language models.
