Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction

Published in Preprint, 2024

Recommended citation: Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, Shafiq Joty (2024). Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction. Preprint.

Abstract

This work studies how early-layer representations can reduce input-token computation for long-context language models.