Long-Context Large Language Models: Efficient Attention, Positional Extrapolation, Streaming Inference, and Memory Optimization
Keywords:
positional encoding, Transformer, efficient attentionAbstract
The context window determines how much information a large language model can process during an individual inference sequence. Expanding this window became a major research priority as language models were increasingly applied to long documents, books, source-code repositories, multi-turn conversations, and retrieval-intensive workflows. Standard Transformer self-attention incurs quadratic memory and computational costs with sequence length, while models trained on shorter sequences frequently fail to generalize to unseen positional ranges. This review examines methods for extending and efficiently processing long contexts through 2024. Transformer-XL introduced segment-level recurrence, while sparse-attention systems including Longformer and BigBird reduced attention complexity. Rotary positional embeddings and attention with linear biases offered alternative positional representations. FlashAttention and FlashAttention-2 improved exact attention efficiency through IO-aware GPU implementations. Positional interpolation, YaRN, and LongRoPE extended pretrained models beyond their original context lengths through modifications to positional encodings. StreamingLLM demonstrated that retaining initial attention-sink tokens alongside a rolling cache could support stable streaming inference. LongLLMLingua addressed another dimension of the problem by compressing prompts before inference. The review also examines key-value cache growth, positional degradation, retrieval within long contexts, attention sinks, memory bandwidth, long-context evaluation, and the difference between nominal context-window size and effective information utilization. Evidence available by 2024 showed that context length had become a systems, architecture, representation, and evaluation problem rather than a single model hyperparameter.
References
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.
Dai Z, Yang Z, Yang Y, Carbonell J, Le QV, Salakhutdinov R. Transformer-XL: attentive language models beyond a fixed-length context. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. p. 2978-2988.
Kitaev N, Kaiser Ł, Levskaya A. Reformer: the efficient Transformer. In: International Conference on Learning Representations. 2020.
Beltagy I, Peters ME, Cohan A. Longformer: the long-document Transformer. arXiv. 2020;2004.05150.
Zaheer M, Guruganesh G, Dubey A, Ainslie J, Alberti C, Ontañón S, et al. Big Bird: Transformers for longer sequences. Adv Neural Inf Process Syst. 2020;33:17283-17297.
Su J, Lu Y, Pan S, Murtadha A, Wen B, Liu Y. RoFormer: enhanced Transformer with rotary position embedding. Neurocomputing. 2024;568:127063.
Press O, Smith NA, Lewis M. Train short, test long: attention with linear biases enables input length extrapolation. In: International Conference on Learning Representations. 2022.
Dao T, Fu D, Ermon S, Rudra A, Ré C. FlashAttention: fast and memory-efficient exact attention with IO-awareness. Adv Neural Inf Process Syst. 2022;35:16344-16359.
Chen S, Wong S, Chen L, Tian Y. Extending context window of large language models via positional interpolation. arXiv. 2023;2306.15595.
Chen Y, Qian S, Tang H, Lai X, Liu Z, Han S, Jia J. LongLoRA: efficient fine-tuning of long-context large language models. In: International Conference on Learning Representations. 2024.
Peng B, Quesnelle J, Fan H, Shippole E. YaRN: efficient context window extension of large language models. In: International Conference on Learning Representations. 2024.
Xiao G, Tian Y, Chen B, Han S, Lewis M. Efficient streaming language models with attention sinks. In: International Conference on Learning Representations. 2024.
Dao T. FlashAttention-2: faster attention with better parallelism and work partitioning. In: International Conference on Learning Representations. 2024.
Ding Y, Zhang LL, Zhang C, Xu Y, Shang N, Xu J, et al. LongRoPE: extending LLM context window beyond 2 million tokens. Proc Mach Learn Res. 2024;235:11091-11104.
Jiang H, Wu Q, Luo X, Li D, Lin CY, Yang Y, Qiu L. LongLLMLingua: accelerating and enhancing LLMs in long context scenarios via prompt compression. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024. p. 1658-1677.
Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, Liang P. Lost in the middle: how language models use long contexts. Trans Assoc Comput Linguist. 2024;12:157-173.
Bai Y, Lv X, Zhang J, Lyu H, Tang J, Huang Z, et al. LongBench: a bilingual, multitask benchmark for long context understanding. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024.
Zhang X, Chen C, Li X, Lin T, Ma Z, Wang C, et al. Infinite-LLM: efficiently deploying LLM service for long context with distattention and distributed KVCache. arXiv. 2024;2401.02669.
Liu X, Dong P, Hu X, Chu X. LongGenBench: long-context generation benchmark. In: Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. p. 865-883.
Gu A, Dao T. Mamba: linear-time sequence modeling with selective state spaces. arXiv. 2023;2312.00752.