Hardware-Aware Efficient Inference for Long-Context and Reasoning Language Models: Low-Bit Quantization, KV-Cache Compression, and Speculative Decoding
Keywords:
AnyBCQ, ChanMix, speculative decoding, KV cache, quantizationAbstract
The deployment cost of large language models increasingly depends not only on parameter count but also on context length, reasoning-token generation, key-value cache growth, accelerator memory bandwidth, batch size, and decoding strategy. These pressures are amplified in reasoning models, which may generate thousands of intermediate tokens, and in long-context systems, where persistent key-value states can dominate runtime memory consumption. This review examines hardware-aware language-model inference techniques through September 2026. Post-training quantization methods including GPTQ, SmoothQuant, AWQ, and low-bit weight formats substantially reduced parameter-memory requirements. BitNet investigated extremely low-bit model representations trained natively rather than quantized after training. FlashAttention and FlashAttention-2 improved exact attention by restructuring memory access, while vLLM's PagedAttention improved key-value-cache management for concurrent serving. StreamingLLM and related long-context methods highlighted the importance of cache organization for persistent generation. KIVI and subsequent KV-cache quantization methods targeted the rapidly growing inference-state memory associated with long contexts. Speculative decoding reduced latency by allowing smaller draft models to propose tokens verified by a larger target model, while Medusa and EAGLE expanded multi-token prediction and draft-generation strategies. By 2026, hardware-aware inference increasingly involved adaptive precision and sensitivity-aware cache compression. ChanMix introduced mixed-precision allocation across key-value-cache channels for long-context inference. AnyBCQ proposed flexible hardware-oriented binary-coded quantization enabling runtime multi-precision operation. Hierarchical Block Quantization explored block-level hardware efficiency alongside second-level scaling, while asynchronous test-time scaling demonstrated that reasoning throughput can also be improved through speculative and asynchronous execution. The review argues that model quality and inference efficiency must increasingly be co-designed across algorithms, numeric formats, kernels, cache architectures, and serving systems.
References
Dettmers T, Lewis M, Belkada Y, Zettlemoyer L. LLM.int8(): 8-bit matrix multiplication for Transformers at scale. Adv Neural Inf Process Syst. 2022;35:30318-30332.
Frantar E, Ashkboos S, Hoefler T, Alistarh D. GPTQ: accurate post-training quantization for generative pre-trained Transformers. In: International Conference on Learning Representations. 2023.
Xiao G, Lin J, Seznec M, Wu H, Demouth J, Han S. SmoothQuant: accurate and efficient post-training quantization for large language models. Proc Mach Learn Res. 2023;202:38087-38099.
Lin J, Tang J, Tang H, Yang S, Chen WM, Wang WC, et al. AWQ: activation-aware weight quantization for on-device LLM compression and acceleration. Proc Mach Learn Syst. 2024;6.
Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L. QLoRA: efficient finetuning of quantized LLMs. Adv Neural Inf Process Syst. 2023;36.
Ma S, Wang H, Ma L, Wang L, Wang W, Huang S, et al. The era of 1-bit LLMs: all large language models are in 1.58 bits. arXiv. 2024;2402.17764.
Ma S, Wang H, Huang S, Zhang X, Hu Y, Song T, et al. BitNet b1.58 2B4T technical report. arXiv. 2025;2504.12285.
Dao T, Fu D, Ermon S, Rudra A, Ré C. FlashAttention: fast and memory-efficient exact attention with IO-awareness. Adv Neural Inf Process Syst. 2022;35:16344-16359.
Dao T. FlashAttention-2: faster attention with better parallelism and work partitioning. In: International Conference on Learning Representations. 2024.
Kwon W, Li Z, Zhuang S, Sheng Y, Zheng L, Yu CH, et al. Efficient memory management for large language model serving with PagedAttention. In: ACM Symposium on Operating Systems Principles. 2023.
Xiao G, Tian Y, Chen B, Han S, Lewis M. Efficient streaming language models with attention sinks. In: International Conference on Learning Representations. 2024.
Liu Z, Yuan J, Jin H, Zhong S, Xu Z, Braverman V, et al. KIVI: a tuning-free asymmetric 2bit quantization for KV cache. Proc Mach Learn Res. 2024.
Leviathan Y, Kalman M, Matias Y. Fast inference from Transformers via speculative decoding. Proc Mach Learn Res. 2023;202:19274-19286.
Chen C, Borgeaud S, Irving G, Lespiau JB, Sifre L, Jumper J. Accelerating large language model decoding with speculative sampling. arXiv. 2023;2302.01318.
Cai T, Li Y, Geng Z, Peng H, Lee J, Chen D, Dao T. Medusa: simple LLM inference acceleration framework with multiple decoding heads. arXiv. 2024;2401.10774.
Li Y, Wei F, Zhang C, Zhang H. EAGLE: speculative sampling requires rethinking feature uncertainty. Proc Mach Learn Res. 2024.
Liao C, Wen Z. Channel-aware mixed-precision quantization for efficient long-context inference. In: International Conference on Learning Representations. 2026.
Park G, Bae J, Kwon B, Kim B, Kwon SJ, Lee D. AnyBCQ: hardware efficient flexible binary-coded quantization for multi-precision LLMs. In: International Conference on Learning Representations. 2026.
Chen CT, Han D, Mun H, Hyun J, Raha A, Agarwal A, et al. HBQ: hierarchical scaling block quantization with hardware-efficiency-aware design for accurate LLM inference. arXiv. 2026;2609.00450.
Xiong J, Chen Q, Ye F, Wan Z, Zheng C, Zhao C, et al. ATTS: asynchronous test-time scaling via conformal prediction. In: International Conference on Learning Representations. 2026