Parameter-Efficient Adaptation and Quantization of Large Language Models: From Adapters and Prompt Tuning to LoRA and QLoRA

Authors

  • Martin Vogel Author

Keywords:

efficient AI, model compression, large language models

Abstract

The rapid increase in the parameter count of pretrained language models created significant barriers to conventional task-specific fine-tuning. Updating every parameter of a multi-billion-parameter model requires substantial accelerator memory, storage, training time, and energy, while maintaining separate full model copies for multiple downstream tasks further increases deployment cost. Parameter-efficient fine-tuning seeks to adapt large pretrained models while modifying only a small fraction of their parameters. This review examines the evolution of efficient model adaptation through 2023. Adapter layers introduced compact trainable modules between frozen network layers, while BitFit demonstrated that useful adaptation could sometimes be achieved by updating only bias terms. Prefix tuning and prompt tuning shifted adaptation toward trainable continuous vectors inserted into the model's input or attention context. Low-Rank Adaptation introduced trainable low-rank decompositions of weight updates and became an increasingly influential approach for large language models. AdaLoRA dynamically allocated rank budgets according to parameter importance, while unified frameworks attempted to characterize common principles across multiple parameter-efficient methods. Alongside adaptation techniques, post-training and training-aware quantization reduced model memory requirements. LLM.int8() enabled large Transformer inference using mixed-precision decomposition, GPTQ provided accurate post-training weight quantization, and SmoothQuant shifted quantization difficulty from activations toward weights. QLoRA subsequently combined 4-bit quantized pretrained models with LoRA adapters, enabling efficient fine-tuning of very large language models on substantially reduced hardware. The review examines memory consumption, trainable parameter count, quantization error, downstream performance, optimizer requirements, inference overhead, and deployment trade-offs.

References

Houlsby N, Giurgiu A, Jastrzebski S, Morrone B, de Laroussilhe Q, Gesmundo A, et al. Parameter-efficient transfer learning for NLP. Proc Mach Learn Res. 2019;97:2790-2799.

Pfeiffer J, Vulić I, Gurevych I, Ruder S. MAD-X: an adapter-based framework for multi-task cross-lingual transfer. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. 2020. p. 7654-7673.

Pfeiffer J, Rücklé A, Poth C, Kamath A, Vulić I, Ruder S, Gurevych I. AdapterHub: a framework for adapting Transformers. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 2020. p. 46-54.

Li XL, Liang P. Prefix-tuning: optimizing continuous prompts for generation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. 2021. p. 4582-4597.

Lester B, Al-Rfou R, Constant N. The power of scale for parameter-efficient prompt tuning. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. p. 3045-3059.

Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, et al. LoRA: low-rank adaptation of large language models. In: International Conference on Learning Representations. 2022.

Ben Zaken E, Goldberg Y, Ravfogel S. BitFit: simple parameter-efficient fine-tuning for Transformer-based masked language-models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 2022. p. 1-9.

Liu X, Ji K, Fu Y, Tam W, Du Z, Yang Z, Tang J. P-Tuning v2: prompt tuning can be comparable to fine-tuning universally across scales and tasks. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 2022. p. 61-68.

He J, Zhou C, Ma X, Berg-Kirkpatrick T, Neubig G. Towards a unified view of parameter-efficient transfer learning. In: International Conference on Learning Representations. 2022.

Mao Y, Mathias L, Hou R, Almahairi A, Ma H, Han J, et al. UniPELT: a unified framework for parameter-efficient language model tuning. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 2022. p. 6253-6264.

Dettmers T, Lewis M, Belkada Y, Zettlemoyer L. LLM.int8(): 8-bit matrix multiplication for Transformers at scale. Adv Neural Inf Process Syst. 2022;35:30318-30332.

Dettmers T, Zettlemoyer L. The case for 4-bit precision: k-bit inference scaling laws. Proc Mach Learn Res. 2023;202:7750-7774.

Frantar E, Ashkboos S, Hoefler T, Alistarh D. GPTQ: accurate post-training quantization for generative pre-trained Transformers. In: International Conference on Learning Representations. 2023.

Xiao G, Lin J, Seznec M, Wu H, Demouth J, Han S. SmoothQuant: accurate and efficient post-training quantization for large language models. Proc Mach Learn Res. 2023;202:38087-38099.

Zhang Q, Chen M, Bukharin A, Karampatziakis N, He P, Cheng Y, et al. AdaLoRA: adaptive budget allocation for parameter-efficient fine-tuning. In: International Conference on Learning Representations. 2023.

Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L. QLoRA: efficient finetuning of quantized LLMs. Adv Neural Inf Process Syst. 2023;36.

Malladi S, Gao T, Nichani E, Damian A, Lee J, Chen D, Arora S. Fine-tuning language models with just forward passes. Adv Neural Inf Process Syst. 2023;36.

Zhang R, Han J, Liu C, Zhou A, Lu P, Li H, et al. LLaMA-Adapter: efficient fine-tuning of language models with zero-init attention. arXiv. 2023;2303.16199.

Liu H, Tam D, Muqeeth M, Mohta J, Huang T, Bansal M, Raffel C. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Adv Neural Inf Process Syst. 2022;35:1950-1965.

Dettmers T, Lewis M, Shleifer S, Zettlemoyer L. 8-bit optimizers via block-wise quantization. In: International Conference on Learning Representations. 2022.

Published

2023-06-01