Transformer-Based Language Models for Text Classification: From Attention to BERT, XLNet, RoBERTa, and Efficient Distillation
Keywords:
Natural language processing, BERT, attentionAbstract
Natural language processing experienced a major methodological transition between 2017 and 2020 as attention-based neural architectures and large-scale language-model pretraining progressively replaced task-specific recurrent and convolutional approaches. The Transformer architecture eliminated recurrent computation from sequence modeling and provided the foundation for a new generation of contextual language representations. This review examines the evolution of transformer-based language modeling and its implications for text classification through 2020. Particular attention is given to BERT and its bidirectional masked-language modeling framework, XLNet and generalized autoregressive pretraining, RoBERTa and optimized BERT training procedures, ALBERT and parameter-sharing strategies, ELECTRA and replaced-token detection, and T5's unified text-to-text framework. Earlier representation approaches including Word2Vec, GloVe, ELMo, sequence-to-sequence learning, and attention-based recurrent networks are discussed to establish the technical progression toward pretrained Transformers. Fine-tuning strategies, domain adaptation, transfer learning, computational requirements, input-length limitations, and model compression are evaluated. Efficient alternatives including DistilBERT and TinyBERT illustrate growing interest in deploying transformer models under restricted computational resources. The review concludes that pretrained contextual representations substantially reduced the requirement for task-specific architectures while introducing new challenges related to computational cost, model size, interpretability, domain adaptation, and sustainable deployment.
References
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.
Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. Stroudsburg: Association for Computational Linguistics; 2019. p. 4171-4186.
Radford A, Narasimhan K, Salimans T, Sutskever I. Improving language understanding by generative pre-training. OpenAI Technical Report. 2018.
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners. OpenAI Technical Report. 2019.
Yang Z, Dai Z, Yang Y, Carbonell J, Salakhutdinov RR, Le QV. XLNet: generalized autoregressive pretraining for language understanding. Adv Neural Inf Process Syst. 2019;32:5753-5763.
Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, et al. RoBERTa: a robustly optimized BERT pretraining approach. arXiv. 2019;1907.11692.
Lan Z, Chen M, Goodman S, Gimpel K, Sharma P, Soricut R. ALBERT: a lite BERT for self-supervised learning of language representations. In: International Conference on Learning Representations. 2020.
Clark K, Luong MT, Le QV, Manning CD. ELECTRA: pre-training text encoders as discriminators rather than generators. In: International Conference on Learning Representations. 2020.
Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J Mach Learn Res. 2020;21(140):1-67.
Peters ME, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, Zettlemoyer L. Deep contextualized word representations. In: Proceedings of NAACL-HLT. 2018. p. 2227-2237.
Howard J, Ruder S. Universal language model fine-tuning for text classification. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. 2018. p. 328-339.
Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. arXiv. 2013;1301.3781.
Pennington J, Socher R, Manning CD. GloVe: global vectors for word representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. 2014. p. 1532-1543.
Sutskever I, Vinyals O, Le QV. Sequence to sequence learning with neural networks. Adv Neural Inf Process Syst. 2014;27:3104-3112.
Bahdanau D, Cho K, Bengio Y. Neural machine translation by jointly learning to align and translate. In: International Conference on Learning Representations. 2015.
Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735-1780.
Kim Y. Convolutional neural networks for sentence classification. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. 2014. p. 1746-1751.
Wang A, Singh A, Michael J, Hill F, Levy O, Bowman SR. GLUE: a multi-task benchmark and analysis platform for natural language understanding. In: International Conference on Learning Representations. 2019.
Sun C, Qiu X, Xu Y, Huang X. How to fine-tune BERT for text classification? In: Sun M, Huang X, Ji H, Liu Z, Liu Y, editors. Chinese Computational Linguistics. Cham: Springer; 2019. p. 194-206.
Gururangan S, Marasović A, Swayamdipta S, Lo K, Beltagy I, Downey D, Smith NA. Don't stop pretraining: adapt language models to domains and tasks. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 8342-8360.
Sanh V, Debut L, Chaumond J, Wolf T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv. 2019;1910.01108.
Jiao X, Yin Y, Shang L, Jiang X, Chen X, Li L, et al. TinyBERT: distilling BERT for natural language understanding. In: Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. p. 4163-4174.