Large Language Models and Instruction-Following Artificial Intelligence: Scaling, Alignment, Reasoning, and Tool-Augmented Learning

Authors

  • Fatima Al-Khatib Author

Keywords:

LLM, AI, Instructions, deep learning

Abstract

Large language models underwent a substantial transformation between large-scale autoregressive pretraining and the emergence of instruction-following conversational systems. By 2023, scaling model parameters, training data, and computational resources had produced increasingly general language models capable of in-context learning, reasoning, code generation, question answering, and instruction following. This review examines major developments underlying large language models through 2023. Transformer architectures and autoregressive language modeling are first considered as the technological foundation for model scaling. GPT-3 demonstrated strong few-shot and in-context learning behavior without conventional task-specific fine-tuning, while subsequent scaling-law research investigated relationships among model size, dataset size, training compute, and performance. Chinchilla scaling emphasized the importance of balancing model parameters with training-token volume. Instruction tuning subsequently improved generalization to previously unseen tasks, while reinforcement learning from human feedback provided a mechanism for aligning model outputs with human preferences. Chain-of-thought prompting, self-consistency, zero-shot reasoning, and ReAct demonstrated that prompting strategies could significantly influence reasoning behavior without modifying model parameters. Toolformer further explored self-supervised integration of external tools. The emergence of LLaMA and GPT-4 in 2023 demonstrated continuing improvements in general-purpose language understanding and generation. This review analyzes scaling, instruction tuning, human-feedback alignment, in-context learning, reasoning, tool use, evaluation, and safety. Remaining challenges included hallucination, bias, unreliable reasoning, prompt sensitivity, opaque training data, computational requirements, alignment uncertainty, and inadequate evaluation of rapidly emerging capabilities.

References

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.

Radford A, Narasimhan K, Salimans T, Sutskever I. Improving language understanding by generative pre-training. OpenAI Technical Report. 2018.

Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners. OpenAI Technical Report. 2019.

Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. Adv Neural Inf Process Syst. 2020;33:1877-1901.

Kaplan J, McCandlish S, Henighan T, Brown TB, Chess B, Child R, et al. Scaling laws for neural language models. arXiv. 2020;2001.08361.

Hoffmann J, Borgeaud S, Mensch A, Buchatskaya E, Cai T, Rutherford E, et al. Training compute-optimal large language models. Adv Neural Inf Process Syst. 2022;35:30016-30030.

Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, et al. Training language models to follow instructions with human feedback. Adv Neural Inf Process Syst. 2022;35:27730-27744.

Wei J, Bosma M, Zhao VY, Guu K, Yu AW, Lester B, et al. Finetuned language models are zero-shot learners. In: International Conference on Learning Representations. 2022.

Chung HW, Hou L, Longpre S, Zoph B, Tay Y, Fedus W, et al. Scaling instruction-finetuned language models. arXiv. 2022;2210.11416.

Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, et al. Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst. 2022;35:24824-24837.

Kojima T, Gu SS, Reid M, Matsuo Y, Iwasawa Y. Large language models are zero-shot reasoners. Adv Neural Inf Process Syst. 2022;35:22199-22213.

Wang X, Wei J, Schuurmans D, Le Q, Chi E, Narang S, et al. Self-consistency improves chain of thought reasoning in language models. In: International Conference on Learning Representations. 2023.

Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, Cao Y. ReAct: synergizing reasoning and acting in language models. In: International Conference on Learning Representations. 2023.

Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, Lomeli M, Hambro E, et al. Toolformer: language models can teach themselves to use tools. Adv Neural Inf Process Syst. 2023;36.

Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, et al. LLaMA: open and efficient foundation language models. arXiv. 2023;2302.13971.

Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. Llama 2: open foundation and fine-tuned chat models. arXiv. 2023;2307.09288.

OpenAI. GPT-4 technical report. arXiv. 2023;2303.08774.

Bai Y, Kadavath S, Kundu S, Askell A, Kernion J, Jones A, et al. Constitutional AI: harmlessness from AI feedback. arXiv. 2022;2212.08073.

Wei J, Tay Y, Bommasani R, Raffel C, Zoph B, Borgeaud S, et al. Emergent abilities of large language models. Trans Mach Learn Res. 2022.

Liang P, Bommasani R, Lee T, Tsipras D, Soylu D, Yasunaga M, et al. Holistic evaluation of language models. Trans Mach Learn Res. 2023.

Published

2023-06-01