Adaptive and Self-Reflective Retrieval-Augmented Generation for Grounded Large Language Models

Authors

  • Ibrahim Khalil Author

Keywords:

Adaptive-RAG, RAG, Self-RAG, RAPTOR

Abstract

Retrieval-Augmented Generation became an important architecture for reducing exclusive reliance on the parametric knowledge of large language models. Conventional RAG systems typically retrieve a fixed number of passages for every query and pass them to a generator regardless of whether retrieval is necessary, whether the retrieved evidence is relevant, or whether conflicting documents are present. Research through 2024 increasingly addressed these limitations through adaptive retrieval, hierarchical document representation, retrieval-aware fine-tuning, corrective retrieval, and self-reflection. This review traces the development from sparse information retrieval and dense passage retrieval to REALM, conventional RAG, Fusion-in-Decoder, RETRO, Atlas, Self-RAG, RAPTOR, Adaptive-RAG, Corrective RAG, and Retrieval-Augmented Fine-Tuning. Self-RAG integrates special reflection tokens that allow a language model to determine when retrieval is appropriate and to critique retrieved evidence and generated content. RAPTOR recursively clusters and summarizes documents to permit retrieval at multiple levels of abstraction. Adaptive-RAG varies retrieval strategy according to estimated query complexity, while Corrective RAG evaluates the quality of retrieved evidence before generation. RAFT trains models to identify relevant evidence and disregard distractor documents in domain-specific settings. This review compares retrieval triggering, retriever-generator integration, hierarchical indexing, evidence selection, domain adaptation, citation, factuality, latency, and evaluation. Persistent limitations include retrieval failure, stale corpora, contradictory evidence, citation inaccuracies, context overload, domain mismatch, and hallucinated synthesis despite access to relevant evidence.

References

Robertson SE, Zaragoza H. The probabilistic relevance framework: BM25 and beyond. Found Trends Inf Retr. 2009;3(4):333-389.

Reimers N, Gurevych I. Sentence-BERT: sentence embeddings using Siamese BERT-networks. In: Proceedings of EMNLP-IJCNLP. 2019. p. 3982-3992.

Karpukhin V, Oguz B, Min S, Lewis P, Wu L, Edunov S, et al. Dense passage retrieval for open-domain question answering. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. 2020. p. 6769-6781.

Guu K, Lee K, Tung Z, Pasupat P, Chang MW. REALM: retrieval-augmented language model pre-training. Proc Mach Learn Res. 2020;119:3929-3938.

Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv Neural Inf Process Syst. 2020;33:9459-9474.

Khattab O, Zaharia M. ColBERT: efficient and effective passage search via contextualized late interaction over BERT. In: Proceedings of the 43rd International ACM SIGIR Conference. 2020. p. 39-48.

Izacard G, Grave E. Leveraging passage retrieval with generative models for open domain question answering. In: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics. 2021. p. 874-880.

Thakur N, Reimers N, Rücklé A, Srivastava A, Gurevych I. BEIR: a heterogeneous benchmark for zero-shot evaluation of information retrieval models. Adv Neural Inf Process Syst. 2021;34:14086-14100.

Borgeaud S, Mensch A, Hoffmann J, Cai T, Rutherford E, Millican K, et al. Improving language models by retrieving from trillions of tokens. Proc Mach Learn Res. 2022;162:2206-2240.

Izacard G, Lewis P, Lomeli M, Hosseini L, Petroni F, Schick T, et al. Atlas: few-shot learning with retrieval augmented language models. J Mach Learn Res. 2023;24(251):1-43.

Gao L, Ma X, Lin J, Callan J. Precise zero-shot dense retrieval without relevance labels. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. 2023. p. 1762-1777.

Muennighoff N, Tazi N, Magne L, Reimers N. MTEB: massive text embedding benchmark. In: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. p. 2014-2037.

Asai A, Wu Z, Wang Y, Sil A, Hajishirzi H. Self-RAG: learning to retrieve, generate, and critique through self-reflection. In: International Conference on Learning Representations. 2024.

Sarthi P, Abdullah S, Tuli A, Khanna S, Goldie A, Manning CD. RAPTOR: recursive abstractive processing for tree-organized retrieval. In: International Conference on Learning Representations. 2024.

Jeong S, Baek J, Cho S, Hwang SJ, Park J. Adaptive-RAG: learning to adapt retrieval-augmented large language models through question complexity. In: Proceedings of NAACL-HLT. 2024. p. 7036-7050.

Yan SQ, Gu JC, Zhu Y, Ling ZH. Corrective retrieval augmented generation. arXiv. 2024;2401.15884.

Zhang T, Patil SG, Jain N, Shen S, Zaharia M, Stoica I, Gonzalez JE. RAFT: adapting language model to domain specific RAG. arXiv. 2024;2403.10131.

Jiang H, Wu Q, Luo X, Li D, Lin CY, Yang Y, Qiu L. LongLLMLingua: accelerating and enhancing LLMs in long context scenarios via prompt compression. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024. p. 1658-1677.

Gao Y, Xiong Y, Gao X, Jia K, Pan J, Bi Y, et al. Retrieval-augmented generation for large language models: a survey. arXiv. 2023;2312.10997.

Ram O, Levine Y, Dalmedigos I, Muhlgay D, Shashua A, Leyton-Brown K, Shoham Y. In-context retrieval-augmented language models. Trans Assoc Comput Linguist. 2023;11:1316-1331.

Published

2024-06-01