Dense Information Retrieval and Retrieval-Augmented Language Models: Architectures for Grounded Knowledge-Intensive Artificial Intelligence

Authors

  • Jonathan Reed Author

Keywords:

question answering, knowledge grounding, semantic search, language models

Abstract

Although large pretrained language models encode substantial information within their parameters, purely parametric knowledge introduces limitations related to factual accuracy, knowledge freshness, provenance, and efficient updating. Retrieval-augmented architectures address these limitations by combining neural language models with external document collections. This review examines the evolution of dense retrieval and retrieval-augmented language modeling through 2023. Traditional sparse retrieval using term-frequency representations is considered as a foundation before examining neural semantic retrieval. Dense Passage Retrieval demonstrated that independently encoded queries and passages could support effective open-domain question answering through vector similarity. REALM incorporated retrieval into language-model pretraining, while Retrieval-Augmented Generation combined a pretrained sequence generator with retrieved documents treated as latent evidence. ColBERT introduced late interaction to preserve token-level matching while retaining retrieval efficiency. Fusion-in-Decoder improved generation over multiple retrieved passages, whereas RETRO demonstrated that retrieval could be integrated into very large language-model pretraining. Atlas further explored few-shot learning using retrieval-augmented language models. In 2023, Hypothetical Document Embeddings demonstrated zero-shot dense retrieval through language-model-generated pseudo-documents, while retrieval increasingly became central to grounding instruction-following systems. This review examines indexing, embedding models, vector similarity search, retriever training, reranking, document chunking, context fusion, factual grounding, retrieval latency, and knowledge updating. Challenges include retrieval errors, evidence conflicts, hallucinated synthesis, embedding-domain mismatch, stale indexes, evaluation methodology, and determining when retrieval should override parametric knowledge.

References

Robertson SE, Zaragoza H. The probabilistic relevance framework: BM25 and beyond. Found Trends Inf Retr. 2009;3(4):333-389.

Johnson J, Douze M, Jégou H. Billion-scale similarity search with GPUs. IEEE Trans Big Data. 2019;7(3):535-547.

Lee K, Chang MW, Toutanova K. Latent retrieval for weakly supervised open domain question answering. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. p. 6086-6096.

Karpukhin V, Oguz B, Min S, Lewis P, Wu L, Edunov S, et al. Dense passage retrieval for open-domain question answering. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. 2020. p. 6769-6781.

Guu K, Lee K, Tung Z, Pasupat P, Chang MW. REALM: retrieval-augmented language model pre-training. Proc Mach Learn Res. 2020;119:3929-3938.

Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv Neural Inf Process Syst. 2020;33:9459-9474.

Khattab O, Zaharia M. ColBERT: efficient and effective passage search via contextualized late interaction over BERT. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 2020. p. 39-48.

Izacard G, Grave E. Leveraging passage retrieval with generative models for open domain question answering. In: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics. 2021. p. 874-880.

Xiong L, Xiong C, Li Y, Tang KW, Liu J, Bennett PN, et al. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In: International Conference on Learning Representations. 2021.

Thakur N, Reimers N, Rücklé A, Srivastava A, Gurevych I. BEIR: a heterogeneous benchmark for zero-shot evaluation of information retrieval models. Adv Neural Inf Process Syst. 2021;34:14086-14100.

Gao T, Yao X, Chen D. SimCSE: simple contrastive learning of sentence embeddings. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. p. 6894-6910.

Borgeaud S, Mensch A, Hoffmann J, Cai T, Rutherford E, Millican K, et al. Improving language models by retrieving from trillions of tokens. Proc Mach Learn Res. 2022;162:2206-2240.

Santhanam K, Khattab O, Saad-Falcon J, Potts C, Zaharia M. ColBERTv2: effective and efficient retrieval via lightweight late interaction. In: Proceedings of NAACL-HLT. 2022. p. 3715-3734.

Ni J, Qu C, Lu J, Dai Z, Hernandez Abrego G, Ma J, et al. Large dual encoders are generalizable retrievers. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. p. 9844-9855.

Izacard G, Lewis P, Lomeli M, Hosseini L, Petroni F, Schick T, et al. Atlas: few-shot learning with retrieval augmented language models. J Mach Learn Res. 2023;24(251):1-43.

Gao L, Ma X, Lin J, Callan J. Precise zero-shot dense retrieval without relevance labels. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. 2023. p. 1762-1777.

Muennighoff N, Tazi N, Magne L, Reimers N. MTEB: massive text embedding benchmark. In: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. p. 2014-2037.

Ram O, Levine Y, Dalmedigos I, Muhlgay D, Shashua A, Leyton-Brown K, Shoham Y. In-context retrieval-augmented language models. Trans Assoc Comput Linguist. 2023;11:1316-1331.

Nogueira R, Cho K. Passage re-ranking with BERT. arXiv. 2019;1901.04085.

Reimers N, Gurevych I. Sentence-BERT: sentence embeddings using Siamese BERT-networks. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. 2019. p. 3982-3992

Published

2023-06-01