Persistent Memory and Experience Learning for Autonomous Language Agents: Retrieval, Consolidation, Reflection, and Selective Forgetting

Authors

  • Omar Rahman Author

Keywords:

long-term memory, autonomous agents, Agent memory

Abstract

Long-lived autonomous agents require mechanisms for retaining and reusing information that extends beyond a single model context window. While conventional retrieval-augmented generation primarily supplies external factual documents to a language model, agent memory introduces additional requirements involving experience accumulation, temporal ordering, personalization, reflection, contradiction handling, selective retention, and learning from past actions. This review examines persistent memory for language-model agents through September 2026. Early retrieval-based architectures, episodic simulations, reflective agents, and external-memory systems established mechanisms for storing interaction histories and reintroducing them into later reasoning. Generative Agents demonstrated the use of memory streams, retrieval, reflection, and planning in persistent simulated environments. Reflexion used linguistic self-feedback as episodic memory for iterative task improvement, while Voyager accumulated reusable skills during open-ended embodied interaction. MemGPT conceptualized external memory management through hierarchical movement between limited context and persistent storage. By 2025–2026, research increasingly shifted from demonstrating that memory could help agents toward evaluating what effective memory actually requires. MemoryAgentBench emphasized accurate retrieval, test-time learning, long-range understanding, and selective forgetting. AMemGym introduced interactive on-policy evaluation for memory-driven assistants. Retrieval-Augmented LLM Agents combined training with retrieved trajectories to improve transfer to unseen tasks. Agent Learning via Early Experience investigated learning from environment transitions and self-reflection without relying entirely on conventional reinforcement learning. DreamGym scaled agent learning using synthesized experience and simulated environment dynamics. The review analyzes memory writing, retrieval policy, reflection, consolidation, forgetting, provenance, personalization, and privacy. Persistent memory is emerging as a core architectural layer separating short-lived language interaction from adaptive autonomous behavior.

References

Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv Neural Inf Process Syst. 2020;33:9459-9474.

Karpukhin V, Oguz B, Min S, Lewis P, Wu L, Edunov S, et al. Dense passage retrieval for open-domain question answering. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. 2020. p. 6769-6781.

Park JS, O'Brien JC, Cai CJ, Morris MR, Liang P, Bernstein MS. Generative agents: interactive simulacra of human behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 2023.

Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S. Reflexion: language agents with verbal reinforcement learning. Adv Neural Inf Process Syst. 2023;36.

Wang G, Xie Y, Jiang Y, Mandlekar A, Xiao C, Zhu Y, et al. Voyager: an open-ended embodied agent with large language models. Trans Mach Learn Res. 2024.

Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, Cao Y. ReAct: synergizing reasoning and acting in language models. In: International Conference on Learning Representations. 2023.

Packer C, Fang V, Patil SG, Lin K, Wooders S, Gonzalez JE. MemGPT: towards LLMs as operating systems. arXiv. 2023;2310.08560.

Zhong W, Guo L, Gao Q, Ye H, Wang Y. MemoryBank: enhancing large language models with long-term memory. arXiv. 2023;2305.10250.

Zhao A, Huang D, Xu Q, Lin M, Liu YJ, Huang G. ExpeL: LLM agents are experiential learners. arXiv. 2023;2308.10144.

Wang L, Ma C, Feng X, Zhang Z, Yang H, Zhang J, et al. A survey on large language model based autonomous agents. Front Comput Sci. 2024;18:186345.

Hu Y, Wang Y, McAuley J. Evaluating memory in LLM agents via incremental multi-turn interactions. In: International Conference on Learning Representations. 2026.

Cheng J, Ru D, Qiu L, Li Y, Cao X, Song Y, Cai X. AMemGym: interactive memory benchmarking for assistants in long-horizon conversations. In: International Conference on Learning Representations. 2026.

Du P. Memory for autonomous LLM agents: mechanisms, evaluation, and emerging frontiers. arXiv. 2026;2603.07670.

Ferraz TP, Deffayet R, Nikoulina V, Déjean H, Clinchant S. Retrieval-augmented LLM agents: learning to learn from experience. arXiv. 2026;2603.18272.

Zhang K, Chen X, Liu B, Xue T, Liao Z, Liu Z, et al. Agent learning via early experience. In: International Conference on Machine Learning. 2026.

Chen Z, Zhao Z, Zhang K, Liu B, Qi Q, Wu Y, et al. Scaling agent learning via experience synthesis. In: International Conference on Learning Representations. 2026.

Bursa O. A dynamic retrieval-augmented generation system with selective memory and remembrance. arXiv. 2026;2601.02428.

Wang R, Chen Y, Wang Y, Wu C, Fang J, Cai X, et al. AgentNoiseBench: benchmarking robustness of tool-using LLM agents under noisy condition. arXiv. 2026;2602.11348.

Asai A, Wu Z, Wang Y, Sil A, Hajishirzi H. Self-RAG: learning to retrieve, generate, and critique through self-reflection. In: International Conference on Learning Representations. 2024.

Sarthi P, Abdullah S, Tuli A, Khanna S, Goldie A, Manning CD. RAPTOR: recursive abstractive processing for tree-organized retrieval. In: International Conference on Learning Representations. 2024.

Published

2026-06-01