Abstract
Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.
Community
๐ก A Comprehensive Survey on Architectural-Level Memory in Large Language Models
Summary:
This survey from Tsinghua, NUS, and Bosch AI provides a unified theoretical framework for understanding architectural-level memory in LLMs, explicitly distinguishing it from external agent-based memory systems. The authors introduce a novel 3D taxonomy categorizing memory mechanisms by Representation (Implicit vs. Explicit), Update Dynamics (Offline vs. Online), and Persistence (Short-term vs. Long-term). It systematically maps the paradigm shift from computational byproducts like attention KV Caches and recurrent hidden states to explicitly addressable modules, such as Titans, TTT, and Engram. Highly recommended for researchers focusing on long-context scaling, hybrid architectures, and algorithm-hardware co-design.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents (2026)
- Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling (2026)
- Dynamic Linear Attention (2026)
- Lifelong In-Context Learning with Transformers Requires Parametric Forms of Attention (2026)
- A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets (2026)
- Are We Ready For An Agent-Native Memory System? (2026)
- Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.25380 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper