Papers
arxiv:2609.39661

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Published on Sep 30
· Submitted by
zhentao tan
on Oct 1
Authors:
,
,
,
,

Abstract

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model. We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.

Community

Paper author Paper submitter

In this survey, we examine efficient sequence architectures through the concept of model-internal contextual memory. We introduce a five-dimensional analytical lens—Memory Representation, Memory Update, Access, Readout, and Integration—to connect research on memory compression, sparse attention, recurrent states, structured state dynamics, and hybrid architectures. Drawing on 59 release-level records across 14 major model lineages and a comparison of 11 high-performing open-weight models, we show how memory processing is increasingly coordinated not only across time, but also across network depth. We hope this framework provides a common language for comparing existing architectures and inspires future designs based on stateful, multidimensional, and selectively routed memory.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.39661
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39661 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39661 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39661 in a Space README.md to link it from this page.

Collections including this paper 1