CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory
Abstract
Linear RNNs offer linear-time sequence processing and constant-memory decoding, but their fixed-size recurrent states must accommodate all past key--value associations. Existing forgetting mechanisms and Delta Rule updates reduce interference by selectively clearing or correcting the state, yet earlier associations can still become difficult to retrieve. We introduce CyFA (Cyclic Flow Attention), a Linear RNN with relative-time-partitioned memory. At each step, a learned clock controls the cyclic transport applied jointly to the key and value states before the current key--value pair enters the age-zero slot, thereby organizing stored associations across relative-time slots. We further derive an exact change to absolute-clock coordinates that expresses CyFA as two scalar-decay linear attention recurrences and enables efficient chunk-wise training. Across 400M--1.4B pretraining experiments with matched recurrent-state sizes, CyFA improves recall-intensive performance while maintaining competitive language modeling and high computational efficiency. At 400M, CyFA outperforms KDA on FDA (42.60 vs. 26.07) while requiring only 46.7% and 48.3% of KDA's forward and backward core-operator execution times, respectively. Our code is publicly available at https://github.com/Chyxx/CyclicFlowAttention{this https URL}.
Get this paper in your agent:
hf papers read 2609.36259 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 3
cyxxxxxxxxxx/cyfa-800M-30B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper