KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
Abstract
Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose KeyRec, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.
Community
Excited to share KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding!
KeyRec is a training-free framework that organizes video streams into a dense recent cache and a bounded key-event memory, with query-adaptive fixed-budget readout. Across four benchmarks and three VLM backbones, it achieves the best compressed performance in 13/15 settings while using only 10% of the dense decoder-facing visual-token budget.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- StreamFlow: Dynamic Memory Flows for Streaming Video Understanding (2026)
- Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding (2026)
- CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding (2026)
- ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding (2026)
- PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding (2026)
- FlashBack: Knowing When to Remember in Streaming Vision-Language Models (2026)
- SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.32182 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper