Papers
arxiv:2609.40195

MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories

Published on Sep 30
· Submitted by
Guangzhi Xiong
on Oct 1
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.

Community

Paper author Paper submitter

We introduced MemLife, an agentic memory system for long-term egocentric video that constructs time- and entity-anchored, first-person episodes and accesses them through time-scoped retrieval and chronological evidence organization. We further proposed MemOpt, which applies reinforcement learning only to the memory writer through the FIRM objective for faithful, informative, and retrievable memories. MemLife outperforms state-of-the-art training-free systems, while MemOpt provides further gains that transfer across writer and reader backbones, memory systems, and out-of-domain video and question distributions. Together, the results demonstrate that learning what to remember is an effective and generalizable approach to long-term video question answering.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.40195
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.40195 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.40195 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.40195 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.