Memento-8B

Built with Llama.

Memento augments Llama-3.1-8B-Instruct with Dynamic Memory, Query-related Memory Selection and a learned visual connector for proactive responses to long video streams. The checkpoint contains the LoRA adapter and trained visual and memory modules; the base Llama and SigLIP weights are downloaded separately.

Model sources

Model details

  • Language model: Llama-3.1-8B-Instruct
  • Vision encoder: SigLIP-Large-Patch16-384
  • Video input: 2 FPS, with one global and nine spatial tokens per frame
  • Training data: Memento-54K

Use

Install the linked Memento repository and its pinned dependencies, then download the assets:

python scripts/download_assets.py --kind model --output checkpoints/Memento-8B
python scripts/download_assets.py --kind base --output checkpoints/Llama-3.1-8B-Instruct
python scripts/download_assets.py --kind vision --output checkpoints/siglip-large-patch16-384

Launch the interactive video demo:

python -m pip install -r requirements-demo.txt
CUDA_VISIBLE_DEVICES=0 python -m demo.server

Open http://127.0.0.1:7860/ to upload a video and interact with the model. See the demo guide and training and evaluation guide for complete instructions. Inference requires Memento's custom streaming engine.

Training and evaluation

The model is fine-tuned for one epoch on Memento-54K using LoRA (rank 128, alpha 256), a learning rate of 1e-4 and a 32,768-token context. See training_config.json for the full configuration and the code repository for training and evaluation instructions.

Limitations

The model may miss events, respond at the wrong time or generate inaccurate descriptions. Performance depends on the video domain and query.

License

The model is subject to the Llama 3.1 Community License and Acceptable Use Policy. Those terms include conditions on commercial use, attribution and redistribution. Code and annotation licenses are separate and do not override the model's upstream terms.

Citation

@inproceedings{memento2026iclr,
  title = {Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video},
  booktitle = {ICLR},
  year = {2026},
  url = {https://openreview.net/forum?id=FtdbdoGbk3}
}
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for liarzone/Memento-8B

Adapter
(2918)
this model

Dataset used to train liarzone/Memento-8B

Space using liarzone/Memento-8B 1