Instructions to use liarzone/Memento-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use liarzone/Memento-8B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct") model = PeftModel.from_pretrained(base_model, "liarzone/Memento-8B") - Notebooks
- Google Colab
- Kaggle
Memento-8B
Built with Llama.
Memento augments Llama-3.1-8B-Instruct with Dynamic Memory, Query-related Memory Selection and a learned visual connector for proactive responses to long video streams. The checkpoint contains the LoRA adapter and trained visual and memory modules; the base Llama and SigLIP weights are downloaded separately.
Model sources
Model details
- Language model: Llama-3.1-8B-Instruct
- Vision encoder: SigLIP-Large-Patch16-384
- Video input: 2 FPS, with one global and nine spatial tokens per frame
- Training data: Memento-54K
Use
Install the linked Memento repository and its pinned dependencies, then download the assets:
python scripts/download_assets.py --kind model --output checkpoints/Memento-8B
python scripts/download_assets.py --kind base --output checkpoints/Llama-3.1-8B-Instruct
python scripts/download_assets.py --kind vision --output checkpoints/siglip-large-patch16-384
Launch the interactive video demo:
python -m pip install -r requirements-demo.txt
CUDA_VISIBLE_DEVICES=0 python -m demo.server
Open http://127.0.0.1:7860/ to upload a video and interact with the model. See the demo guide and training and evaluation guide for complete instructions. Inference requires Memento's custom streaming engine.
Training and evaluation
The model is fine-tuned for one epoch on Memento-54K using LoRA (rank 128, alpha 256), a learning rate of 1e-4 and a 32,768-token context. See training_config.json for the full configuration and the code repository for training and evaluation instructions.
Limitations
The model may miss events, respond at the wrong time or generate inaccurate descriptions. Performance depends on the video domain and query.
License
The model is subject to the Llama 3.1 Community License and Acceptable Use Policy. Those terms include conditions on commercial use, attribution and redistribution. Code and annotation licenses are separate and do not override the model's upstream terms.
Citation
@inproceedings{memento2026iclr,
title = {Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video},
booktitle = {ICLR},
year = {2026},
url = {https://openreview.net/forum?id=FtdbdoGbk3}
}
- Downloads last month
- 17
Model tree for liarzone/Memento-8B
Base model
meta-llama/Llama-3.1-8B