YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Model weights for RVAS: Referring Video Active Exploration and Segmentation (ICML 2026).
LESA (Language Environment-aware Segmentation Assistant) performs action planning and referring video segmentation. A lightweight controller invokes the multimodal language model as frames arrive. The generated [SEG] embedding conditions a shared SAM2 model for segmentation and tracking, with backward mask revision when new evidence becomes available.
Usage
Follow the installation instructions, then download the model from the repository root:
from huggingface_hub import snapshot_download
snapshot_download('FudanCVL/LESA', local_dir='weights')
The package contains four safetensors shards, model configuration, and tokenizer files. LoRA parameters are merged, and SAM2 is included. Model classes are provided by the source repository.
Run inference on a directory of ordered RGB frames:
CUDA_VISIBLE_DEVICES=0 python -m lesa.infer \
--frames data/example/frames \
--expression 'The person wearing a blue shirt.' \
--video-id example --exp-id 1 \
--output outputs/example
The entry point uses the default command-line settings and loads the checkpoint in weights/. See the source repository for RVAS data preparation, training, benchmark inference, and evaluation.
Citation
@inproceedings{hu2026rvas,
title = {RVAS: Referring Video Active Exploration and Segmentation},
author = {Hu, Hengrui and Gao, Weiwei and Zhang, Zipei and Ding, Henghui},
booktitle = {Proceedings of the International Conference on Machine Learning},
year = {2026}
}
License
See LICENSE and the source repository's acknowledgements. Upstream components and pretrained weights remain subject to their respective licenses and terms.
- Downloads last month
- 7