|
Download README.md from alisalam/LookWhen: direct link, hf CLI and curl.
- Browser
- Download file 1.82 kB
-
https://huggingface.co/alisalam/LookWhen/resolve/main/README.md
- Command line
-
hf download hf://alisalam/LookWhen/README.md
-
curl -L -o README.md https://huggingface.co/alisalam/LookWhen/resolve/main/README.md
1.82 kB
| license: mit | |
| pipeline_tag: video-classification | |
| tags: | |
| - video | |
| - action-recognition | |
| - efficient-inference | |
| - token-selection | |
| library_name: pytorch | |
| # LookWhen ViT-B/16 | |
| Pre-trained checkpoint for **"LookWhen? Fast Video Recognition by Learning When, Where, | |
| and What to Compute"** ([arXiv:2605.06809](https://arxiv.org/abs/2605.06809)). | |
| LookWhen splits video recognition in two: a shallow **selector** scores every patch of a | |
| 2x-downscaled (2x in each dimension) clip, and a deep **extractor** processes only the top-K of them while still | |
| predicting features of the whole video. | |
| This is the model used throughout the paper: ViT-B/16, pre-trained for 20 epochs on | |
| Kinetics-400 and SSv2 by distilling InternVideo2's video token and DINOv3's frame and | |
| patch tokens, with selection supervised by top1-distance token uniqueness. It was trained | |
| over sparsity 70-95%, so it runs anywhere in that range without retraining. | |
| ## Usage | |
| ```bash | |
| pip install git+https://github.com/alisalamatian1/LookWhen.git | |
| ``` | |
| ```python | |
| import torch | |
| from lookwhen import LookWhen | |
| from lookwhen.data.video import load_clip | |
| model = LookWhen.from_pretrained().cuda().eval() # downloads this checkpoint | |
| clip = load_clip("video.mp4", num_frames=16, img_size=224).unsqueeze(0).cuda() | |
| out = model(clip, keep_ratio=0.1) # keep_ratio = 1 - sparsity | |
| out["video"] # (1, 768) video embedding | |
| ``` | |
| Code, training, and evaluation: https://github.com/alisalamatian1/LookWhen | |
| ## Citation | |
| ```bibtex | |
| @article{salamatian2026lookwhen, | |
| title = {LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute}, | |
| author = {Salamatian, Ali and Fuller, Anthony and Sarkar, Pritam and | |
| Green, James R. and Sigal, Leonid and Shelhamer, Evan}, | |
| journal = {arXiv preprint arXiv:2605.06809}, | |
| year = {2026} | |
| } | |
| ``` | |