PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders Paper • 2603.25398 • Published Mar 26 • 3
LVMT: Video Mask Transformer for Long-term Video Segmentation Paper • 2609.34895 • Published 7 days ago • 3
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens Paper • 2604.04913 • Published Apr 6 • 10
view article Article How I contributed a new model to the Transformers library using Codex nielsr • Mar 30 • 53
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model Paper • 2602.17807 • Published Feb 19 • 7
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think Paper • 2409.11355 • Published Sep 17, 2024 • 31