Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision Paper • 2608.16812 • Published Aug 17 • 50
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal Paper • 2608.05565 • Published Aug 6 • 23
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published Jul 30 • 62
Tooony133/dinov3-vits16-pretrain-lvd1689m Image Feature Extraction • 21.6M • Updated Jun 19 • 2.86k • 2
TimeLens2 Collection Generalist Video Temporal Grounding with Multimodal LLMs • 8 items • Updated Jul 28 • 15
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 120
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Paper • 2604.24764 • Published Apr 27 • 121