Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 16 days ago • 84
Bernini: Latent Semantic Planning for Video Diffusion Paper • 2605.22344 • Published May 21 • 20
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives Paper • 2605.12496 • Published May 12 • 31
WRBench: Current World Models Lack a Persistent State Core Collection WRBench public release: paper, prompts, videos, scores, human labels, and leaderboard. • 6 items • Updated 21 days ago • 3
Memento: Reconstruct to Remember for Consistent Long Video Generation Paper • 2606.14667 • Published Jun 12 • 18
VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing Paper • 2605.30117 • Published May 28 • 2
Pelican-Unified 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action Paper • 2605.15153 • Published May 14 • 1
VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing Paper • 2605.30117 • Published May 28 • 2