Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 11 days ago • 41
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations Paper • 2609.20817 • Published 11 days ago • 40
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published 21 days ago • 146
Generative Late-Interaction Embeddings For Visual Document Retrieval Paper • 2609.11808 • Published 18 days ago • 27
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? Paper • 2609.05324 • Published 24 days ago • 28
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model Paper • 2609.09158 • Published 20 days ago • 22
FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow Paper • 2609.03563 • Published 25 days ago • 19
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 27 days ago • 120
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 160
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Paper • 2608.05850 • Published Aug 6 • 23
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published Aug 6 • 103
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published Aug 4 • 26