WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Paper • 2607.23909 • Published 4 days ago • 5
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 18 days ago • 77
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published 18 days ago • 44
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published 23 days ago • 26
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Paper • 2607.07386 • Published 23 days ago • 13
SiamJEPA: On the Role of Siamese Student Encoders in JEPA Paper • 2607.04044 • Published 27 days ago • 1
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better Paper • 2607.04884 • Published 25 days ago • 9
Unified Audio Intelligence Without Regressing on Text Intelligence Paper • 2607.05196 • Published 25 days ago • 23
view post Post 262 @alvarobartt published a step-by-step guide to deploy zai-org/GLM-5.2 on AMD GPUs, using newly released features in Microsoft Foundry. The FP8 model fits on a single node of MI300X GPUs, which cuts the bill in half vs. H100.https://alvarobartt.com/goal-glm-5.2-on-foundry/ See translation 👀 1 1 + Reply
Duration Aware Scheduling for ASR Serving Under Workload Drift Paper • 2603.11273 • Published Mar 11 • 3
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models Paper • 2606.03748 • Published Jun 2 • 22
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Paper • 2605.27295 • Published May 26 • 23
Geometric Context Transformer for Streaming 3D Reconstruction Paper • 2604.14141 • Published Apr 15 • 37
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens Paper • 2604.04913 • Published Apr 6 • 12
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios Paper • 2603.28130 • Published Mar 30 • 11
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders Paper • 2603.19209 • Published Mar 19 • 6
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Paper • 2603.14482 • Published Mar 15 • 37