Submitted by Jianzong Wu 25 LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing Peking University 90 2
Submitted by XuWan 7 The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs Peking University 0 2
Submitted by yunyangge 21 OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning Peking University 68 2
Submitted by Hao Liang 55 DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Peking University 241 2
Submitted by Ji Shi 7 RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting Peking University 113 1
Submitted by yfdeng 14 StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Peking University 1
Submitted by mengfanxu 12 GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding Peking University 1
Submitted by Ivan Tang 25 VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction Peking University 2
Submitted by xwm 56 Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Peking University 39 3
Submitted by Zeyu Zhang 2 Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction Peking University 13 1
Submitted by Zeyu Zhang 9 PresentAgent-2: Towards Generalist Multimodal Presentation Agents Peking University 18 1
Submitted by Gongbo Zhang 50 Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models Peking University 71 3