kiyoxi2022 's Collections paper-reading
updated
Paper
• 2605.18747
• Published • 225
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper
• 2605.12500
• Published • 194
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper
• 2604.27660
• Published • 171
PhysBrain 1.0 Technical Report
Paper
• 2605.15298
• Published • 145
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper
• 2605.20025
• Published • 191
MMSkills: Towards Multimodal Skills for General Visual Agents
Paper
• 2605.13527
• Published • 122
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
Paper
• 2605.06130
• Published • 116
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
Paper
• 2605.18739
• Published • 116
Qwen-Image-2.0 Technical Report
Paper
• 2605.10730
• Published • 118
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
Paper
• 2605.18233
• Published • 93
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper
• 2605.00658
• Published • 86
Lance: Unified Multimodal Modeling by Multi-Task Synergy
Paper
• 2605.18678
• Published • 79
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Paper
• 2605.23902
• Published • 47
Image-Text-to-Text
• 27B • Updated • 9.77k
• 81
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the
LLM Era
Paper
• 2503.12329
• Published • 28
Text-to-Image
• 8B • Updated • 68.3k
• • 665
GenClaw: Code-Driven Agentic Image Generation
Paper
• 2605.30248
• Published • 41
Text Generation
• Updated • 166
• 43
deepseek-ai/DeepSeek-V4-Pro
Text Generation
• 862B • Updated • 1.66M
• • 5.31k
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
Paper
• 2605.02641
• Published • 1
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
Paper
• 2605.28816
• Published • 433
Image-Text-to-Text
• 9B • Updated • 99
• 9
Text-to-Image
• Updated • 3
• 1
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper
• 2604.13016
• Published • 114
Image-Text-to-Video
• Updated • 218
• 289
Trust Region On-Policy Distillation
Paper
• 2606.01249
• Published • 48
65B • Updated • 101k
• 217
VLM3: Vision Language Models Are Native 3D Learners
Paper
• 2605.30561
• Published • 26
Cosmos 3: Omnimodal World Models for Physical AI
Paper
• 2606.02800
• Published • 141
Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
Paper
• 2606.04527
• Published • 28
Personal AI Agent for Camera Roll VQA
Paper
• 2606.05275
• Published • 20
Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing
Paper
• 2606.05172
• Published • 1
Representation Forcing for Bottleneck-Free Unified Multimodal Models
Paper
• 2605.31604
• Published • 63
Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Paper
• 2606.07433
• Published • 21
Latent Spatial Memory for Video World Models
Paper
• 2606.09828
• Published • 71
CoVEBench: Can Video Editing Models Handle Complex Instructions?
Paper
• 2606.08415
• Published • 52
Trajectory-Refined Distillation
Paper
• 2606.08432
• Published • 7
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement
Learning
Paper
• 2509.22647
• Published • 37
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
Paper
• 2602.11401
• Published
HiDream-ai/HiDream-O1-Image
Image-Text-to-Image
• 9B • Updated • 15.6k
• 514
Kwai Keye-VL-2.0 Technical Report
Paper
• 2606.10651
• Published • 194
Paper
• 2606.13392
• Published • 153
InterleaveThinker: Reinforcing Agentic Interleaved Generation
Paper
• 2606.13679
• Published • 83
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
Paper
• 2606.13432
• Published • 113
Rethinking RAG in Long Videos: What to Retrieve and How to Use It?
Paper
• 2606.13141
• Published • 36
Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
Paper
• 2606.09076
• Published • 66
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Paper
• 2606.14777
• Published • 215
VisualClaw: A Real-Time, Personalized Agent for the Physical World
Paper
• 2606.16295
• Published • 28
Memento: Reconstruct to Remember for Consistent Long Video Generation
Paper
• 2606.14667
• Published • 18
PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions
Paper
• 2606.14832
• Published • 12
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
Paper
• 2606.18249
• Published • 14
Sumi: Open Uniform Diffusion Language Model from Scratch
Paper
• 2606.19005
• Published • 12
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
Paper
• 2606.19534
• Published • 65
DanceOPD: On-Policy Generative Field Distillation
Paper
• 2606.27377
• Published • 81
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Paper
• 2606.26907
• Published • 49
Qwen-Image-2.0-RL Technical Report
Paper
• 2606.27608
• Published • 52
deepseek-ai/DeepSeek-V4-Pro-DSpark
Text Generation
• 889B • Updated • 47.8k
• 510
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Paper
• 2606.26740
• Published • 82
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
Paper
• 2607.00248
• Published • 32
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
Paper
• 2607.05373
• Published • 65
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Paper
• 2607.04438
• Published • 64
Wan-Streamer v0.2: Higher Resolution, Same Latency
Paper
• 2607.04443
• Published • 40
Text-to-Image
• Updated • 1
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Paper
• 2607.07675
• Published • 64
Video Generation Models are General-Purpose Vision Learners
Paper
• 2607.09024
• Published • 84
Trust Region Policy Distillation
Paper
• 2607.04751
• Published • 34
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Paper
• 2607.13125
• Published • 136
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
Paper
• 2607.12625
• Published • 80
Registers Matter for Pixel-Space Diffusion Transformers
Paper
• 2605.16147
• Published • 26
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Paper
• 2606.29538
• Published • 139
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Paper
• 2607.16401
• Published • 43
Bridging Supervised Learning and Reinforcement Learning in Math
Reasoning
Paper
• 2505.18116
• Published • 5
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Paper
• 2607.19064
• Published • 68
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Paper
• 2607.19191
• Published • 293
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Paper
• 2603.06507
• Published • 2
Representation Alignment for Generation: Training Diffusion Transformers
Is Easier Than You Think
Paper
• 2410.06940
• Published • 14
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring
and Expert-Level Understanding
Paper
• 2507.14533
• Published • 8