Shijie Geng
makitanikaze
AI & ML interests
None yet
Recent Activity
upvoted a paper 10 days ago
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation upvoted a paper 10 days ago
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses updated a collection about 2 months ago
vlmOrganizations
None yet
rl
gui agent
-
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
Paper • 2505.21496 • Published • 38 -
Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
Paper • 2506.04614 • Published • 19 -
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
Paper • 2506.09790 • Published • 53 -
WebSailor: Navigating Super-human Reasoning for Web Agent
Paper • 2507.02592 • Published • 127
diffusion
vlm
-
Images are Worth Variable Length of Representations
Paper • 2506.03643 • Published • 4 -
PyVision: Agentic Vision with Dynamic Tooling
Paper • 2507.07998 • Published • 33 -
Scaling RL to Long Videos
Paper • 2507.07966 • Published • 160 -
Douyin Multimodal Embedding Model Technical Report
Paper • 2608.02148 • Published • 15
world model
diffusion
rl
vlm
-
Images are Worth Variable Length of Representations
Paper • 2506.03643 • Published • 4 -
PyVision: Agentic Vision with Dynamic Tooling
Paper • 2507.07998 • Published • 33 -
Scaling RL to Long Videos
Paper • 2507.07966 • Published • 160 -
Douyin Multimodal Embedding Model Technical Report
Paper • 2608.02148 • Published • 15
gui agent
-
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
Paper • 2505.21496 • Published • 38 -
Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
Paper • 2506.04614 • Published • 19 -
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
Paper • 2506.09790 • Published • 53 -
WebSailor: Navigating Super-human Reasoning for Web Agent
Paper • 2507.02592 • Published • 127