How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Paper • 2605.27310 • Published May 26 • 18
Assessing and Learning Alignment of Unimodal Vision and Language Models Paper • 2412.04616 • Published Dec 5, 2024
Exploring the Best Practices of Query Expansion with Large Language Models Paper • 2401.06311 • Published Jan 12, 2024 • 1
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents Paper • 2609.06702 • Published 24 days ago • 26
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Paper • 2607.12625 • Published Jul 15 • 56
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Paper • 2605.27310 • Published May 26 • 18
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Paper • 2605.27310 • Published May 26 • 18
RiT: Vanilla Diffusion Transformers Suffice in Representation Space Paper • 2605.21981 • Published May 21 • 8
Communicating about Space: Language-Mediated Spatial Integration Across Partial Views Paper • 2603.27183 • Published Mar 28 • 18
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings Paper • 2603.13594 • Published Mar 13 • 150
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs Paper • 2602.00462 • Published Jan 31 • 21
Grounding Computer Use Agents on Human Demonstrations Paper • 2511.07332 • Published Nov 10, 2025 • 107
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning Paper • 2505.20046 • Published May 26, 2025 • 18
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Paper • 2505.04921 • Published May 8, 2025 • 187