DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation Paper • 2609.33485 • Published 8 days ago • 3
DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation Paper • 2609.33485 • Published 8 days ago • 3
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching Paper • 2609.35673 • Published 7 days ago • 34
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching Paper • 2609.35673 • Published 7 days ago • 34
PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences Paper • 2308.13924 • Published Aug 26, 2023
NoLiMa: Long-Context Evaluation Beyond Literal Matching Paper • 2502.05167 • Published Feb 7, 2025 • 16
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models Paper • 2501.00874 • Published Jan 1, 2025 • 13
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Paper • 2412.15484 • Published Dec 20, 2024 • 15
VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs Paper • 2407.01863 • Published Jul 2, 2024
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions Paper • 2603.03646 • Published Mar 4 • 8
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions Paper • 2603.03646 • Published Mar 4 • 8
PageGuide: Browser extension to assist users in navigating a webpage and locating information Paper • 2604.23772 • Published Apr 26 • 7
SketchVLM: Vision language models can annotate images to explain thoughts and guide users Paper • 2604.22875 • Published Apr 23 • 38
ViT-AdaLA: Adapting Vision Transformers with Linear Attention Paper • 2603.16063 • Published Mar 17 • 2
ViT-AdaLA: Adapting Vision Transformers with Linear Attention Paper • 2603.16063 • Published Mar 17 • 2