Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 16 days ago • 302
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Paper • 2411.11066 • Published Nov 17, 2024 • 1
Weakly Supervised Face Naming with Symmetry-Enhanced Contrastive Loss Paper • 2210.08957 • Published Oct 17, 2022 • 1
Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks Paper • 2403.09377 • Published Mar 14, 2024 • 1
Visually-Aware Context Modeling for News Image Captioning Paper • 2308.08325 • Published Aug 16, 2023 • 1
Alleviating Exposure Bias in Diffusion Models through Sampling with Shifted Time Steps Paper • 2305.15583 • Published May 24, 2023 • 2