Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
xiaokun sun's picture

xiaokun sun

yscsb
7
  • sunxk2020

AI & ML interests

None yet

Organizations

None yet

upvoted a paper 3 months ago

Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration

Paper • 2309.01131 • Published Sep 3, 2023 • 2
upvoted 6 papers 4 months ago

VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model

Paper • 2505.03739 • Published May 6, 2025 • 10

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

Paper • 2512.12633 • Published Dec 14, 2025 • 2

RISE-Video: Can Video Generators Decode Implicit World Rules?

Paper • 2602.05986 • Published Feb 5 • 28

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

Paper • 2603.06577 • Published Mar 6 • 50

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Paper • 2604.05015 • Published Apr 6 • 236

When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition

Paper • 2603.16256 • Published Mar 17 • 2
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs