Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

SeViLA

community
Activity Feed Request to join this org

AI & ML interests

None defined yet.

Yu's profile picture Abhay Zala's profile picture

Shoubin 
submitted a paper to Daily Papers 4 months ago

Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

Paper • 2603.22529 • Published Mar 23 • 7
Shoubin 
submitted a paper to Daily Papers 5 months ago

VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting

Paper • 2603.14659 • Published Mar 15 • 6
Shoubin 
submitted a paper to Daily Papers 6 months ago

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

Paper • 2602.08236 • Published Feb 9 • 9
Shoubin 
authored a paper over 1 year ago

Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel

Paper • 2412.08467 • Published Dec 11, 2024 • 6
Shoubin 
updated a Space almost 2 years ago
Runtime error
Agents
12

SeViLA Demo

⛓
12

abhayzala 
authored a paper over 2 years ago

Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

Paper • 2404.09967 • Published Apr 15, 2024 • 21
abhayzala 
authored a paper almost 3 years ago

VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Paper • 2309.15091 • Published Sep 26, 2023 • 35
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs