PhoneBuddy-4B
Phone-use GUI agent - screenshot + task to next action
None defined yet.
Phone-use GUI agent - screenshot + task to next action
GUI grounding with VISTA-9B β predict click coordinates
Multi-view visual reasoning VLM based on Qwen3-VL 4B
Object and Material Selection VLM
Document-parsing VLM (1.2B) by KoreaDeep
Vietnamese text-to-speech with Kokoro TTS
Interleaved text and image generation with SenseNova-U1
Scientific object generator (molecules, proteins, materials)
Real-time audio-visual social world model (22B)
One-step autoregressive image-to-video generation
Predicts click coordinates on GUI screenshots
SP3 Spherical Priors for Plug-and-Play Image Restoration
Ground localized defects in AI-generated images
Speech-to-text transcription with MOSS-Transcribe-preview-2B
Phone-use GUI agent that predicts actions from screenshots
Enhance text-to-image with the Krea2 Enhancer LoRA on Turbo
Krea-2-Turbo with IdeoKrea style LoRA demo
Multilingual speech recognition in 30+ languages
Video understanding with InternVideo3 MΒ²LA multimodal LLM
Semantic-First Diffusion text-to-image in 4 steps
PP-OCRv6 medium recognition β text recognition from images
Detect text regions in images with PP-OCRv6 medium
3D-aware object insertion with the DIRECT model (ICML 2026)
ASR for 52 languages and dialects