naver-clova-ix/donut-base
Image-to-Text β’ Updated β’ 45.4k β’ 255
Audio Conditioned LipSync with Latent Diffusion Models
Generate consistent image sequences from text and photos
Import a portrait, click to move the head!
Edit images with sketches, colors, and text prompts
Line Art Colorization with Precise Reference Following
Track, rank and evaluate open LLMs and chatbots
Explore and submit LLM benchmarks
Transcribe audio or YouTube video into text
ALA