Detect and label objects in images and videos
Convert voice to another voice
Generate audio from text using VITS model