OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 5 days ago • 26
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Paper • 2606.07639 • Published Jun 1 • 3
laion/moss-tts-local-transformer-4.55b-voice-acting Text-to-Speech • 4B • Updated 13 days ago • 2.93k • 1
MOSS Transcribe Collection A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription. • 4 items • Updated 20 days ago • 13
OpenMOSS-Team/MOSS-Transcribe-preview-2B Automatic Speech Recognition • 2B • Updated Jun 26 • 2.8k • 44