TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning Paper • 2404.09275 • Published Apr 14, 2024
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting Paper • 2608.07693 • Published 6 days ago
BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition Paper • 2505.00059 • Published Apr 30, 2025