HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published Aug 5 • 41
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment Paper • 2603.23883 • Published Mar 25 • 5
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference Paper • 2603.22053 • Published Mar 23 • 3
AlignBench: Benchmarking Fine-Grained Image-Text Alignment with Synthetic Image-Caption Pairs Paper • 2511.20515 • Published Nov 25, 2025 • 5
AgroBench: Vision-Language Model Benchmark in Agriculture Paper • 2507.20519 • Published Jul 28, 2025 • 8
Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention Paper • 2509.09116 • Published Sep 11, 2025
ActionVOS: Actions as Prompts for Video Object Segmentation Paper • 2407.07402 • Published Jul 10, 2024