DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles Updated Jan 10, 2025 • 33 • 6
HuRo: Robotizing Human Videos for Scalable VLA Pretraining Paper • 2609.10706 • Published 10 days ago • 28
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 11 days ago • 90
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Paper • 2609.25165 • Published 7 days ago • 73
xinyuzhou2000/Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Model Viewer • Updated Sep 17, 2023 • 18k • 85 • 7
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 7 days ago • 155
ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition Audio Classification • 0.3B • Updated Oct 24, 2024 • 20.1k • 258
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 10 days ago • 149
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design Paper • 2609.16251 • Published 14 days ago • 14
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 11 days ago • 185
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 11 days ago • 54
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 11 days ago • 57
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation Paper • 2609.12397 • Published 11 days ago • 46