Kev Collection Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own • 6 items • Updated about 13 hours ago • 31
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 4 days ago • 10
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 30 days ago • 104
view article Article **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** nvidia • 9 days ago • 66
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 12 days ago • 55
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 30 days ago • 114
TerraMind: Large-Scale Generative Multimodality for Earth Observation Paper • 2504.11171 • Published Apr 15, 2025 • 3
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 25 days ago • 166
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 30 days ago • 144
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 30 days ago • 186
K2 Horizon Collection K2 Horizon models, datasets, and supporting resources • 24 items • Updated about 2 hours ago • 137
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published Sep 1 • 53
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Paper • 2608.24845 • Published Aug 25 • 16
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Paper • 2608.24053 • Published Aug 25 • 71