$φ$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models Paper • 2602.22601 • Published Feb 26
Running Featured 1.4k FineWeb: decanting the web for the finest text data at scale 🍷 1.4k Explore and download the FineWeb web‑scale text dataset
Running on CPU Upgrade Featured 3.25k The Smol Training Playbook 📚 3.25k The secrets to building world-class LLMs
Running 3.95k The Ultra-Scale Playbook 🌌 3.95k The ultimate guide to training LLM on large GPU Clusters
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Paper • 2502.05173 • Published Feb 7, 2025 • 64
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 125
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding Paper • 2501.12380 • Published Jan 21, 2025 • 82
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning Paper • 2501.12948 • Published Jan 22, 2025 • 457
Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise Paper • 2501.08331 • Published Jan 14, 2025 • 20