Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs Paper • 2608.26684 • Published Aug 27 • 24
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 28 days ago • 45
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation Paper • 2606.02479 • Published Jun 1 • 25
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published 24 days ago • 108
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting Paper • 2603.21304 • Published Mar 22 • 34
Open-RS Collection Model weights & datasets in the paper "Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t" • 8 items • Updated Mar 21, 2025 • 13
view article Article Fine-tuning SmolLM with Group Relative Policy Optimization (GRPO) by following the Methodologies prithivMLmods • Feb 17, 2025 • 30
LLaMo: Large Language Model-based Molecular Graph Assistant Paper • 2411.00871 • Published Oct 31, 2024 • 22