LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published Aug 18 • 26
What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models Paper • 2503.24235 • Published Mar 31, 2025 • 55
view article Article Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial open-r1 • Jan 31, 2025 • 52
Running 602 Scaling test-time compute 📈 602 Boost LLM answers with flexible test‑time search strategies
Meta Llama 3 Collection This collection hosts the transformers and original repos of the Meta Llama 3 and Llama Guard 2 releases • 5 items • Updated Dec 6, 2024 • 1.01k