nguyenthilaitrieulong/DeepSeek-R1-Distill-Llama-70B-abliterated Text Generation • 71B • Updated 3 days ago • 368 • 1
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 8 days ago • 301
baohao/OPD_Search_Qwen3-8B_SFT-RL_to_Search_Qwen3-4B-Instruct-2507_SFT 4B • Updated 14 days ago • 20 • 1
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 22 days ago • 209
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Paper • 2607.05691 • Published Jul 6 • 4
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning Paper • 2606.32017 • Published Jun 30 • 12
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Paper • 2606.19980 • Published Jun 18 • 15