llmfan46/Qwen3.5-40B-Claude-4.5-Opus-High-Reasoning-Thinking-uncensored-heretic-GGUF Image-Text-to-Text • 39B • Updated Apr 13 • 2.39k • 18
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning Paper • 2609.35505 • Published 9 days ago • 26
llmfan46/Qwen3.5-40B-Claude-4.5-Opus-High-Reasoning-Thinking-uncensored-heretic Image-Text-to-Text • 40B • Updated Apr 13 • 290 • 17
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models Paper • 2609.32607 • Published 11 days ago • 154
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 15 days ago • 103
InternW0: A Foundational Physical World Model for Efficient Real-World Interactions Paper • 2609.27656 • Published 14 days ago • 15
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 16 days ago • 55