srinivas2002/rollout_eval_act_v2v3_20260729_20260729_212521 Viewer • Updated about 21 hours ago • 4.12k • 7 • 1
AIOpsInSpace/DeepSeek-R1-Distill-Llama-8B-Abliterated-Patched Text Generation • 8B • Updated 9 days ago • 603 • 1
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Paper • 2606.30124 • Published Jun 29 • 8
WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Paper • 2607.02517 • Published 29 days ago • 33
ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning Paper • 2606.14697 • Published Jun 12 • 8
Alkatt/eval_demo_LAVLA_Async_VI_whiteblackcube_v2_2026_06_02_white_v2 Viewer • Updated Jun 2 • 4.49k • 26 • 1
Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs Paper • 2505.11277 • Published May 16, 2025 • 64