FlowRL: Matching Reward Distributions for LLM Reasoning Paper • 2509.15207 • Published Sep 18, 2025 • 119
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning Paper • 2509.22761 • Published Sep 26, 2025 • 1
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Paper • 2608.02585 • Published 3 days ago • 22
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts Paper • 2512.24885 • Published Dec 31, 2025 • 5
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts Paper • 2512.24885 • Published Dec 31, 2025 • 5
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts Paper • 2512.24885 • Published Dec 31, 2025 • 5
TongSIM: A General Platform for Simulating Intelligent Machines Paper • 2512.20206 • Published Dec 23, 2025 • 28
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Paper • 2512.07461 • Published Dec 8, 2025 • 80
FlowRL: Matching Reward Distributions for LLM Reasoning Paper • 2509.15207 • Published Sep 18, 2025 • 119
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling Paper • 2506.08672 • Published Jun 10, 2025 • 30
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space Paper • 2505.13308 • Published May 19, 2025 • 27
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space Paper • 2505.13308 • Published May 19, 2025 • 27 • 4
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space Paper • 2505.13308 • Published May 19, 2025 • 27 • 4
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space Paper • 2505.13308 • Published May 19, 2025 • 27