Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Paper • 2605.30159 • Published May 28 • 6
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Paper • 2606.03503 • Published Jun 2 • 25