Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
Abstract
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify Decision--Timestamp Mismatch: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce AlignOPSD, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate AlignOPSD with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. AlignOPSD outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
Community
We introduce AlignOPSD, a framework for improving on-policy distillation in long-horizon agents.
The key insight is that temporal alignment does not necessarily imply decision alignment. Existing approaches typically assign supervision based on timestamps, while real agent decisions often unfold across multiple interaction steps. This creates a hidden mismatch between where supervision is provided and where decisions are actually made.
AlignOPSD first aligns supervision with functional decisions and then performs hierarchical credit assignment over decision spans. By moving beyond timestamp-level alignment, our approach provides more reliable training signals for long-horizon agent learning.
We evaluate AlignOPSD on several interactive agent benchmarks and show consistent improvements over existing baselines.
Get this paper in your agent:
hf papers read 2609.33391 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper