Temporal decomposition

Let’s Decompose Path into States and Actions

The policy gradient from 6.2 contains ∇θ log P(τ(i); θ). We can compute this trajectory log-probability gradient by adding contributions from the individual actions along the path.

The superscript (i) identifies the sampled trajectory; t identifies a time step within it. At step t, the agent observes st(i), takes action ut(i), and reaches st+1(i). Here each path includes the next state after its last action.

Likelihood Ratio Gradient Estimate

The following expression provides an unbiased estimate of the gradient, and we can compute it without access to a dynamics model:

g^=1m∑i=1m↑Average m trajectories∇θlogP(τ(i);θ)↑Trajectory log-probability gradientR(τ(i))↑Observed total return

Collect m trajectories using the current policy πθ. For each trajectory, multiply its log-probability gradient by its observed total return R(τ(i)). Average these vectors to obtain ĝ, the estimate used for a policy update.

Here, the temporal decomposition gives:

∇θlogP(τ(i);θ)=∑t=0H∇θlogπθ(ut(i)|st(i))↑No dynamics model required

The trajectory gradient is the sum of action log-probability gradients. Each term can be computed from the policy network and an observed state-action pair. The environment supplies the sampled transitions and rewards.

Substituting this into the estimate:

g^=1m∑i=1m(∑t=0H∇θlogπθ(ut(i)|st(i)))R(τ(i))

We differentiate the policy’s log-probabilities with respect to θ. The sampled states, actions, and returns are held fixed during this calculation; no derivative of the reward function or environment dynamics is needed.

Unbiased means

E[g^]↑Expected estimate=∇θU(θ)↑True gradient

Keep θ fixed and imagine repeatedly collecting fresh batches of trajectories from πθ. The expectation averages ĝ over all those possible batches. It equals the true policy gradient ∇θU(θ), under the assumptions used in the derivation.

A single batch can still give a poor estimate, including the wrong sign for a gradient component. “Unbiased” describes the average across repeated samples, rather than the accuracy of any one estimate. Even a one-trajectory estimate can be unbiased.

Unbiased but very noisy

Different sampled paths have different actions, transitions, and returns. Their return-weighted gradients can therefore vary greatly between batches. A high-variance estimate makes update directions less reliable, even when its expectation is correct.

For independent trajectories with finite variance, averaging more trajectories reduces the variance of the estimate, but requires more interaction with the environment.

Fixes that lead to real-world practicality: