Temporal decomposition
Let’s Decompose Path into States and Actions
The policy gradient from 6.2 contains ∇θ log P(τ(i); θ). We can compute this trajectory log-probability gradient by adding contributions from the individual actions along the path.
The superscript (i) identifies the sampled trajectory; t identifies a time step within it. At step t, the agent observes st(i), takes action ut(i), and reaches st+1(i). Here each path includes the next state after its last action.
-
Factor the trajectory probability
With Markov dynamics and a policy that depends on the current state, a path’s probability factors into the probabilities of its actions and state transitions, multiplied across time.
- Dynamics model: P(st+1(i) | st(i), ut(i)) is the probability of reaching the observed next state after taking that action.
- Policy: πθ(ut(i) | st(i)) is the probability that the policy assigns to the action actually taken in that state.
The initial-state probability is omitted here. If the initial-state distribution is fixed with respect to θ, its log-probability contributes zero to the gradient.
-
Turn the product into sums of logs
Apply log(ab) = log a + log b. The product over time becomes a sum, which we separate into log transition probabilities and log action probabilities. This is an exact algebraic rewrite.
-
Remove the dynamics gradient
The environment dynamics do not depend on the policy parameters θ. For the fixed states and actions in this sampled path, each transition log-probability therefore has zero derivative with respect to θ. Only the policy terms remain.
Likelihood ratio changes probabilities of experienced paths. When differentiating the likelihood of one observed path, we hold its sampled states and actions fixed. Changing θ changes the probability assigned to that path and therefore how likely it is to occur in future runs.
A path derivative follows how actions and subsequent states change with θ, and differentiates the return through those dependencies when the required derivatives are available. The likelihood ratio calculation instead differentiates the log-probability of a fixed sampled path; it requires no derivatives through the environment or reward function.
-
Differentiate each action log-probability
Differentiation is linear, so the gradient of the finite sum equals the sum of the gradients. For every observed state-action pair, compute ∇θ log πθ(ut(i) | st(i)) using the policy network, then add these vectors over time.
No dynamics model required: we need sampled experience and the policy’s action log-probabilities. We can obtain the gradient without knowing the transition probabilities or differentiating through the environment. The dynamics still determine the sampled paths; their θ derivatives are what disappear from this calculation.
Likelihood Ratio Gradient Estimate
The following expression provides an unbiased estimate of the gradient, and we can compute it without access to a dynamics model:
Collect m trajectories using the current policy πθ. For each trajectory, multiply its log-probability gradient by its observed total return R(τ(i)). Average these vectors to obtain ĝ, the estimate used for a policy update.
Here, the temporal decomposition gives:
The trajectory gradient is the sum of action log-probability gradients. Each term can be computed from the policy network and an observed state-action pair. The environment supplies the sampled transitions and rewards.
Substituting this into the estimate:
- Sum over t: add contributions from the actions within one trajectory.
- R(τ(i)): in this version, every action in that trajectory receives the same total-return weight.
- Average over i: combine the contributions of the m sampled trajectories.
We differentiate the policy’s log-probabilities with respect to θ. The sampled states, actions, and returns are held fixed during this calculation; no derivative of the reward function or environment dynamics is needed.
Unbiased means
Keep θ fixed and imagine repeatedly collecting fresh batches of trajectories from πθ. The expectation averages ĝ over all those possible batches. It equals the true policy gradient ∇θU(θ), under the assumptions used in the derivation.
A single batch can still give a poor estimate, including the wrong sign for a gradient component. “Unbiased” describes the average across repeated samples, rather than the accuracy of any one estimate. Even a one-trajectory estimate can be unbiased.
Unbiased but very noisy
Different sampled paths have different actions, transitions, and returns. Their return-weighted gradients can therefore vary greatly between batches. A high-variance estimate makes update directions less reliable, even when its expectation is correct.
For independent trajectories with finite variance, averaging more trajectories reduces the variance of the estimate, but requires more interaction with the environment.
Fixes that lead to real-world practicality:
-
Baseline
Subtract a reference return so that the weight reflects performance relative to that reference. An appropriate baseline can reduce variance while preserving the expected gradient. A state-dependent baseline may depend on the current state, but must not depend on the action whose log-probability is being differentiated.
-
Temporal structure
An action cannot influence rewards already received. Weight its log-probability gradient using rewards from that time onward. Removing past rewards from that weight preserves the expected gradient and can reduce noise.
-
Next: Trust region / natural gradient
These methods address how to turn a gradient estimate into a controlled policy update. A trust region limits how far the policy distribution can change in one update. A natural gradient accounts for the geometry of that distribution when choosing the update direction.