Policy Optimization

Consider a control policy parametrized by parameter vector θ

A policy tells the agent how to choose actions. We represent it as πθ, where θ is the collection of adjustable parameters. In a neural network, these are the weights and biases. Changing θ changes how the agent behaves.

Policy optimization adjusts θ to increase the policy’s expected total reward. The diagram shows the interaction loop: observe the state, choose an action, receive the next state and reward, and use the rewards to improve the policy.

A state enters the policy network, which sends an action to the environment. The environment returns the next state and reward. Rewards guide updates to the policy’s parameters theta.

Maximize expected total reward

maxθ↑Choose θE[∑t=0HR(st)|πθ]↑Expected total reward under πθ

For each candidate θ, imagine running its policy many times and averaging the total rewards. We want θ with the highest average. Changing an early action can affect later states and rewards, so optimizing the policy considers these future consequences too.

Stochastic policy class (smooths out the problem)

πθ(u|s)↑Action distribution:probability of action u in state s

Here u is an action and s is the current state. For discrete actions, the policy assigns a probability to each available action, and these probabilities sum to 1. The agent samples an action from this distribution. For continuous actions, πθ(u | s) describes a probability density.

For example, in one state the policy might choose left with probability 0.3 and right with probability 0.7. The same state can therefore lead to different actions on different visits. Adjusting θ changes those probabilities.

Why does this help optimization? With a differentiable stochastic policy, small parameter changes can gradually shift action probabilities. Averaging the returns over those possible actions can make the objective smoother, even when individual actions are discrete.

Consider a one-step task where right gives reward 10 and left gives 0. If the probability of right is p, the expected reward is 10p. Moving p from 0.5 to 0.6 raises the expected reward from 5 to 6. We improve the average return by shifting probabilities gradually. Sampling also provides exploration while the policy learns.

Why Policy Optimization

Work partially addressing the continuous-action challenge: