Policy Optimization
Consider a control policy parametrized by parameter vector θ
A policy tells the agent how to choose actions. We represent it as πθ, where θ is the collection of adjustable parameters. In a neural network, these are the weights and biases. Changing θ changes how the agent behaves.
Policy optimization adjusts θ to increase the policy’s expected total reward. The diagram shows the interaction loop: observe the state, choose an action, receive the next state and reward, and use the rewards to improve the policy.
Maximize expected total reward
- max over θ: search for parameter values that make the policy perform as well as possible.
- R(st): the reward associated with the state at time t, using the slide’s state-based reward notation.
- Sum from t = 0 to H: add rewards across the finite horizon, so the objective accounts for the whole sequence of decisions.
- E and | πθ: average over trajectories produced by following this policy. Different action samples, initial states, or environment outcomes can produce different total rewards.
For each candidate θ, imagine running its policy many times and averaging the total rewards. We want θ with the highest average. Changing an early action can affect later states and rewards, so optimizing the policy considers these future consequences too.
Stochastic policy class (smooths out the problem)
Here u is an action and s is the current state. For discrete actions, the policy assigns a probability to each available action, and these probabilities sum to 1. The agent samples an action from this distribution. For continuous actions, πθ(u | s) describes a probability density.
For example, in one state the policy might choose left with probability 0.3 and right with probability 0.7. The same state can therefore lead to different actions on different visits. Adjusting θ changes those probabilities.
Why does this help optimization? With a differentiable stochastic policy, small parameter changes can gradually shift action probabilities. Averaging the returns over those possible actions can make the objective smoother, even when individual actions are discrete.
Consider a one-step task where right gives reward 10 and left gives 0. If the probability of right is p, the expected reward is 10p. Moving p from 0.5 to 0.6 raises the expected reward from 5 to 6. We improve the average return by shifting probabilities gradually. Sampling also provides exploration while the policy learns.
Why Policy Optimization
-
Often π can be simpler than Q or V
A policy needs to describe how to act. A value function needs to predict expected future returns. Sometimes a good action rule is easier to represent and learn than those numerical predictions.
For example, in robotic grasping, a policy might learn to move toward an object, align the gripper, and close its fingers. Predicting the expected return for every possible position, grasp angle, and finger movement can require a much more complicated function. This is a reason to learn π directly; it does not mean that every policy is simple.
-
V: doesn’t prescribe actions
V(s) tells us how good a state is, but it does not tell us which action to take there. A high value for the current state still leaves the question: which movement should the robot make?
To choose an action using V alone through one-step look-ahead, we need a dynamics model and rewards. For each candidate action, predict the possible next states, add the immediate reward to the discounted V of each next state, and average over those outcomes. Then choose the action with the largest expected value. This is the one-step Bellman back-up in the slide.
-
Q: need to be able to efficiently solve
Qθ(s, u) predicts the return of taking action u in state s. The arg max returns the action with the highest prediction. If there are only a few discrete actions, we can evaluate them all and pick the best.
Continuous or high-dimensional actions make this difficult. A robot may choose real-valued torques for many joints at once. There are infinitely many possible action vectors, so we cannot enumerate them. A general neural Q-function can also have many local peaks, making the search for its best action expensive at every decision.
A directly learned policy can produce an action, or the parameters of an action distribution, with a network evaluation. The agent can then execute or sample an action without repeatedly solving this maximization problem.
Work partially addressing the continuous-action challenge:
- NAF: Gu, Lillicrap, Sutskever, Levine, ICML 2016.
- Input Convex NNs: Amos, Xu, Kolter, arXiv 2016.
- Deep Energy Q: Haarnoja, Tang, Abbeel, Levine, ICML 2017.