Note / maximum-entropy-formulation.html
SteveZeyuZhang
Publish Note study website
5e1c603
Raw History Blame Contribute Delete
46.9 kB
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="color-scheme" content="light">
<title>Maximum entropy formulation</title>
<link rel="stylesheet" href="styles.css">
</head>
<body>
<main class="topic chapter-page">
<a class="back" href="index.html" aria-label="Home">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" aria-hidden="true">
<path d="m14 6-6 6 6 6M8 12h12" />
</svg>
</a>
<h1 class="topic-title">Maximum entropy formulation</h1>
<section class="value-intro" aria-labelledby="near-optimal-solutions">
<h2 id="near-optimal-solutions">What if we could find a distribution over near-optimal solutions?</h2>
<ul class="policy-explanation">
<li><strong>More robust policy:</strong> if the environment changes, the distribution over near-optimal solutions might still have some good ones for the new situation</li>
<li><strong>More robust learning:</strong> if we can retain a distribution over near-optimal solutions our agent will collect more interesting exploratory data during learning</li>
</ul>
</section>
<section class="value-case entropy-definition" aria-labelledby="entropy">
<h2 id="entropy"><a class="reference-link" href="https://ocw.mit.edu/courses/6-02-introduction-to-eecs-ii-digital-communication-systems-fall-2012/4bcb4b872e8e98459ce71e77d4fe9c42_MIT6_02F12_chap02.pdf#page=4" target="_blank" rel="noopener noreferrer">Entropy</a></h2>
<p class="mdp-explanation">Entropy = measure of uncertainty over random variable X. Here, X is discrete.</p>
<p class="mdp-explanation">It also describes the ideal average number of bits needed to encode an outcome of X.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Entropy is the probability-weighted average information over all outcomes">
<math display="block" aria-label="H of X equals the sum over i of p of x i times log base two of one divided by p of x i, which equals minus the sum over i of p of x i times log base two of p of x i">
<mrow><mi>ℋ</mi><mo>(</mo><mi>X</mi><mo>)</mo></mrow><mo>=</mo><munder><mo>∑</mo><mi>i</mi></munder>
<munder accentunder="false">
<mrow><mi>p</mi><mo>(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>)</mo></mrow>
<mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em">
<mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr>
<mtr><mtd><mtext>Probability</mtext></mtd></mtr>
</mtable></mstyle>
</munder>
<munder accentunder="false">
<mrow><msub><mi>log</mi><mn>2</mn></msub><mfrac><mn>1</mn><mrow><mi>p</mi><mo>(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>)</mo></mrow></mfrac></mrow>
<mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em">
<mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr>
<mtr><mtd><mtext>Information (bits)</mtext></mtd></mtr>
</mtable></mstyle>
</munder>
<mo>=</mo><mo>−</mo><munder><mo>∑</mo><mi>i</mi></munder><mrow><mi>p</mi><mo>(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>)</mo></mrow><msub><mi>log</mi><mn>2</mn></msub><mrow><mi>p</mi><mo>(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>)</mo></mrow>
</math>
</div>
<ul class="policy-explanation">
<li><math aria-label="p of x i"><mi>p</mi><mo>(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>)</mo></math> is the probability of outcome <math aria-label="x i"><msub><mi>x</mi><mi>i</mi></msub></math>.</li>
<li><math aria-label="log base two of one divided by p of x i"><msub><mi>log</mi><mn>2</mn></msub><mfrac><mn>1</mn><mrow><mi>p</mi><mo>(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>)</mo></mrow></mfrac></math> measures how much information that outcome brings. A rare result is more surprising, so it carries more information.</li>
<li>Multiplying by each outcome’s probability and summing gives the average information: entropy. The two forms are equal because <math aria-label="log base two of one divided by p equals minus log base two of p"><msub><mi>log</mi><mn>2</mn></msub><mfrac><mn>1</mn><mi>p</mi></mfrac><mo>=</mo><mo>−</mo><msub><mi>log</mi><mn>2</mn></msub><mi>p</mi></math>.</li>
</ul>
<p class="mdp-explanation">For a fair coin, each result has probability 0.5 and brings log₂(2) = 1 bit, so the entropy is 0.5 × 1 + 0.5 × 1 = 1 bit. If the result is certain, its probability is 1 and log₂(1) = 0, so the entropy is 0.</p>
<p class="mdp-explanation">The base-2 logarithm makes the unit bits. Zero-probability outcomes contribute 0. With efficient coding of long sequences of independent outcomes, entropy is the limit on the minimum average number of bits per outcome.</p>
</section>
<section class="value-case entropy-examples" aria-labelledby="entropy-examples">
<h2 id="entropy-examples">Examples</h2>
<h3>Binary random variable</h3>
<p class="mdp-explanation">Let p = Pr(X = 1), so Pr(X = 0) = 1 − p.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Entropy of a binary random variable">
<math display="block" aria-label="H of X equals minus p times log base two of p minus one minus p times log base two of one minus p">
<mi>ℋ</mi><mo>(</mo><mi>X</mi><mo>)</mo><mo>=</mo><mo>−</mo><mi>p</mi><msub><mi>log</mi><mn>2</mn></msub><mi>p</mi><mo>−</mo><mo>(</mo><mn>1</mn><mo>−</mo><mi>p</mi><mo>)</mo><msub><mi>log</mi><mn>2</mn></msub><mo>(</mo><mn>1</mn><mo>−</mo><mi>p</mi><mo>)</mo>
</math>
</div>
<figure class="entropy-binary-plot">
<img src="assets/maximum-entropy/binary-entropy.svg" alt="Binary entropy curve: entropy is zero at probabilities zero and one and reaches one bit at probability one half.">
</figure>
<p class="mdp-explanation">At p = 0 or p = 1, the result is certain and entropy is 0. At p = 0.5, both results are equally likely, so uncertainty is greatest and entropy reaches 1 bit. The curve is symmetric: swapping the two outcomes does not change the entropy.</p>
<h3>Five-outcome distributions</h3>
<div class="entropy-distributions">
<figure>
<img src="assets/maximum-entropy/balanced-distribution.svg" alt="Five outcomes with probabilities 0.25, 0.25, 0.25, 0.125, and 0.125.">
<figcaption>More balanced: H = 2.25 bits</figcaption>
</figure>
<figure>
<img src="assets/maximum-entropy/concentrated-distribution.svg" alt="Five outcomes with probabilities 0.75, 0.0625, 0.0625, 0.0625, and 0.0625.">
<figcaption>More concentrated: H ≈ 1.3 bits</figcaption>
</figure>
</div>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The more balanced distribution">
<math display="block" aria-label="p of S equals 0.25, 0.25, 0.25, 0.125, 0.125">
<mi>p</mi><mo>(</mo><mi>S</mi><mo>)</mo><mo>=</mo><mo>{</mo><mn>0.25</mn><mo>,</mo><mn>0.25</mn><mo>,</mo><mn>0.25</mn><mo>,</mo><mn>0.125</mn><mo>,</mo><mn>0.125</mn><mo>}</mo>
</math>
</div>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Entropy of the more balanced distribution">
<math display="block" aria-label="H equals three times 0.25 times log base two of four plus two times 0.125 times log base two of eight, which equals 1.5 plus 0.75, which equals 2.25 bits">
<mi>H</mi><mo>=</mo><mn>3</mn><mo>×</mo><mn>0.25</mn><mo>×</mo><msub><mi>log</mi><mn>2</mn></msub><mn>4</mn><mo>+</mo><mn>2</mn><mo>×</mo><mn>0.125</mn><mo>×</mo><msub><mi>log</mi><mn>2</mn></msub><mn>8</mn><mo>=</mo><mn>1.5</mn><mo>+</mo><mn>0.75</mn><mo>=</mo><mn>2.25</mn><mspace width="0.25em"/><mtext>bits</mtext>
</math>
</div>
<p class="mdp-explanation">Three outcomes have probability 0.25, giving 2 bits each; two have probability 0.125, giving 3 bits each. The factors 3 and 2 count the repeated probabilities.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The more concentrated distribution">
<math display="block" aria-label="p of S equals 0.75, 0.0625, 0.0625, 0.0625, 0.0625">
<mi>p</mi><mo>(</mo><mi>S</mi><mo>)</mo><mo>=</mo><mo>{</mo><mn>0.75</mn><mo>,</mo><mn>0.0625</mn><mo>,</mo><mn>0.0625</mn><mo>,</mo><mn>0.0625</mn><mo>,</mo><mn>0.0625</mn><mo>}</mo>
</math>
</div>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Entropy of the more concentrated distribution">
<math display="block" aria-label="H equals 0.75 times log base two of four thirds plus four times 0.0625 times log base two of sixteen, which is approximately 0.3 plus one, which equals 1.3 bits">
<mi>H</mi><mo>=</mo><mn>0.75</mn><mo>×</mo><msub><mi>log</mi><mn>2</mn></msub><mo>(</mo><mfrac><mn>4</mn><mn>3</mn></mfrac><mo>)</mo><mo>+</mo><mn>4</mn><mo>×</mo><mn>0.0625</mn><mo>×</mo><msub><mi>log</mi><mn>2</mn></msub><mn>16</mn><mo>≈</mo><mn>0.3</mn><mo>+</mo><mn>1</mn><mo>=</mo><mn>1.3</mn><mspace width="0.25em"/><mtext>bits</mtext>
</math>
</div>
<p class="mdp-explanation">The most likely outcome has probability 0.75, so its information is log₂(4/3). The four rare outcomes each have probability 1/16 and bring 4 bits. The exact entropy is about 1.3113 bits. This distribution is easier to predict because one outcome dominates, so its entropy is lower.</p>
</section>
<section class="value-case maximum-entropy-mdp" aria-labelledby="maximum-entropy-mdp">
<h2 id="maximum-entropy-mdp"><a class="reference-link" href="https://arxiv.org/abs/1801.01290" target="_blank" rel="noopener noreferrer">Maximum Entropy MDP</a></h2>
<h3>Regular formulation:</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The regular objective maximizes expected total reward">
<math display="block" aria-label="Maximize over policy pi the expectation of the sum from t equals zero to H of r t">
<munder><mo>max</mo><mi>π</mi></munder><mi>E</mi><mrow><mo>[</mo><munderover><mo>∑</mo><mrow><mi>t</mi><mo>=</mo><mn>0</mn></mrow><mi>H</mi></munderover><msub><mi>r</mi><mi>t</mi></msub><mo>]</mo></mrow>
</math>
</div>
<p class="mdp-explanation">Find the policy π with the highest expected total reward. The sum adds rewards rₜ from time 0 through H, and E averages over the trajectories produced by the policy and the environment. This finite-horizon formula uses undiscounted rewards.</p>
<h3>Max-ent formulation:</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The maximum entropy objective adds an entropy bonus at every time step">
<math display="block" aria-label="Maximize over policy pi the expectation of the sum from t equals zero to H of r t plus beta times the entropy of the full action distribution pi given state s t">
<munder><mo>max</mo><mi>π</mi></munder><mi>E</mi><mrow><mo minsize="3em" maxsize="3em">[</mo><munderover><mo>∑</mo><mrow><mi>t</mi><mo>=</mo><mn>0</mn></mrow><mi>H</mi></munderover>
<mrow><mo stretchy="false">(</mo>
<munder accentunder="false">
<msub><mi>r</mi><mi>t</mi></msub>
<mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em">
<mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr>
<mtr><mtd><mtext>Reward</mtext></mtd></mtr>
</mtable></mstyle>
</munder><mo>+</mo>
<munder accentunder="false">
<mi>β</mi>
<mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em">
<mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr>
<mtr><mtd><mtext>Weight</mtext></mtd></mtr>
</mtable></mstyle>
</munder>
<munder accentunder="false">
<mrow><mi>ℋ</mi><mo>(</mo><mi>π</mi><mo>(</mo><mo>·</mo><mo>|</mo><msub><mi>s</mi><mi>t</mi></msub><mo>)</mo><mo>)</mo></mrow>
<mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em">
<mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr>
<mtr><mtd><mtext>Policy entropy</mtext></mtd></mtr>
</mtable></mstyle>
</munder><mo stretchy="false">)</mo>
</mrow><mo minsize="3em" maxsize="3em">]</mo>
</mrow>
</math>
</div>
<p class="mdp-explanation">At every time step, add an entropy bonus to the reward, then sum both terms over time. This favors policies that earn reward while keeping a more diverse action distribution.</p>
<ul class="policy-explanation">
<li><math aria-label="pi of dot given s t"><mi>π</mi><mo>(</mo><mo>·</mo><mo>|</mo><msub><mi>s</mi><mi>t</mi></msub><mo>)</mo></math> is the full distribution over actions at state sₜ. The dot means all possible actions, rather than one particular action.</li>
<li><math aria-label="entropy of pi of dot given s t"><mi>ℋ</mi><mo>(</mo><mi>π</mi><mo>(</mo><mo>·</mo><mo>|</mo><msub><mi>s</mi><mi>t</mi></msub><mo>)</mo><mo>)</mo></math> applies the entropy formula above to that action distribution. Always choosing one action gives entropy 0; spreading probability more evenly among a fixed set of actions gives higher entropy.</li>
<li>β ≥ 0 controls the weight of the entropy bonus. β = 0 recovers the regular objective. A larger β places more weight on entropy relative to reward; the policy still balances both terms.</li>
</ul>
<p class="mdp-explanation">H in the sum is the horizon’s last time index; <math aria-label="script H"><mi>ℋ</mi></math> is the entropy function. They have different meanings.</p>
<h3>Example</h3>
<p class="mdp-explanation">Consider a one-step task with two actions, each giving reward 1. Always choosing the first action gives entropy 0, so the max-ent objective is 1. Choosing each action with probability 0.5 gives entropy 1 bit, so the objective is 1 + β. With β &gt; 0, the second policy is preferred. If the actions have different rewards, their reward difference also affects the choice.</p>
</section>
<section class="value-case entropy-one-step-policy" aria-labelledby="maxent-one-step-policy">
<h2 id="maxent-one-step-policy"><a class="reference-link" href="https://arxiv.org/abs/1801.01290" target="_blank" rel="noopener noreferrer">Max-ent for 1-step problem</a></h2>
<p class="mdp-explanation">The state is fixed and there is one decision. We choose a probability for each action; there are finitely many actions, their rewards are finite, and <math aria-label="beta is positive"><mi>β</mi><mo>&gt;</mo><mn>0</mn></math>.</p>
<p class="mdp-explanation">Here, <math aria-label="log equals natural log"><mo>log</mo><mo>=</mo><mo>ln</mo></math>, so entropy is measured in nats. Using base-2 logarithms changes the scale of <math aria-label="beta"><mi>β</mi></math>.</p>
<ol class="contraction-list">
<li>
<h3>Choose the distribution</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="One-step maximum entropy objective">
<math display="block" aria-label="Maximize over the action probabilities pi of a the expected reward plus beta times the entropy of the action distribution">
<munder><mo>max</mo><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow></munder>
<mrow><mi mathvariant="normal">E</mi><mo>[</mo><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>]</mo><mo>+</mo><mi>β</mi><mi>ℋ</mi><mrow><mo>(</mo><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo><mo>)</mo></mrow></mrow>
</math>
</div>
<p class="mdp-explanation">The unknowns are the action probabilities. The objective balances expected reward with the entropy of their distribution.</p>
</li>
<li>
<h3>Expand reward and entropy</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Expanded one-step objective">
<math display="block" aria-label="Maximize over pi of a the sum over actions of pi of a times r of a, minus beta times the sum over actions of pi of a times log pi of a">
<munder><mo>max</mo><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow></munder>
<munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mrow><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow>
<mo>−</mo><mi>β</mi><munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>log</mo><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow>
</math>
</div>
<p class="mdp-explanation">Expected reward weights each reward by its action probability. Entropy supplies the second sum. Valid probabilities must satisfy:</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Probability constraints">
<math display="block" aria-label="Pi of a is nonnegative for every action, and the sum over actions of pi of a equals one">
<mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>≥</mo><mn>0</mn><mspace width="0.5em" /><mo>∀</mo><mi>a</mi><mo>,</mo><mspace width="1em" />
<munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>=</mo><mn>1</mn>
</math>
</div>
</li>
<li>
<h3>Enforce normalization</h3>
<div class="mdp-equation entropy-long-equation" tabindex="0" role="region" aria-label="Lagrangian objective with the normalization constraint">
<math display="block" aria-label="Maximize over pi of a and minimize over lambda the Lagrangian; equivalently, maximize over pi of a and minimize over lambda the entire expanded reward, entropy, and normalization expression">
<munder><mo>max</mo><mrow><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow></munder><munder><mo>min</mo><mi>λ</mi></munder>
<mrow><mi>ℒ</mi><mo>(</mo><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo><mo>,</mo><mi>λ</mi><mo>)</mo></mrow><mo>=</mo>
<munder><mo>max</mo><mrow><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow></munder><munder><mo>min</mo><mi>λ</mi></munder><mrow><mo>[</mo>
<munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mrow><mi>r</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow>
<mo>−</mo><mi>β</mi><munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>log</mo><mrow><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow>
<mo>+</mo><mi>λ</mi><mrow><mo>(</mo><munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>−</mo><mn>1</mn><mo>)</mo></mrow>
<mo>]</mo></mrow>
</math>
</div>
<p class="mdp-explanation"><math aria-label="lambda"><mi>λ</mi></math> is a constraint multiplier. It can be any real number: if the probabilities do not sum to 1, it can make the Lagrangian arbitrarily negative. When they sum to 1, its added term is 0.</p>
</li>
<li>
<h3>Differentiate one probability</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Stationarity with respect to an action probability">
<math display="block" aria-label="The partial derivative of the Lagrangian with respect to pi of a equals zero">
<mfrac><mrow><mo>∂</mo><mi>ℒ</mi></mrow><mrow><mo>∂</mo><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow></mfrac><mo>=</mo><mn>0</mn>
</math>
</div>
<p class="mdp-explanation">With positive <math aria-label="beta"><mi>β</mi></math>, the optimal probabilities are positive. Set the derivative for each action probability to 0, treating the other probabilities and <math aria-label="lambda"><mi>λ</mi></math> as fixed.</p>
</li>
<li>
<h3>Expand the derivative</h3>
<div class="mdp-equation entropy-long-equation" tabindex="0" role="region" aria-label="Derivative of the entire Lagrangian expression">
<math display="block" aria-label="The partial derivative with respect to pi of a of the entire bracket containing expected reward, minus beta times the sum of pi log pi, and the lambda normalization term equals zero">
<mfrac><mo>∂</mo><mrow><mo>∂</mo><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow></mfrac>
<mrow><mo>[</mo>
<munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mrow><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow>
<mo>−</mo><mi>β</mi><munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>log</mo><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow>
<mo>+</mo><mi>λ</mi><mrow><mo>(</mo><munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>−</mo><mn>1</mn><mo>)</mo></mrow>
<mo>]</mo></mrow><mo>=</mo><mn>0</mn>
</math>
</div>
<p class="mdp-explanation">Vary one component <math aria-label="pi of a"><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></math> of the distribution. Only that action’s term in each sum changes.</p>
</li>
<li>
<h3>Apply the product rule</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Expanded stationarity condition">
<math display="block" aria-label="R of a minus beta log pi of a minus beta plus lambda equals zero; the separate minus beta comes from the product rule">
<mrow><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>−</mo><mi>β</mi><mo>log</mo><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow>
<munder accentunder="false"><mrow><mo>−</mo><mi>β</mi></mrow><mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em"><mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr><mtr><mtd><mtext>Product rule</mtext></mtd></mtr></mtable></mstyle></munder>
<mo>+</mo><mi>λ</mi><mo>=</mo><mn>0</mn>
</math>
</div>
<p class="mdp-explanation">The reward term gives <math aria-label="r of a"><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></math>. Since <math aria-label="The derivative of x log x is log x plus one"><mfrac><mrow><mi mathvariant="normal">d</mi><mrow><mo>(</mo><mi>x</mi><mo>log</mo><mi>x</mi><mo>)</mo></mrow></mrow><mrow><mi mathvariant="normal">d</mi><mi>x</mi></mrow></mfrac><mo>=</mo><mo>log</mo><mi>x</mi><mo>+</mo><mn>1</mn></math>, the entropy term gives both negative terms. The constraint term gives <math aria-label="lambda"><mi>λ</mi></math>.</p>
</li>
<li>
<h3>Isolate the log probability</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Rearranged stationarity condition">
<math display="block" aria-label="Beta times log pi of a equals r of a minus beta plus lambda">
<mi>β</mi><mo>log</mo><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>=</mo><mrow><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>−</mo><mi>β</mi><mo>+</mo><mi>λ</mi>
</math>
</div>
<p class="mdp-explanation">Move the log term to the other side; the same <math aria-label="lambda"><mi>λ</mi></math> applies to every action.</p>
</li>
<li>
<h3>Exponentiate</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Unnormalized exponential action probabilities">
<math display="block" aria-label="Pi of a equals exp of one over beta times r of a minus beta plus lambda">
<mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>=</mo><mo>exp</mo><mrow><mo>[</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mrow><mo>(</mo><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>−</mo><mi>β</mi><mo>+</mo><mi>λ</mi><mo>)</mo></mrow><mo>]</mo></mrow>
</math>
</div>
<p class="mdp-explanation">Divide by <math aria-label="beta"><mi>β</mi></math>, then apply exp to undo the natural logarithm. The factor <math aria-label="C equals exp of lambda minus beta divided by beta"><mi>C</mi><mo>=</mo><mo>exp</mo><mrow><mo>(</mo><mfrac><mrow><mi>λ</mi><mo>−</mo><mi>β</mi></mrow><mi>β</mi></mfrac><mo>)</mo></mrow></math> is common to every action.</p>
</li>
<li>
<h3>Differentiate the multiplier</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Stationarity with respect to lambda">
<math display="block" aria-label="The partial derivative of the Lagrangian with respect to lambda equals zero">
<mfrac><mrow><mo>∂</mo><mi>ℒ</mi></mrow><mrow><mo>∂</mo><mi>λ</mi></mrow></mfrac><mo>=</mo><mn>0</mn>
</math>
</div>
<p class="mdp-explanation">Only the normalization term contains <math aria-label="lambda"><mi>λ</mi></math>, so this derivative recovers the probability constraint.</p>
</li>
<li>
<h3>Make the probabilities sum to 1</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Normalization recovered from the multiplier derivative">
<math display="block" aria-label="The sum over actions of pi of a minus one equals zero">
<munder><mo>∑</mo><mi>a</mi></munder><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>−</mo><mn>1</mn><mo>=</mo><mn>0</mn>
</math>
</div>
<p class="mdp-explanation">Let <math aria-label="Z"><mi>Z</mi></math> be the sum of the exponential reward weights. Each weight is multiplied by the same <math aria-label="C"><mi>C</mi></math>, so their total probability is <math aria-label="C times Z"><mi>C</mi><mi>Z</mi></math>. Normalization gives <math aria-label="C times Z equals one, so C equals one over Z"><mi>C</mi><mi>Z</mi><mo>=</mo><mn>1</mn><mo>⇒</mo><mi>C</mi><mo>=</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac></math>.</p>
</li>
<li>
<h3>Normalize the exponential weights</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Normalized maximum entropy policy">
<math display="block" aria-label="Pi of a equals one over Z times exp of one over beta times r of a">
<mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow><mo>=</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mo>exp</mo><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow>
</math>
</div>
<p class="mdp-explanation">Divide every positive weight by their common total. The result is a valid action distribution, with higher-reward actions receiving higher probability.</p>
</li>
<li>
<h3>Define the partition function</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Partition function for normalization">
<math display="block" aria-label="Z equals the sum over actions of exp of one over beta times r of a">
<mi>Z</mi><mo>=</mo><munder><mo>∑</mo><mi>a</mi></munder><mo>exp</mo><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow>
</math>
</div>
<p class="mdp-explanation"><math aria-label="Z"><mi>Z</mi></math> is the total exponential weight across all actions. It is positive and finite under the stated assumptions.</p>
</li>
</ol>
<p class="mdp-explanation">For two actions with rewards 1 and 0, <math aria-label="beta equals one"><mi>β</mi><mo>=</mo><mn>1</mn></math> gives probabilities approximately 0.731 and 0.269. A smaller <math aria-label="beta"><mi>β</mi></math> favors higher-reward actions more strongly; a larger <math aria-label="beta"><mi>β</mi></math> makes the probabilities more even.</p>
</section>
<section class="value-case entropy-one-step-value" aria-labelledby="maxent-one-step-value">
<h2 id="maxent-one-step-value"><a class="reference-link" href="https://arxiv.org/abs/1702.08165" target="_blank" rel="noopener noreferrer">Max-ent for 1-step problem: value</a></h2>
<p class="mdp-explanation">Use the optimal policy derived above, with <math aria-label="Beta is positive"><mrow><mi>β</mi><mo>&gt;</mo><mn>0</mn></mrow></math> and natural logarithms.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Pi of a equals one over Z times exp of one over beta times r of a, where Z equals the sum over actions a of exp of one over beta times r of a">
<math display="block" aria-label="Pi of a equals one over Z times exp of one over beta times r of a, where Z equals the sum over actions a of exp of one over beta times r of a"><mrow><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>=</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mo>,</mo><mspace width="0.6em" /><mi>Z</mi><mo>=</mo><munder><mo>∑</mo><mi>a</mi></munder><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow></mrow></math>
</div>
<div class="mdp-equation entropy-long-equation" tabindex="0" role="region" aria-label="V equals the sum over actions a of one over Z times exp of one over beta times r of a times r of a, minus beta times the sum over actions a of one over Z times exp of one over beta times r of a times log of that probability">
<math display="block" aria-label="V equals the sum over actions a of one over Z times exp of one over beta times r of a times r of a, minus beta times the sum over actions a of one over Z times exp of one over beta times r of a times log of that probability"><mrow><mi>V</mi><mo>=</mo><munder><mo>∑</mo><mi>a</mi></munder><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>−</mo><mi>β</mi><munder><mo>∑</mo><mi>a</mi></munder><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mi>log</mi><mrow><mo>[</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mo>]</mo></mrow></mrow></math>
</div>
<p class="mdp-explanation">Substitute the optimal policy into the reward plus entropy objective. V is the maximized value of this full objective.</p>
<div class="mdp-equation entropy-long-equation" tabindex="0" role="region" aria-label="V equals the sum over actions a of the optimal probability times the quantity r of a minus beta times log of exp of one over beta times r of a, minus beta times the sum over actions a of the optimal probability times log of one over Z">
<math display="block" aria-label="V equals the sum over actions a of the optimal probability times the quantity r of a minus beta times log of exp of one over beta times r of a, minus beta times the sum over actions a of the optimal probability times log of one over Z"><mrow><mi>V</mi><mo>=</mo><munder><mo>∑</mo><mi>a</mi></munder><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mrow><mo>[</mo><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>−</mo><mi>β</mi><mi>log</mi><mrow><mo>(</mo><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mo>)</mo></mrow><mo>]</mo></mrow><mo>−</mo><mi>β</mi><munder><mo>∑</mo><mi>a</mi></munder><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow><mi>log</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mo>)</mo></mrow></mrow></math>
</div>
<p class="mdp-explanation">Use <math aria-label="Log of u v equals log of u plus log of v"><mrow><mi>log</mi><mrow><mo>(</mo><mi>u</mi><mi>v</mi><mo>)</mo></mrow><mo>=</mo><mi>log</mi><mrow><mo>(</mo><mi>u</mi><mo>)</mo></mrow><mo>+</mo><mi>log</mi><mrow><mo>(</mo><mi>v</mi><mo>)</mo></mrow></mrow></math> to split the logarithm, then combine the reward with the exponential term. <math aria-label="Log of one over Z"><mrow><mi>log</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mo>)</mo></mrow></mrow></math> is the same for every action.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="V equals zero minus beta times log of one over Z times the sum over actions a of one over Z times exp of one over beta times r of a">
<math display="block" aria-label="V equals zero minus beta times log of one over Z times the sum over actions a of one over Z times exp of one over beta times r of a; the reward terms cancel to zero and the sum of probabilities equals one"><mrow><mi>V</mi><mo>=</mo><munder accentunder="false"><mn>0</mn><mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em"><mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr><mtr><mtd><mtext>Reward terms cancel</mtext></mtd></mtr></mtable></mstyle></munder><mo>−</mo><mi>β</mi><mi>log</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mo>)</mo></mrow><munder accentunder="false"><mrow><munder><mo>∑</mo><mi>a</mi></munder><mfrac><mn>1</mn><mi>Z</mi></mfrac><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow></mrow><mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em"><mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr><mtr><mtd><mtext>Probabilities sum to 1</mtext></mtd></mtr></mtable></mstyle></munder></mrow></math>
</div>
<p class="mdp-explanation">The logarithm reverses the exponential, so <math aria-label="R of a minus beta times one over beta times r of a equals zero"><mrow><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>−</mo><mi>β</mi><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>=</mo><mn>0</mn></mrow></math>. Pull the remaining constant factor outside the sum.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="V equals minus beta times log of one over Z">
<math display="block" aria-label="V equals minus beta times log of one over Z"><mrow><mi>V</mi><mo>=</mo><mo>−</mo><mi>β</mi><mi>log</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mo>)</mo></mrow></mrow></math>
</div>
<p class="mdp-explanation">The probabilities are normalized: <math aria-label="The sum over actions a of pi of a equals one"><mrow><munder><mo>∑</mo><mi>a</mi></munder><mi>π</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>=</mo><mn>1</mn></mrow></math>. The remaining sum therefore becomes 1.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="V equals beta times log of the sum over actions a of exp of one over beta times r of a">
<math display="block" aria-label="V equals beta times log of the sum over actions a of exp of one over beta times r of a"><mrow><mi>V</mi><mo>=</mo><mi>β</mi><mi>log</mi><munder><mo>∑</mo><mi>a</mi></munder><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><mi>r</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow><mo>)</mo></mrow></mrow></math>
</div>
<p class="mdp-explanation">Use <math aria-label="Log of one over Z equals minus log of Z"><mrow><mi>log</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac><mo>)</mo></mrow><mo>=</mo><mo>−</mo><mi>log</mi><mrow><mo>(</mo><mi>Z</mi><mo>)</mo></mrow></mrow></math>, then replace Z with its defining sum.</p>
<p class="mdp-explanation">The policy probabilities form a softmax distribution. V is a scalar soft maximum: beta times log-sum-exp.</p>
</section>
<section class="value-case entropy-value-iteration" aria-labelledby="maxent-value-iteration">
<h2 id="maxent-value-iteration"><a class="reference-link" href="https://arxiv.org/abs/1702.08165" target="_blank" rel="noopener noreferrer">Max-ent Value Iteration</a></h2>
<p class="mdp-explanation">For a finite action set, the one-step solution also works when an action’s score includes the future. Here, β &gt; 0 and log means the natural logarithm.</p>
<h3>1. Count the remaining decision steps</h3>
<p class="mdp-explanation">Here k counts remaining decision steps; H is the last time index. A sum from time 0 through H contains H + 1 decisions. With no decisions left, the value is 0:</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The value with no decision steps remaining is zero">
<math display="block" aria-label="V zero of s equals zero">
<msub><mi>V</mi><mn>0</mn></msub><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow><mo>=</mo><mn>0</mn>
</math>
</div>
<p class="mdp-explanation">Vₖ(s) is the best expected total reward plus entropy bonus over the last k decisions, starting from state s:</p>
<div class="mdp-equation entropy-long-equation" tabindex="0" role="region" aria-label="The optimal maximum entropy value with k remaining decision steps">
<math display="block" aria-label="V k of s equals the maximum over policies pi of the expected sum from t equals H minus k plus one through H of r of s t and a t plus beta times the entropy of the full action distribution pi given s t, conditional on s at time H minus k plus one equaling s">
<msub><mi>V</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow><mo>=</mo>
<munder><mo>max</mo><mi>π</mi></munder><mi mathvariant="normal">E</mi><mo>[</mo>
<munderover><mo>∑</mo><mrow><mi>t</mi><mo>=</mo><mi>H</mi><mo>−</mo><mi>k</mi><mo>+</mo><mn>1</mn></mrow><mi>H</mi></munderover>
<mrow><mo>(</mo><mi>r</mi><mo>(</mo><msub><mi>s</mi><mi>t</mi></msub><mo>,</mo><msub><mi>a</mi><mi>t</mi></msub><mo>)</mo><mo>+</mo><mi>β</mi><mi>ℋ</mi><mo>(</mo><mi>π</mi><mo>(</mo><mo>·</mo><mo>|</mo><msub><mi>s</mi><mi>t</mi></msub><mo>)</mo><mo>)</mo><mo>)</mo>
</mrow><mo>|</mo><msub><mi>s</mi><mrow><mi>H</mi><mo>−</mo><mi>k</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>=</mo><mi>s</mi><mo>]</mo>
</math>
</div>
<p class="mdp-explanation">The expectation averages over action choices and environment transitions. ℋ applies to the full action distribution, and its bonus is included at every decision. In the sum, π denotes a sequence of policies, with π(·|sₜ) meaning the distribution used at time t. Each policy can depend on how many steps remain.</p>
<p class="mdp-explanation">Separate the current decision from the remaining k − 1 decisions:</p>
<div class="mdp-equation entropy-long-equation" tabindex="0" role="region" aria-label="Maximum entropy value recursion">
<math display="block" aria-label="V k of s equals the maximum over pi of the expectation of r of s and a plus beta times the entropy of pi given s plus V k minus one of the next state s prime">
<msub><mi>V</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow><mo>=</mo>
<munder><mo>max</mo><mi>π</mi></munder><mi mathvariant="normal">E</mi><mo>[</mo>
<mi>r</mi><mrow><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow><mo>+</mo><mi>β</mi><mi>ℋ</mi><mrow><mo>(</mo><mi>π</mi><mo>(</mo><mo>·</mo><mo>|</mo><mi>s</mi><mo>)</mo><mo>)</mo></mrow><mo>+</mo>
<msub><mi>V</mi><mrow><mi>k</mi><mo>−</mo><mn>1</mn></mrow></msub><mrow><mo>(</mo><msup><mi>s</mi><mo>′</mo></msup><mo>)</mo></mrow><mo>]</mo>
</math>
</div>
<p class="mdp-explanation">Sample a from π(·|s), then sample s′ from P(·|s,a). The three terms are immediate reward, current entropy bonus, and the best future soft value. Starting from V₀, compute this for k = 1, …, H + 1. This formulation is undiscounted.</p>
<h3>2. Define the action value</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The action value combines immediate reward and optimal future soft value">
<math display="block" aria-label="Q k of s and a equals the expected r of s and a plus V k minus one of the next state, conditional on s and a">
<msub><mi>Q</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow><mo>=</mo>
<mi mathvariant="normal">E</mi><mo minsize="2.2em" maxsize="2.2em">[</mo><munder accentunder="false"><mrow><mi>r</mi><mrow><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow></mrow><mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em"><mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr><mtr><mtd><mtext>Reward now</mtext></mtd></mtr></mtable></mstyle></munder><mo>+</mo>
<munder accentunder="false"><mrow><msub><mi>V</mi><mrow><mi>k</mi><mo>−</mo><mn>1</mn></mrow></msub><mrow><mo>(</mo><msup><mi>s</mi><mo>′</mo></msup><mo>)</mo></mrow></mrow><mstyle scriptlevel="0"><mtable class="proof-label" rowspacing="0.1em"><mtr><mtd><mo class="proof-arrow">↑</mo></mtd></mtr><mtr><mtd><mtext>Future soft value</mtext></mtd></mtr></mtable></mstyle></munder><mo>|</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo minsize="2.2em" maxsize="2.2em">]</mo>
</math>
</div>
<p class="mdp-explanation">Fix the current state and action. Only the next-state transition is averaged here. Qₖ includes the current reward and optimal future rewards and entropy bonuses; it leaves the current entropy bonus outside.</p>
<h3>3. Reuse the one-step solution</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The remaining optimization is the one-step entropy problem with Q as the action score">
<math display="block" aria-label="V k of s equals the maximum over pi of the expectation over actions of Q k of s and a plus beta times the entropy of pi given s">
<msub><mi>V</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow><mo>=</mo>
<munder><mo>max</mo><mi>π</mi></munder><mi mathvariant="normal">E</mi><mo>[</mo>
<msub><mi>Q</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow><mo>+</mo><mi>β</mi><mi>ℋ</mi><mrow><mo>(</mo><mi>π</mi><mo>(</mo><mo>·</mo><mo>|</mo><mi>s</mi><mo>)</mo><mo>)</mo></mrow><mo>]</mo>
</math>
</div>
<p class="mdp-explanation">Here max over π chooses the current state’s action distribution; the future is already optimized by Vₖ₋₁. The expectation is only over a ∼ π(·|s). The entropy is the same for every sampled action. Replace the reward score r(s,a) in the one-step derivation with Qₖ(s,a): it already accounts for the optimal soft future.</p>
<h3>4. Compute the soft value and policy</h3>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The optimal soft value is beta times log sum exp of the action values divided by beta">
<math display="block" aria-label="V k of s equals beta times the natural logarithm of the sum over actions a of exp of one divided by beta times Q k of s and a">
<msub><mi>V</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow><mo>=</mo><mi>β</mi><mi mathvariant="normal">log</mi>
<munder><mo>∑</mo><mi>a</mi></munder><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><msub><mi>Q</mi><mi>k</mi></msub><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo><mo>)</mo>
</mrow></math>
</div>
<p class="mdp-explanation">Ordinary value iteration uses the largest action value, maxₐ Qₖ(s,a). The entropy bonus replaces that maximum with log-sum-exp, which combines all action scores.</p>
<div class="mdp-equation" tabindex="0" role="region" aria-label="The optimal policy is a normalized exponential distribution over action values">
<math display="block" aria-label="Pi k of a given s equals one divided by Z times exp of one divided by beta times Q k of s and a">
<msub><mi>π</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>a</mi><mo>|</mo><mi>s</mi><mo>)</mo></mrow><mo>=</mo><mfrac><mn>1</mn><mi>Z</mi></mfrac>
<mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><msub><mi>Q</mi><mi>k</mi></msub><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo><mo>)</mo>
</mrow></math>
</div>
<div class="mdp-equation" tabindex="0" role="region" aria-label="Z normalizes the action probabilities and equals the exponential of the soft value divided by beta">
<math display="block" aria-label="Z equals the sum over actions of exp of one divided by beta times Q k of s and a, which equals exp of V k of s divided by beta">
<mi>Z</mi><mo>=</mo><munder><mo>∑</mo><mi>a</mi></munder><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mn>1</mn><mi>β</mi></mfrac><msub><mi>Q</mi><mi>k</mi></msub><mo>(</mo><mi>s</mi><mo>,</mo><mi>a</mi><mo>)</mo><mo>)</mo>
</mrow><mo>=</mo><mi mathvariant="normal">exp</mi><mrow><mo>(</mo><mfrac><mrow><msub><mi>V</mi><mi>k</mi></msub><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow></mrow><mi>β</mi></mfrac><mo>)</mo>
</mrow></math>
</div>
<p class="mdp-explanation">Z depends on s and k and makes the action probabilities sum to 1. Instead of an argmax action, πₖ is a softmax distribution: higher Qₖ gets higher probability. Use πₖ when k steps remain; after one action, k decreases by 1. As β approaches 0 from above, the policy concentrates on greedy actions.</p>
</section>
</main>
</body>
</html>