Leplanner / code /Offline_Goal-Directed Controller Loss.md
nottygian's picture
Push code package
872cf4d verified
|
Raw History Blame Contribute Delete
31.1 kB

The attached setup changes one important assumption: Push-T in LeWM uses continuous actions, grouped into five-action blocks, so a discrete action head or Gumbel-Softmax is not appropriate here. I’ll align the controller with the frozen encoder/predictor contract and keep the objective focused on latent goal reaching plus the minimum safeguards needed for offline training.The attached paper changes the design materially: LeWM Push-T uses continuous actions, not discrete actions. It also groups five environment actions into one world-model transition. Therefore, remove the discrete action classifier, Gumbel-Softmax, entropy loss, and action-embedding expectation from the earlier proposal.Your controller should learn an amortized iterative MPC solver: it predicts and repeatedly improves a continuous action plan using the frozen LeWM predictor.ObjectiveGiven the current observation history and a goal image, produce a plan that minimizes the same latent terminal distance used by LeWM:$$\min_\theta \left|\hat x_H-x_G\right|2^2$$The controller is not trained to reproduce dataset actions. Dataset trajectories provide current-state/goal pairs and an action-support model, while controller gradients come from the consequences predicted by LeWM.This means the learned controller can discover a different valid action sequence from the one in the dataset.1. Fixed LeWM componentsUse the pretrained LeWM:$$x_t=E(o_t)$$$$\hat x{t+1}=M(x_{t-N+1:t},b_t)$$where: $E$ is the frozen image encoder; $M$ is the frozen latent predictor; $N=3$ for Push-T; $b_t$ is one block of five continuous environment actions. Let the native environment action dimension be $d_a$. Do not hard-code it before checking the checkpoint and dataset. One action block has shape$$b_t\in\mathbb R^{5\times d_a}.$$With the paper's planning horizon $H=5$, the full controller plan has shape$$A=\left(b_0,\ldots,b_{H-1}\right) \in\mathbb R^{H\times5\times d_a}.$$This represents 25 environment actions.Keep both pretrained networks frozen. The predictor must remain differentiable with respect to its action input:python .spinner { transform-origin: center; overflow: visible; animation: spinner-rotate var(--spinner-speed) linear infinite; } @keyframes spinner-rotate { to { transform: rotate(360deg); } } Run CodeRunning...Complete!FailedCopy codeCopied!for parameter in encoder.parameters(): parameter.requires_grad_(False)

for parameter in predictor.parameters(): parameter.requires_grad_(False)

encoder.eval() predictor.eval()

      .spinner {
        transform-origin: center;
        overflow: visible;
        animation: spinner-rotate var(--spinner-speed) linear infinite;
      }
      @keyframes spinner-rotate {
        to { transform: rotate(360deg); }
      }
    You may use torch.no_grad() when encoding real images, but not around predictor rollouts during controller training. Otherwise the controller receives no gradient.2. Controller architectureUse two small trainable transformers with shared parameters across refinement iterations.InputsEncode:$$C_0=(x_{t-2},x_{t-1},x_t),\qquad x_G=E(o_G).$$The goal should be a future observation from the same offline trajectory, initially one to five latent transitions ahead.Action-plan representationRepresent every action block with a hidden plan token:$$Y^{(k)}=(y_0^{(k)},\ldots,y_{H-1}^{(k)}),

\qquad y_j^{(k)}\in\mathbb R^{d_{\text{model}}}.$$A small MLP converts each token into five continuous actions:$$u_j^{(k)}=W_a y_j^{(k)}+b_a,$$$$b_j^{(k)}

a_{\text{center}} + a_{\text{scale}}\tanh(u_j^{(k)}).$$The tanh transformation enforces Push-T action bounds directly. There is then no need for an action-bound penalty.Initial planInitialize the plan using learned plan-query tokens conditioned on the current history and goal:$$Y^{(0)}

G_\theta \left( Y_{\text{query}}, C_0, x_G \right).$$Avoid initializing every training plan from exactly zero because this can produce overly symmetric updates. Learned queries are preferable.Frozen consequence rolloutAt refinement iteration $k$, map $Y^{(k)}$ to action blocks and roll them out through LeWM:$$C_0^{(k)}=C_0,$$$$\hat x_{j+1}^{(k)}

M\left(C_j^{(k)},b_j^{(k)}\right),$$$$C_{j+1}^{(k)}

\operatorname{shift} \left( C_j^{(k)},\hat x_{j+1}^{(k)} \right).$$This gives$$\hat X^{(k)}

\left( \hat x_1^{(k)},\ldots,\hat x_H^{(k)} \right).$$The frozen LeWM predictor is the actual consequence model. Do not train another transformer to duplicate next-latent prediction.Consequence transformerYour trainable consequence transformer should interpret LeWM's rollout rather than independently predict it:$$R^{(k)}

F_\theta \left( C_0, x_G, Y^{(k)}, \hat X^{(k)} \right).$$For each horizon position, provide tokens containing projections of:$$\left[ y_j^{(k)}, \hat x_{j+1}^{(k)}, \hat x_{j+1}^{(k)}-x_G, \left|\hat x_{j+1}^{(k)}-x_G\right|_2^2 \right].$$These consequence tokens tell the action transformer what the current plan is predicted to cause.Action-refinement transformerThe second transformer outputs corrections to the plan tokens:$$\Delta Y^{(k)}

G_\theta \left( Y^{(k)},R^{(k)},C_0,x_G \right),$$$$Y^{(k+1)}

Y^{(k)} + \eta_k\Delta Y^{(k)}.$$Use learned or fixed update scales $\eta_k$, constrained to a reasonable range. Share $F_\theta$ and $G_\theta$ across all refinement iterations. That gives the recursive TRM-like behavior you want without increasing parameter count with $K$.A practical starting configuration is: world-model horizon $H=5$; internal refinements $K=3$; controller width $256$; 4 transformer layers in each module; 4 or 8 attention heads; pre-layer normalization; dropout $0.1$. Keep $H$ and $K$ conceptually separate: $H$ is imagined environment time, while $K$ is controller reasoning time.3. Main goal lossLeWM was explicitly evaluated using Euclidean latent distance, so use that same metric. Since SIGReg makes each latent coordinate approximately unit scale, divide by latent dimension $D$:$$d_j^{(k)}

\frac{1}{D} \left| \hat x_j^{(k)}-x_G \right|2^2.$$The principal loss for refinement $k$ is:$$\mathcal L{\mathrm{terminal}}^{(k)}

d_H^{(k)}.$$This is the most important term and exactly matches the planning objective from the paper.Because you requested cumulative horizon supervision, add a small late-weighted path loss:$$\mathcal L_{\mathrm{path}}^{(k)}

\sum_{j=1}^{H-1} w_jd_j^{(k)},$$where$$w_j

\frac{(j/H)^2} {\sum_{q=1}^{H-1}(q/H)^2}.$$Then:$$\mathcal L_{\mathrm{goal}}^{(k)}

d_H^{(k)} + \alpha\mathcal L_{\mathrm{path}}^{(k)}.$$Start with:$$\alpha=0.05.$$Keep this coefficient small because Push-T may require temporary detours or contact configurations that do not monotonically reduce latent goal distance. If performance is worse than terminal-only training, set $\alpha=0$. Do not add a separate progress loss because it is redundant with the path term and can encourage greedy behavior.4. Iterative refinement lossApply the outcome loss after every refinement, with more weight on later iterations:$$\mathcal L_{\mathrm{refine}}

\frac{ \sum_{k=0}^{K}\rho_k \mathcal L_{\mathrm{goal}}^{(k)} }{ \sum_{k=0}^{K}\rho_k }, \qquad \rho_k=2^k.$$This has two benefits: early plans remain usable if computation is stopped early; every refinement receives a direct learning signal. Use full backpropagation through all $K$ refinements initially. If memory becomes a problem, use activation checkpointing around frozen predictor rollouts before considering truncated gradients.5. Offline action-support lossPure optimization through a frozen offline world model can exploit prediction errors. The controller may generate bounded actions that are valid numerically but poorly represented in the Push-T dataset.Train a separate continuous behavior-density model$$\beta_\psi(b\mid C)$$on real LeWM latent histories and real five-action blocks. A small conditional Gaussian-mixture model is sufficient. This model is not the controller.Calculate per-action-dimension negative log likelihood:$$r(C,b)

-\frac{1}{5d_a} \log\beta_\psi(b\mid C).$$Let $c_{95}$ be the 95th percentile of $r(C,b)$ on held-out real dataset transitions. Penalize only plans outside this support:$$\mathcal L_{\mathrm{support}}

\frac{1}{KH} \sum_{k=1}^{K} \sum_{j=0}^{H-1} \left[ \max \left( 0, r(C_j^{(k)},b_j^{(k)})-c_{95} \right) \right]^2.$$This is deliberately not ordinary behavioral cloning: no penalty is applied inside the broad dataset-supported region; the controller is not asked to reproduce the recorded action; any supported plan that reaches the goal can be selected. A full KL penalty against $\beta$ would pull the controller much more strongly toward behavior cloning, so the thresholded support loss is preferable.6. Final controller lossThe simplified recommended objective is:$$\boxed{ \mathcal L_{\mathrm{controller}}

\mathcal L_{\mathrm{refine}} + \lambda_{\mathrm{support}} \mathcal L_{\mathrm{support}} }$$or explicitly:$$\boxed{ \mathcal L_{\mathrm{controller}}

\frac{ \sum_{k=0}^{K} 2^k \left[ d_H^{(k)} + \alpha \sum_{j=1}^{H-1}w_jd_j^{(k)} \right] }{ \sum_{k=0}^{K}2^k } + \lambda_{\mathrm{support}} \mathcal L_{\mathrm{support}}. }$$Start with:$$\alpha=0.05, \qquad \lambda_{\mathrm{support}}=0.01.$$The loss terms need comparable empirical scales, so inspect their batch means before fixing $\lambda_{\mathrm{support}}$. The desired behavior is that the support term is nearly zero for most plans and activates on clear outliers.You do not need: dataset-action MSE; discrete-action cross entropy; Gumbel-Softmax; entropy regularization; a success classifier; progress loss; action cost; a second learned world model; SIGReg on the controller. SIGReg was needed to train the LeWM encoder. The encoder is now frozen, so applying SIGReg to controller internals has no clear purpose.7. Offline training dataFrom every Push-T trajectory, sample:$$(o_{t-2},o_{t-1},o_t,o_g).$$Choose $g$ from a future point in the same trajectory. Use a curriculum: Start with goals 1-2 latent transitions ahead. Increase to 1-3 transitions. Finish with 1-5 transitions, corresponding to at most 25 environment steps. This guarantees that sampled goals are reachable under the dataset dynamics while never telling the controller which actions were used to reach them.Cache frozen encoder outputs if image augmentation is not being used. This substantially reduces controller-training cost.8. Training procedureFor each minibatch: Encode the three-frame history and goal using frozen $E$. Initialize $Y^{(0)}$. Produce bounded action blocks. Roll out the plan through frozen $M$. Calculate consequence tokens. Refine the plan using the action transformer. Repeat for $K$ iterations. Calculate latent goal losses after each iteration. Add the thresholded support penalty. Backpropagate only into the two controller transformers and action head. A sensible starting setup is: AdamW with learning rate $10^{-4}$; weight decay $10^{-4}$; gradient clipping at $1.0$; batch size $16$ or $32$; $K=3$; horizon curriculum ending at $H=5$; mixed precision; activation checkpointing through LeWM rollouts. 9. Execution with MPCAt test time: Encode the latest three observations. Run $K$ controller refinements. Obtain five predicted action blocks. Execute only the first five-action block. Observe the resulting frames. Update the three-frame history and replan. Replanning after one block is more robust than executing the entire 25-action plan, especially because autoregressive latent errors increase with horizon.10. Evaluation and monitoringTraining loss alone is insufficient because the controller and its objective use the same frozen model. Monitor:$$\text{TerminalDistance}

\frac{1}{D}|\hat x_H-x_G|^2,$$$$\text{RefinementGain}_k

d_H^{(k-1)}-d_H^{(k)},$$$$\text{SupportViolationRate}

\Pr[r(C,b)>c_{95}].$$Also compare against the paper's CEM planner:$$\text{CEMRegret}

d_H^{\mathrm{controller}}-d_H^{\mathrm{CEM}}.$$Most importantly, periodically execute the controller in the real Push-T simulator and measure: environment success rate; final physical block position and angle error; success versus goal distance; success after each refinement count; inference time versus CEM; success when replanning every block. Use the same held-out initial-state/goal pairs for controller and CEM. Do not monitor agreement with dataset actions because alternate valid plans are explicitly allowed.11. Necessary baselines and ablationsFor a defensible research result, compare: pretrained LeWM + CEM; direct one-pass controller, $K=0$; iterative controller, $K=1,2,3,5$; terminal-only loss; terminal plus path loss; with and without support regularization; execute one block versus execute full plan; controller initialized randomly versus CEM warm-started, if later explored. The strongest claim available from this design is not guaranteed optimal control. It is: The controller amortizes LeWM latent-space trajectory optimization into a small fixed number of learned consequence-guided refinement steps. Its output is the best plan found under the frozen LeWM approximation and finite horizon, not a mathematically guaranteed globally optimal action. The main research comparison is whether it approaches CEM success while requiring substantially fewer world-model evaluations.

My assessmentThis is a real and conceptually clean finding. Your diagnosis is convincing.The result is not actually contrary to MPC theory. Receding-horizon MPC behaves well only when its objective contains appropriate stage costs, terminal constraints, terminal value functions, or another mechanism ensuring progress. Your controller currently solves:$$\min_{a_{1:H}} d(\hat z_H,z_G),$$which means: “Reach the goal exactly $H$ blocks from now.” After executing one block, replanning resets the deadline back to $H$. The controller therefore keeps postponing arrival. It has learned a good fixed-horizon open-loop planner, but not yet a time-consistent closed-loop goal policy.The full-plan result is exactly consistent with that interpretation: Full execution: the original five-block deadline is preserved, so the plan succeeds. Replanning every block: the deadline is repeatedly shifted forward, so the controller approaches asymptotically. More execution budget does not help: strong evidence for a closed-loop fixed point rather than ordinary slow behavior. I would call this receding-horizon procrastination or horizon-reset procrastination rather than literally Zeno behavior, although “Zeno-like” is a good intuitive description.1. What the results establishYour main conclusions are currently:Refinement is useful$$36%\rightarrow50%\rightarrow52%$$from $K=0$ to $K=2$.That supports the central architectural hypothesis: imagined-consequence refinement improves the initial plan.Refinement saturates quickly$K=2$ and $K=3$ both produce $52%$, while $K=5$ drops to $42%$.Because $K=5$ exceeds the trained unroll depth, this does not show that additional reasoning is fundamentally harmful. It shows that the learned recurrent update is not guaranteed to be contractive outside its training depth.At present, $K=2$ is the Pareto-optimal receding-horizon configuration: equal success to $K=3$ with fewer predictor rows. You should test full-plan execution for every $K$, however, before selecting the final configuration.Open-loop execution matches the training objectiveThe $88%$ result is not a suspicious outlier anymore. It shows that the controller has successfully learned to produce a complete five-block plan.The existing objective is dynamically inconsistent under replanningThe one-block-away profile is especially strong evidence:$$[0.090,;0.053,;0.035,;0.023,;0.014].$$For a goal that is already reachable after one block, the learned planner nevertheless reserves most of the progress for later blocks. That is precisely the behavior induced by fixed terminal-horizon training.2. I would not use min-over-blocks aloneA loss such as$$\mathcal L_{\min}

\min_{1\le j\le H}d_j$$would eliminate the requirement to arrive specifically at block five, but it introduces new problems: The controller can touch the goal at one block and immediately leave. Only the minimum-distance block receives useful gradient. The identity of the minimum block can switch abruptly during training. It does not explicitly reward stable goal occupancy. It may produce plans whose final state is far from the goal. A soft minimum is smoother,$$\operatorname{softmin}{\tau}(d{1:H})

-\tau\log \sum_{j=1}^{H}\exp(-d_j/\tau),$$but still rewards “touch once and leave.”Therefore, I recommend the arrival-and-hold objective, using the known temporal offset of every relabeled goal.3. Recommended new loss: horizon-matched arrival and holdFor every offline training sample, you already know how far in the future its goal was selected.Let$$q_i\in{1,\ldots,H}$$be the number of world-model blocks between the current state and goal for sample $i$.For example: goal five environment actions ahead: $q_i=1$; goal ten actions ahead: $q_i=2$; goal twenty-five actions ahead: $q_i=5$. For refinement iteration $k$, define:$$d_{i,j}^{(k)}

\frac{1}{D} \left| \hat z_{i,j}^{(k)}-z_{G_i} \right|2^2.$$Replace the old fixed terminal loss$$d{i,H}^{(k)}$$with:$$\boxed{ \mathcal J_i^{(k)}

d_{i,q_i}^{(k)} + \lambda_{\mathrm{hold}} \frac{ \sum_{j=q_i+1}^{H}d_{i,j}^{(k)} }{ \max(1,H-q_i) }. }$$When $q_i=H$, define the hold term as zero.InterpretationThe first term says: Reach the goal no later than its demonstrated reachable deadline. The second says: Once the deadline has been reached, remain close to the goal for the rest of the plan. It does not supervise dataset actions. The controller can use any action sequence that reaches and maintains the desired latent state.It also does not prevent earlier arrival. If the controller reaches the goal before $q_i$ and stays there, it still achieves a low loss at and after $q_i$.Example: one-block-away goalFor $q_i=1$,$$\mathcal J_i^{(k)}

d_{i,1}^{(k)} + \frac{\lambda_{\mathrm{hold}}}{4} \left( d_{i,2}^{(k)} +d_{i,3}^{(k)} +d_{i,4}^{(k)} +d_{i,5}^{(k)} \right).$$Now, deliberate one-fifth progress is strongly penalized because the controller must already be at the goal after the first block.The desired learned profile should look more like:$$[0.015,;0.014,;0.014,;0.015,;0.014]$$instead of decreasing only at the fifth block.Example: five-block-away goalFor $q_i=5$,$$\mathcal J_i^{(k)}=d_{i,5}^{(k)}.$$This retains the original objective for the longest-distance examples.During training, the model also sees intermediate state-goal pairs from the same trajectory with $q=4,3,2,1$. Therefore, as MPC moves closer to the goal, the controller should implicitly recognize that less time remains and produce increasingly direct plans.Do not pass $q_i$ to the controller as an input initially. Use it only to index the training loss. If you condition the controller on a deadline and always reset that deadline to five during evaluation, the procrastination problem returns.4. Remove the current late-weighted path loss initiallyYour previous objective included:$$\alpha\sum_{j=1}^{H-1}w_jd_j,$$with higher weights on later blocks. Although $\alpha$ is small, this does not solve the horizon-reset issue and slightly reinforces late arrival.For the first corrected experiment, use only:$$\mathcal J_i^{(k)}

d_{i,q_i}^{(k)} + \lambda_{\mathrm{hold}}\overline d_{i,q_i+1:H}^{(k)}.$$Start with:$$\lambda_{\mathrm{hold}}=0.5.$$Test:$$\lambda_{\mathrm{hold}}\in{0,;0.25,;0.5,;1.0}.$$This ablation has clear interpretations: $0$: arrival deadline only; $0.25$: weak retention; $0.5$: balanced; $1.0$: arrival and retention have similar importance. I would avoid adding progress, soft-minimum, success, and path terms simultaneously. First determine whether the direct horizon-matched correction fixes MPC.5. Complete revised controller objectiveRetain deep supervision over controller refinements:$$\mathcal L_{\mathrm{goal}}

\frac{ \sum_{k=0}^{K}\rho_k \left[ \frac1B\sum_{i=1}^{B}\mathcal J_i^{(k)} \right] }{ \sum_{k=0}^{K}\rho_k }, \qquad \rho_k=2^k.$$Then keep your support penalty:$$\boxed{ \mathcal L_{\mathrm{controller}}

\mathcal L_{\mathrm{goal}} + \lambda_{\mathrm{support}} \mathcal L_{\mathrm{support}}. }$$Explicitly:$$\boxed{ \begin{aligned} \mathcal L_{\mathrm{controller}} ={}& \frac{1}{\sum_k2^k} \sum_{k=0}^{K}2^k \frac1B \sum_{i=1}^{B} \Bigg[ d_{i,q_i}^{(k)} \ &\quad+ \lambda_{\mathrm{hold}} \mathbf 1[q_i<H] \frac{1}{H-q_i} \sum_{j=q_i+1}^{H}d_{i,j}^{(k)} \Bigg] \ &+ \lambda_{\mathrm{support}} \mathcal L_{\mathrm{support}}. \end{aligned} }$$My initial setting would be:$$K_{\mathrm{train}}=3,\qquad \lambda_{\mathrm{hold}}=0.5, \qquad \lambda_{\mathrm{support}}=\text{current value}.$$Do not modify the architecture yet. This is an objective problem, and changing both architecture and loss would make the result harder to interpret.6. Pseudocode for the changed partpython .spinner { transform-origin: center; overflow: visible; animation: spinner-rotate var(--spinner-speed) linear infinite; } @keyframes spinner-rotate { to { transform: rotate(360deg); } } Run CodeRunning...Complete!FailedCopy codeCopied!def horizon_matched_goal_loss( rollout_latents, # [B, K+1, H, D] goal_latent, # [B, D] goal_offset, # [B], values 1,...,H refinement_weights, # [K+1] hold_weight=0.5, ): B, num_refinements, H, D = rollout_latents.shape

# [B, K+1, H]
distances = (
    (rollout_latents - goal_latent[:, None, None, :])
    .pow(2)
    .mean(dim=-1)
)

total = 0.0
weight_sum = 0.0

for k in range(num_refinements):
    batch_loss = 0.0

    for i in range(B):
        # Convert q in {1,...,H} to zero-based index.
        q = int(goal_offset[i].item())
        arrival = distances[i, k, q - 1]

        if q < H:
            hold = distances[i, k, q:].mean()
        else:
            hold = 0.0

        batch_loss = (
            batch_loss
            + arrival
            + hold_weight * hold
        )

    batch_loss = batch_loss / B

    refinement_weight = refinement_weights[k]
    total = total + refinement_weight * batch_loss
    weight_sum += refinement_weight

return total / weight_sum

      .spinner {
        transform-origin: center;
        overflow: visible;
        animation: spinner-rotate var(--spinner-speed) linear infinite;
      }
      @keyframes spinner-rotate {
        to { transform: rotate(360deg); }
      }
    A vectorized version is preferable for actual training, but this illustrates the indexing clearly.7. Balance goal offsets during samplingMake sure the training batch does not accidentally contain mostly $q=4$ or $q=5$ examples.I recommend sampling:$$q\sim \operatorname{Uniform}\{1,2,3,4,5\}$$and then sampling a valid trajectory position for that $q$.Each minibatch should contain roughly equal numbers of:

one-block goals; two-block goals; three-block goals; four-block goals; five-block goals. Otherwise, the terminal-at-five bias may survive simply because long-offset examples dominate the dataset.Log the loss separately:$$\mathcal L_{q=1}, \mathcal L_{q=2}, \ldots, \mathcal L_{q=5}.$$This will immediately show whether the controller learns the intended distance-dependent timing.8. What to expect after retrainingFor goals one block away, the predicted distance should bottom out at block one and remain low:$$d_1\approx d_2\approx\cdots\approx d_5.$$For goals three blocks away, a reasonable profile might be:$$[0.50,;0.18,;0.03,;0.03,;0.04].$$For goals five blocks away:$$[1.10,;0.70,;0.25,;0.07,;0.02].$$The exact values are not important. The important properties are: The goal-offset block $q$ is close to the goal. Distances after $q$ do not rebound substantially. One-block goals are solved mainly by the first block. Replanning every block no longer causes shrinking partial steps. The main success criterion is:$$\operatorname{Success}{\text{exec1}} \approx \operatorname{Success}{\text{exec5}}.$$The closed-loop result does not necessarily have to exceed the open-loop result immediately. First, the large $52%$ versus $88%$ discrepancy should collapse.9. Important CEM comparisonYou should evaluate CEM under both execution schedules: CEM terminal objective, execute one block; CEM terminal objective, execute all five blocks. The attached LeWM implementation says that it executes the entire optimized action sequence before replanning. Therefore, if your current CEM result uses one-block execution, it is not directly the same evaluation configuration as the paper.CEM with the terminal-only objective can suffer from the same horizon-reset procrastination. This is not fundamentally controller-specific.The fair comparison table should eventually contain:PlannerObjectiveExecution blocksCEMfixed terminal $d_H$1CEMfixed terminal $d_H$5Controllerold fixed terminal1Controllerold fixed terminal5Controllerarrival-and-hold1Controllerarrival-and-hold5Optionally, also give CEM the corrected arrival-and-hold objective. Otherwise, a reviewer could argue that the controller was compared against a knowingly time-inconsistent CEM configuration.The current result establishes that the controller beats this CEM configuration, not necessarily the best-tuned CEM baseline.10. Test full-plan execution for every $K$You currently have full execution only for $K=3$. Run:$$K\in{0,1,2,3,5}$$with full-plan execution.This distinguishes two questions: Does refinement improve plan quality? Does execution schedule improve compatibility with the learned objective? A useful table would be:$K$Execute 1Execute 5036?150?252?35288542?It is possible that $K=2$, full execution already achieves approximately $88%$, in which case $K=2$ becomes the strongest configuration.11. Analyze the $K=5$ failure separatelyBecause the controller was trained for $K=3$, extrapolating to $K=5$ is not guaranteed to improve the plan. The recurrent refinement function may overshoot or cycle.Log:$$J^{(0)},J^{(1)},J^{(2)},J^{(3)},J^{(4)},J^{(5)}$$on each episode. Then calculate:$$\Delta J^{(k)}

J^{(k-1)}-J^{(k)}.$$Questions to answer: Does predicted cost improve until $K=3$ and then worsen? Does predicted cost keep improving while real success worsens? Do action updates grow after the trained depth? Does the action plan oscillate? If you later want arbitrary test-time depth, train with a randomized number of refinements:$$K_{\mathrm{train}}\sim\operatorname{Uniform}{1,\ldots,5}.$$You can also add:$$\mathcal L_{\mathrm{monotonic}}

\frac1K \sum_{k=1}^{K} \max\left(0,J^{(k)}-J^{(k-1)}\right),$$but I would not add this in the next experiment. First isolate the arrival-loss correction.For now, simply use $K=2$ or $K=3$ at inference.12. Further diagnostic to strengthen the fixed-point claimMeasure the first-block progress as a function of current goal distance.Let $D_n$ be the true or latent distance before replanning and $D_{n+1}$ the distance after executing the first block. Fit:$$D_{n+1}=cD_n+b.$$If$$0<c<1,\qquad b>0,$$the resulting fixed point is:$$D^*=\frac{b}{1-c}.$$If $D^*$ lies outside the 20-pixel success radius, that quantitatively explains why increasing the execution budget from 50 to 200 does not help.Also measure the first-block action magnitude against distance. A shrinking action magnitude near the goal would provide direct evidence of asymptotic partial-step behavior.13. Statistical reportingWith 50 episodes, one episode changes success by two percentage points. Therefore: $52%$ versus $50%$ is not meaningful without paired outcomes; $K=2$ and $K=3$ should currently be called tied; $88%$ versus $52%$ is large, but should still be reported with uncertainty. Because all methods use identical start-goal pairs, save the per-episode binary outcomes and use a paired comparison, such as McNemar’s test or paired bootstrap confidence intervals.For multiple controller training seeds, use a hierarchical bootstrap: sample training seeds; sample held-out episodes; recompute the success difference. Avoid treating the same 50 episodes across multiple seeds as 150 fully independent episodes.Also distinguish two types of speedup: predictor-transition rows; wall-clock time. From your table, the $K=3$ comparison is approximately:$$\frac{383400}{144}\approx2663\times$$fewer predictor rows, but only:$$\frac{7.04}{0.18}\approx39\times$$faster in wall-clock time because CEM evaluates many rows in parallel.For full-plan execution:$$\frac{383400}{25}\approx15336\times$$fewer rows and:$$\frac{7.04}{0.13}\approx54\times$$wall-clock speedup.Report both numbers rather than calling predictor rows direct model invocations.14. What I would run nextIn priority order:Experiment 1: no retrainingRun the current controller and CEM with execution lengths:$$m\in{1,2,3,4,5}.$$If success rises with $m$, that strongly validates the horizon-reset explanation.Also test full-plan execution across every $K$.Experiment 2: corrected lossTrain:$$\mathcal L

\mathcal L_{\mathrm{deadline+hold}} + \lambda_{\mathrm{support}}\mathcal L_{\mathrm{support}}$$with:$$\lambda_{\mathrm{hold}}=0.5.$$Use balanced $q$-sampling and remove the old path loss.Experiment 3: hold ablationTrain or fine-tune with:$$\lambda_{\mathrm{hold}}\in{0,;0.5,;1.0}.$$The key comparison is deadline-only versus deadline-and-hold.Experiment 4: pending old ablationsThen finish: terminal-only; no-support. These remain useful, but the objective correction is now more important because you have identified a structural failure mode.Bottom lineThe $88%$ result is not an embarrassment or an evaluation artifact. It shows that your controller learned the five-block planning problem very effectively. The $52%$ receding-horizon result reveals that the learned planner is horizon-dependent rather than time-consistent.The cleanest correction is not min-over-blocks. It is:$$\boxed{ \text{reach by the known goal offset }q ;+; \text{remain close after }q. }$$That preserves outcome-based learning, does not clone dataset actions, uses information already available from future-goal relabeling, and directly targets the failure you observed. If it makes execute-one performance approach execute-five performance while retaining the controller’s large compute advantage, that becomes a strong central result rather than merely an implementation fix.