Understanding Multimodality in Generative Behavioral Cloning
Abstract
Behavioral cloning becomes challenging when the same observation admits several valid actions. We study how generative behavioral-cloning policies represent such multimodal expert behavior and identify different bottlenecks across model parameterizations. For latent-variable policies, preserving demonstrated modes requires action-conditioned information in the latent representation. Excessive posterior-prior regularization can suppress this information and prevent the policy from distinguishing demonstrated modes. Weaker or aggregate regularization can preserve mode information, but shifts the challenge to ensuring that the deployment-time prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with a small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal navigation and a physical-robot bimodal manipulation task support these mechanisms. In contrast, our analysis reveals limited conditional multimodality in standard robotic simulation benchmarks, where deterministic regression remains competitive.
Community
A little story about the paper: the idea goes back a while. Last December, I happened to be looking into ACT[Zhao]’s CVAE structure and its surprisingly high KL-regularization coefficient (β = 10 or even 100). Reading InfoVAE(an old gem by Stefano Ermon's group) made me suspect that the latent was little more than unstructured Gaussian noise. Some PCA analysis supported that suspicion. I then started experimenting with MMD and Sinkhorn regularization: I wanted a policy that could encode the supposed multimodality of BC data without the computational burden of diffusion or flow models.
On the popular LIBERO manipulation benchmark, the results were disappointing: a KL-CVAE with β = 100 performed on par with the alternatives. So I removed the CVAE completely and trained a deterministic action transformer. The results were quite shocking: it matched previously reported SOTA results. I couldn’t understand it: since the start of my PhD, the literature (IBC[Florence], Diffusion Policy[Chi], VQ-BeT[Lee], QFAT[Sheebaelhamd]) had convinced me that expressive policies were essential on BC benchmarks. Ariel Rodríguez and I ran further experiments on LIBERO and Meta-World and wrote a quick paper, which got rejected because it didn’t really explain why this was happening.
But I was still determined to find settings where the promised multimodality was observable. I developed a bunch of synthetic benchmarks with clearly separated trajectories from a fixed initial state. At this point, Massimiliano Datres made a huge contribution to shaping the theory, notation, and formalization of the ideas I had in mind. Together, we formalized multimodality in BC and derived bounds linking mode representation to KL regularization in CVAEs and the noise-to-action Lipschitz constant in diffusion/flow samplers.
With Ariel’s help integrating the simulation benchmarks used in the previously mentioned works, I then ran the experiments, expecting deterministic regression to fail badly, and… well, it still performed solidly. We constructed a proof-of-concept bimodal manipulation task on the robot, and there, we finally observed mode collapse.
Eventually, we decided to construct a conditional mode estimator based on GMM clustering for data without ground-truth mode labels. Our estimates suggested that all the simulation benchmarks we tested were pretty much unimodal across the sampled states! Deterministic regression remained competitive, often with lower inference latency. Looking back, this paper is a collection of failures that repeatedly went against what I expected. I’m happy we kept trying to understand them. It makes me confident that failures and the effort to understand them can lead to worthwhile science, and it’s nice that this paper came out of exactly that.
See you in Sydney, and thanks to all co-authors for their contributions.
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper