RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
Abstract
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.
Community
Paper accepted by CoRL 2026
RoboFollow is a simulation benchmark built on RoboTwin for evaluating instruction-conditioned manipulation by embodied agents. Multiple tasks share the same scene configuration, requiring policies to use language to select the appropriate object, spatial relation, motion constraint, or logical branch.
The benchmark provides four scene families, expert demonstration collection, an L0โL3 evaluation protocol, and stage-wise Intent and Execution scores. These complementary measures help distinguish incorrect task selection from incomplete physical execution.
๐ Paper: https://arxiv.org/abs/2609.25636
๐ฅ Project Page: https://mrc-crm.github.io/RoboFollow/
๐น๏ธ Code: https://github.com/AutoLab-SAI-SJTU/RoboFollow
๐ค Huggingface Dataset: https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation (2026)
- The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction (2026)
- HINT: Human-Intent Inception for Long-Horizon Robot Manipulation (2026)
- StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models (2026)
- JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations (2026)
- Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models (2026)
- G0.5: One Autoregressive Stream for Robot Reasoning and Action (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.25636 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper