YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
SpaceVLA
Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors
SpaceVLA is a Unity-based framework for communicating a user's spatial intent to a vision-language-action (VLA) policy. Instead of specifying an entire robot trajectory, the user selects two targets:
- a green grasp anchor on the object;
- a blue placement anchor in the workspace.
The anchors are rendered directly into the robot's RGB observation. A LoRA-fine-tuned OpenVLA-7B policy then predicts incremental 7-DoF actions and executes the pick-and-place task in closed loop.
Research code. This repository contains the Unity environment and the OpenVLA training entry point used in the project. Model checkpoints and the 200-episode demonstration dataset are not included.
Overview
Language commands such as βpick up the mug and place it on the tableβ do not identify the user's preferred grasp region or placement location. SpaceVLA addresses this ambiguity through Visual Intent Anchors: user-authored image-space cues that condition an existing image-based VLA without changing its backbone.
flowchart TD
A["User selects grasp and place targets"] --> B["Unity renders green and blue anchors"]
B --> C["OpenVLA predicts a 7-DoF incremental action"]
C --> D["Unity robot controller executes the action"]
D --> E["Next RGB observation"]
E --> B
Before grasping, both anchors are visible. After grasp completion, the green grasp anchor is removed and only the blue placement anchor remains active during transport and release.
Key Features
- Direct spatial authoring: users specify preferred grasp and placement regions without teleoperation or full trajectory programming.
- RGB-native conditioning: anchors are alpha-composited into the camera image, preserving compatibility with image-conditioned VLA models.
- Closed-loop control: every predicted action is executed in Unity before the next marked observation is rendered.
- 7-DoF action prediction: the policy predicts translation deltas, controller-defined orientation deltas, and gripper width.
- LoRA fine-tuning: OpenVLA-7B is adapted on temporally subsampled, anchor-augmented Unity demonstrations.
Method
Let $I_t \in \mathbb{R}^{H \times W \times 3}$ be the RGB observation at time step $t$. The grasp and placement targets are represented by binary image-space masks $M_g$ and $M_p$. The policy receives the composited observation
During training and inference, the anchor-conditioned policy receives the instruction:
Pick up the object by the green region and place it on the blue region.
Each marked RGB frame is paired with the corresponding 7-DoF incremental action. Action statistics are computed from the training split and used for normalization before the actions are converted into OpenVLA action tokens.
Experimental Results
The dataset contains 200 Unity pick-and-place demonstrations:
- 120 episodes for fine-tuning;
- 80 disjoint episodes for closed-loop evaluation.
| Evaluation condition | Successful episodes | Full-task success | Mean grasp error | Mean placement error |
|---|---|---|---|---|
| Intended anchors | 73 / 80 | 91.25% | 0.5 cm | 0.7 cm |
| Random object anchors | 62 / 80 | 77.5% | 2.6 cm | 3.4 cm |
| No-anchor baseline | 40 / 80 | 50.0% | N/A | N/A |
The results show that the policy uses the spatial positions of the anchors rather than treating them only as visual decorations. In additional trials, anchors placed at arbitrary workspace locations redirected the robot toward the specified positions.
Repository Structure
.
βββ ArmRobot/
β βββ Assets/
β β βββ Meshes/ # Robot and environment meshes
β β βββ Scenes/ # Unity scenes
β β βββ Scripts/ # Robot, gripper, and interaction controllers
β βββ Packages/ # Unity package manifest and lock file
β βββ ProjectSettings/ # Unity project configuration
βββ Assets/ # Additional project assets
βββ train_openvla_visual_intent_no_val.py
βββ .gitignore
Unity-generated directories such as Library, Logs, Temp, Obj, and UserSettings are intentionally excluded from version control.
Getting Started
1. Clone the repository
The repository is currently private, so your GitHub account must have access to the MWS-Physical-AI organization.
git clone git@github.com:MWS-Physical-AI/SpaceVLA-Spatially-Grounded-VLA-for-Robotic-Manipulation-with-User-Authored-Grasp-and-Place-Anchors.git SpaceVLA
cd SpaceVLA
2. Open the Unity project
- Install the Unity Editor version specified in
ArmRobot/ProjectSettings/ProjectVersion.txt. - In Unity Hub, select Add project from disk.
- Open the
ArmRobot/directory. - Open the included robot scene under
ArmRobot/Assets/Scenes/.
Unity reconstructs the ignored Library/ directory during the first import. The initial import can therefore take several minutes.
3. Prepare OpenVLA fine-tuning
Set up an environment compatible with OpenVLA, then provide:
- the OpenVLA-7B base checkpoint;
- the anchor-augmented training data;
- local dataset and output paths expected by the training script.
The repository's training entry point is:
python train_openvla_visual_intent_no_val.py
Before starting a long run, inspect the script and update its dataset, checkpoint, and output paths for your machine. The repository does not currently provide a pinned Python environment or downloadable model weights.
Anchor-Augmented Data
For every recorded demonstration:
- annotate the grasp target in the initial part of the episode;
- annotate the placement target in the target workspace region;
- render both masks before grasp completion;
- remove the grasp mask after the object is grasped;
- pair each resulting RGB observation with the instruction and corresponding 7-DoF action;
- temporally subsample the episode and add the samples to the training data.
The no-anchor baseline uses the original RGB observations and the instruction:
Pick up the object and place it on the table.
Current Limitations
- Anchors are stored as fixed image-space masks under a static camera.
- They are not yet attached to objects or world-space locations in 3D.
- Camera motion or object motion can therefore break the alignment between a mask and its intended target.
- The current repository does not package the training dataset, model checkpoints, or a fully pinned training environment.
Future work includes object-relative and world-space 3D anchors, OpenXR interaction, and AR-based selection of grasp and placement targets for physical robots.
Paper
SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors
- Paper: arXiv:2608.05730
- PDF: arxiv.org/pdf/2608.05730
Citation
If you use this project in your research, please cite:
@article{zinniatullina2026spacevla,
title = {{SpaceVLA}: Spatially Grounded {VLA} for Robotic Manipulation with User-Authored Grasp and Place Anchors},
author = {Zinniatullina, Daniia and Kolomiets, Iaroslav and Konenkov, Mikhail and Altamirano Cabrera, Miguel and Tsetserukou, Dzmitry},
journal = {arXiv preprint arXiv:2608.05730},
year = {2026}
}
Acknowledgments
This work was conducted at the Skolkovo Institute of Science and Technology and the R&D Center, MWS.