---
language:
- en
license: other
library_name: robotstar
pipeline_tag: robotics
tags:
- sign-language-generation
- text-to-motion
- robotics
- motion-generation
datasets:
- how2sign
---
RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots
Yujia Zeng*
Chensheng Peng*
Yuxin Chen*
Alex Shao
Nathan Jew
Masayoshi Tomizuka
UC Berkeley
*Equal contribution
RoboSTAR translates English text or audio into continuous American Sign Language (ASL) motion.
This release contains the text-to-sign-motion component.
## Installation
```bash
git clone https://github.com/zyjOrz/RoboSTAR.git
cd RoboSTAR
conda create -n robostar python=3.10 -y
conda activate robostar
pip install -e .
```
## Inference
Predicted-duration inference:
```bash
python -m robostar.infer \
--model Ivystream/RoboSTAR \
--text "A person explains the plan." \
--length-mode predicted \
--output-dir outputs/predicted
```
For a stable demonstration duration, specify seconds explicitly:
```bash
python -m robostar.infer \
--model Ivystream/RoboSTAR \
--text "A person explains the plan." \
--length-mode seconds \
--length-value 6.0 \
--output-dir outputs/six_seconds
```
Other supported modes are `tokens` and `frames`.
## Data preparation
```bash
python -m robostar.prepare_data \
--train-manifest data/train_raw.jsonl \
--val-manifest data/val_raw.jsonl \
--test-manifest data/test_raw.jsonl \
--output data/prepared
```
## Training
### 1. Train the FSQ tokenizer
```bash
torchrun --standalone --nproc_per_node=8 -m robostar.train_tokenizer \
--config configs/tokenizer.yaml \
--prepared-root data/prepared \
--output experiments/robostar_tokenizer
```
### 2. Export motion tokens
```bash
torchrun --standalone --nproc_per_node=8 -m robostar.export_tokens \
--model experiments/robostar_tokenizer/best.pt \
--prepared-root data/prepared \
--output data/tokens
```
### 3. Build the optional retrieval memory
The isolated-word source used in our experiments is [`akasheroor/American-Sign-Language-Dataset`](https://huggingface.co/datasets/akasheroor/American-Sign-Language-Dataset).
```bash
python -m robostar.retrieval build \
--word-token-jsonl data/word_tokens/train_source_tokens.jsonl \
--output data/retrieval/word2code.json
```
### 4. Build coarse-to-fine caches
```bash
python -m robostar.build_cache \
--token-root data/tokens \
--tokenizer-model experiments/robostar_tokenizer/best.pt \
--base-model google/mt5-large \
--retrieval data/retrieval/word2code.json \
--output data/cache
```
### 5. Train RoboSTAR
```bash
torchrun --standalone --nproc_per_node=8 -m robostar.train_generator \
--config configs/robostar_mt5_large.yaml \
--cache-root data/cache \
--output experiments/robostar_mt5_large
```
## Evaluation
```bash
python -m robostar.evaluate \
--predictions outputs/predictions \
--manifest data/prepared/test.jsonl
```
## 🔗 Citation
If you find our work useful, please consider citing:
```bibtex
@misc{zeng2026robostarnextscaleautoregressivesign,
title={RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots},
author={Yujia Zeng and Chensheng Peng and Yuxin Chen and Alex Shao and Nathan Jew and Masayoshi Tomizuka},
year={2026},
eprint={2609.32250},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.32250},
}
```
## 🙏 Acknowledgements
- [SOKE](https://github.com/2000ZRL/SOKE): Our work follows SOKE in adopting its motion representation and sign retrieval formulation.
- [How2Sign Dataset](https://how2sign.github.io/): Our work uses the How2Sign dataset for training and evaluation.
- [HandMDM](https://imagine.enpc.fr/~leore.bensabath/HandMDM/): Our human-mesh rendering and visualization pipeline is inspired by HandMDM.