File size: 4,977 Bytes
1d2cd4d bf0ed2f 1d2cd4d fea458d 1d2cd4d 9a02c9c 1d2cd4d bf0ed2f 1d2cd4d a47f59b 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d fea458d 1d2cd4d bf0ed2f 1d2cd4d fea458d 1d2cd4d fea458d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 | ---
language:
- en
license: other
library_name: robotstar
pipeline_tag: robotics
tags:
- sign-language-generation
- text-to-motion
- robotics
- motion-generation
datasets:
- how2sign
---
<h1 align="center"><strong>RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots</strong></h1>
<p align="center">
<a href="https://www.yujiazeng.com/">Yujia Zeng<sup>*</sup></a>
<a href="https://pengchensheng.com/">Chensheng Peng<sup>*</sup></a>
<a href="https://thomaschen98.github.io/">Yuxin Chen<sup>*</sup></a>
<a href="https://alexshao.net/">Alex Shao<sup></sup></a>
<a href="https://www.linkedin.com/in/nathan-jew-a31314273/">Nathan Jew<sup></sup></a>
<a href="https://me.berkeley.edu/people/masayoshi-tomizuka/">Masayoshi Tomizuka<sup></sup></a> <br>
<sup></sup>UC Berkeley <br> <sup>*</sup>Equal contribution
</p>
<p align="center">
<a href="https://arxiv.org/abs/2609.32250">
<img src="https://img.shields.io/badge/arXiv-2609.32250-b31b1b" alt="arXiv">
</a>
<a href="https://zyjorz.github.io/RoboSTAR/">
<img src="https://img.shields.io/badge/Project-Page-blue">
</a>
<a href="https://github.com/zyjOrz/RoboSTAR">
<img src="https://img.shields.io/badge/GitHub-Code-black">
</a>
<a href="https://huggingface.co/Ivystream/RoboSTAR">
<img src="https://img.shields.io/badge/🤗_HuggingFace-Model-orange">
</a>
</p>
<p align="center">
<em>
RoboSTAR translates English text or audio into continuous American Sign Language (ASL) motion. <br>
This release contains the text-to-sign-motion component.
</em>
</p>
## Installation
```bash
git clone https://github.com/zyjOrz/RoboSTAR.git
cd RoboSTAR
conda create -n robostar python=3.10 -y
conda activate robostar
pip install -e .
```
## Inference
Predicted-duration inference:
```bash
python -m robostar.infer \
--model Ivystream/RoboSTAR \
--text "A person explains the plan." \
--length-mode predicted \
--output-dir outputs/predicted
```
For a stable demonstration duration, specify seconds explicitly:
```bash
python -m robostar.infer \
--model Ivystream/RoboSTAR \
--text "A person explains the plan." \
--length-mode seconds \
--length-value 6.0 \
--output-dir outputs/six_seconds
```
Other supported modes are `tokens` and `frames`.
## Data preparation
```bash
python -m robostar.prepare_data \
--train-manifest data/train_raw.jsonl \
--val-manifest data/val_raw.jsonl \
--test-manifest data/test_raw.jsonl \
--output data/prepared
```
## Training
### 1. Train the FSQ tokenizer
```bash
torchrun --standalone --nproc_per_node=8 -m robostar.train_tokenizer \
--config configs/tokenizer.yaml \
--prepared-root data/prepared \
--output experiments/robostar_tokenizer
```
### 2. Export motion tokens
```bash
torchrun --standalone --nproc_per_node=8 -m robostar.export_tokens \
--model experiments/robostar_tokenizer/best.pt \
--prepared-root data/prepared \
--output data/tokens
```
### 3. Build the optional retrieval memory
The isolated-word source used in our experiments is [`akasheroor/American-Sign-Language-Dataset`](https://huggingface.co/datasets/akasheroor/American-Sign-Language-Dataset).
```bash
python -m robostar.retrieval build \
--word-token-jsonl data/word_tokens/train_source_tokens.jsonl \
--output data/retrieval/word2code.json
```
### 4. Build coarse-to-fine caches
```bash
python -m robostar.build_cache \
--token-root data/tokens \
--tokenizer-model experiments/robostar_tokenizer/best.pt \
--base-model google/mt5-large \
--retrieval data/retrieval/word2code.json \
--output data/cache
```
### 5. Train RoboSTAR
```bash
torchrun --standalone --nproc_per_node=8 -m robostar.train_generator \
--config configs/robostar_mt5_large.yaml \
--cache-root data/cache \
--output experiments/robostar_mt5_large
```
## Evaluation
```bash
python -m robostar.evaluate \
--predictions outputs/predictions \
--manifest data/prepared/test.jsonl
```
## 🔗 Citation
If you find our work useful, please consider citing:
```bibtex
@misc{zeng2026robostarnextscaleautoregressivesign,
title={RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots},
author={Yujia Zeng and Chensheng Peng and Yuxin Chen and Alex Shao and Nathan Jew and Masayoshi Tomizuka},
year={2026},
eprint={2609.32250},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.32250},
}
```
## 🙏 Acknowledgements
- [SOKE](https://github.com/2000ZRL/SOKE): Our work follows SOKE in adopting its motion representation and sign retrieval formulation.
- [How2Sign Dataset](https://how2sign.github.io/): Our work uses the How2Sign dataset for training and evaluation.
- [HandMDM](https://imagine.enpc.fr/~leore.bensabath/HandMDM/): Our human-mesh rendering and visualization pipeline is inspired by HandMDM.
|