--- language: - en license: other library_name: robotstar pipeline_tag: robotics tags: - sign-language-generation - text-to-motion - robotics - motion-generation datasets: - how2sign ---

RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots

Yujia Zeng* Chensheng Peng* Yuxin Chen* Alex Shao Nathan Jew Masayoshi Tomizuka
UC Berkeley
*Equal contribution

arXiv

RoboSTAR translates English text or audio into continuous American Sign Language (ASL) motion.
This release contains the text-to-sign-motion component.

## Installation ```bash git clone https://github.com/zyjOrz/RoboSTAR.git cd RoboSTAR conda create -n robostar python=3.10 -y conda activate robostar pip install -e . ``` ## Inference Predicted-duration inference: ```bash python -m robostar.infer \ --model Ivystream/RoboSTAR \ --text "A person explains the plan." \ --length-mode predicted \ --output-dir outputs/predicted ``` For a stable demonstration duration, specify seconds explicitly: ```bash python -m robostar.infer \ --model Ivystream/RoboSTAR \ --text "A person explains the plan." \ --length-mode seconds \ --length-value 6.0 \ --output-dir outputs/six_seconds ``` Other supported modes are `tokens` and `frames`. ## Data preparation ```bash python -m robostar.prepare_data \ --train-manifest data/train_raw.jsonl \ --val-manifest data/val_raw.jsonl \ --test-manifest data/test_raw.jsonl \ --output data/prepared ``` ## Training ### 1. Train the FSQ tokenizer ```bash torchrun --standalone --nproc_per_node=8 -m robostar.train_tokenizer \ --config configs/tokenizer.yaml \ --prepared-root data/prepared \ --output experiments/robostar_tokenizer ``` ### 2. Export motion tokens ```bash torchrun --standalone --nproc_per_node=8 -m robostar.export_tokens \ --model experiments/robostar_tokenizer/best.pt \ --prepared-root data/prepared \ --output data/tokens ``` ### 3. Build the optional retrieval memory The isolated-word source used in our experiments is [`akasheroor/American-Sign-Language-Dataset`](https://huggingface.co/datasets/akasheroor/American-Sign-Language-Dataset). ```bash python -m robostar.retrieval build \ --word-token-jsonl data/word_tokens/train_source_tokens.jsonl \ --output data/retrieval/word2code.json ``` ### 4. Build coarse-to-fine caches ```bash python -m robostar.build_cache \ --token-root data/tokens \ --tokenizer-model experiments/robostar_tokenizer/best.pt \ --base-model google/mt5-large \ --retrieval data/retrieval/word2code.json \ --output data/cache ``` ### 5. Train RoboSTAR ```bash torchrun --standalone --nproc_per_node=8 -m robostar.train_generator \ --config configs/robostar_mt5_large.yaml \ --cache-root data/cache \ --output experiments/robostar_mt5_large ``` ## Evaluation ```bash python -m robostar.evaluate \ --predictions outputs/predictions \ --manifest data/prepared/test.jsonl ``` ## 🔗 Citation If you find our work useful, please consider citing: ```bibtex @misc{zeng2026robostarnextscaleautoregressivesign, title={RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots}, author={Yujia Zeng and Chensheng Peng and Yuxin Chen and Alex Shao and Nathan Jew and Masayoshi Tomizuka}, year={2026}, eprint={2609.32250}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2609.32250}, } ``` ## 🙏 Acknowledgements - [SOKE](https://github.com/2000ZRL/SOKE): Our work follows SOKE in adopting its motion representation and sign retrieval formulation. - [How2Sign Dataset](https://how2sign.github.io/): Our work uses the How2Sign dataset for training and evaluation. - [HandMDM](https://imagine.enpc.fr/~leore.bensabath/HandMDM/): Our human-mesh rendering and visualization pipeline is inspired by HandMDM.