UniMate Checkpoints
Pretrained checkpoints for UniMate: One Unified Model to Animate Diverse Skeletons (SIGGRAPH Asia 2026). UniMate is a single text-conditioned flow-matching model that generates motion for skeletons of any topology: animals, humanoids and rigged objects.
Project Page · Paper · Video · Code · Dataset · Interactive Demo
Models
| Model | Training data | Joints per skeleton | Steps | Params | Recommended checkpoint | Status |
|---|---|---|---|---|---|---|
unimate_uniml3d_f60_v2 |
UniML3D, release of 2026-09-27 | 5 to 70 | 100k | 74.1M | checkpoint_step_100000.pt |
Recommended |
unimate_uniml3d_f60_preview |
UniML3D, earlier build | 5 to 60 | 120k | 74.1M | checkpoint_step_120000.pt |
Superseded |
Both models share one architecture: graph attention over joints with AdaLN text conditioning, 10 layers of width 512, and a google/flan-t5-base text encoder. They generate 60-frame clips (2 seconds at 30 fps) and are trained on all three sources of UniML3D: Truebones ZOO animals, Mixamo humanoids and rigged Objaverse-XL objects.
unimate_uniml3d_f60_v2
The recommended model. It is trained with configs/uniml3d_60frames_graph_adaln_v2.json on the UniML3D release of 2026-09-27 (dataset revision faaa817). Compared with the preview model, it:
- covers more skeletons. The joint limit is raised from 60 to 70, so 98.1% of the training clips are kept instead of 86.9%. This includes 70 of the 74 Truebones species.
- samples the datasets more evenly. Datasets are balanced before object types (
sampler_dataset_alpha = 0.25), so Truebones, Mixamo and Objaverse-XL make up about 13%, 24% and 63% of the training samples. Before, the single Mixamo rig got under 1% of them.
Training: 100k steps on 8 NVIDIA H100 GPUs (about 22 hours), batch size 16 per GPU.
unimate_uniml3d_f60_preview
The first public model. It is trained with configs/uniml3d_60frames_graph_adaln.json on an earlier build of UniML3D, made before the 2026-09-27 release, so its training data differ from the current dataset. Training: 120k steps on 6 NVIDIA H100 GPUs (about 23 hours), batch size 16 per GPU. It is kept for reference; use unimate_uniml3d_f60_v2 for new work.
Quick start
1. Install the code. Clone the code repository and set up its unimate environment as described in its README. Run every command below from the repository root.
2. Prepare the dataset features. At sampling time, the model reads each target skeleton (its T-pose, topology and joint names) from dataset/features/<dataset>/. The Hub dataset ships the stage 1-3 annotations only, so build the features locally with stage 4 of the data pipeline. Use the dataset revision the model was trained on:
hf download Linzhan/UniML3D --repo-type dataset \
--revision faaa81773b315247b03f183548e3898dbec2ce80 --local-dir dataset
bash data_process/scripts/run_extract_features.sh truebones # likewise mixamo, objaverse
The Truebones motion files come from a commercial pack and are not redistributed; see the dataset card for how to rebuild them.
3. Download a model. This downloads the recommended checkpoint only:
hf download Linzhan/UniMate \
--include "unimate_uniml3d_f60_v2/*.json" "unimate_uniml3d_f60_v2/*.npy" \
"unimate_uniml3d_f60_v2/checkpoints/checkpoint_step_100000.pt" \
--local-dir outputs
4. Generate motion from text. Write the prompts to a JSON file whose keys are <object_type>-<case_id>:
{
"Horse-0": "A horse gallops forward.",
"mixamo-0": "A person jumps and waves both arms."
}
Then run:
python -m unimate.inference.sample \
--exp_dir outputs/unimate_uniml3d_f60_v2 \
--test_cases_json test_cases.json \
--num_repetitions 3
The script uses the latest checkpoint in checkpoints/ and the EMA weights. The code README covers the other applications: in-betweening, joint-level editing, motion expansion, and driving a rigged mesh with the result.
Training
To resume training from a released checkpoint:
bash scripts/run_train.sh configs/uniml3d_60frames_graph_adaln_v2.json -- \
--resume outputs/unimate_uniml3d_f60_v2/checkpoints/checkpoint_step_100000.pt
To train unimate_uniml3d_f60_v2 from scratch:
accelerate launch --num_processes 8 -m unimate.training.train \
--config configs/uniml3d_60frames_graph_adaln_v2.json
Repository layout
config.json index of the released models
<model>/
config.json resolved training configuration; read by inference
dataset_stats.npy feature normalization statistics
checkpoints/
checkpoint_step_<N>.pt model and EMA weights, optimizer and scheduler state; saved every 10k steps
logs/ TensorBoard training curves
samples/
step_<NNNNNN>/ motions generated during training (step_000000: before training)
<object_type>-<i>_fk.mp4 joints from forward kinematics of the predicted rotations
<object_type>-<i>_ric.mp4 predicted joint positions
Limitations
- Clip length: clips are 60 frames at 30 fps. The expansion application chains clips into longer motion.
- Skeletons: a skeleton must go through the data pipeline (stages 1-4) before the model can animate it. Skeletons with more joints than the model's limit (70 for v2) are outside its training range.
- Data version: UniML3D is actively maintained. Captions and annotations change between releases, so pin the dataset revision a model was trained on if you need to reproduce it. Models trained on later releases will be published here with their data version noted.
License
The UniMate code is released under the MIT License. The training data remain governed by the licenses of their original sources: Adobe's Mixamo terms of use, the per-object licenses of Objaverse-XL, and the commercial license of the Truebones ZOO pack. Please review these terms before using the models.
Citation
@article{mou2026unimate,
title = {UniMate: One Unified Model to Animate Diverse Skeletons},
author = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
journal = {arXiv preprint arXiv:2609.05415},
year = {2026}
}
- Downloads last month
- -
