MotionPersona V1
Forty-four characters generated by this model, each conditioned on the persona and body shape of one annotated participant, following the same straight path.
The trained model of MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles: a single generative locomotion controller whose three controls are a captured motion persona (who is moving), a target body (an SMPL-X body shape that carries the motion) and a performed style (how the character is moving), under a trajectory command. The controller has 36M parameters (175 MB), runs in real time on a CPU (27 ms per block on two threads of a laptop CPU), and is trained on thousands of hours of motion: MotionPersonaX holds 4,200 hours (33 hours of captured takes retargeted onto 128 bodies), and after the evaluation split is held out the VAE trains on 3,700 hours and the prior on 2,500 hours. Training is efficient: both stages finish in about 30 hours on consumer GPUs, under 100 GPU-hours in total (VAE: ~10 h on 2x RTX 4090; prior: ~19 h on 4x RTX 4090).
Code, documentation and the realtime demo: https://github.com/AIGAnimation/MotionPersona · project page: https://motionpersona25.github.io/ · paper: https://arxiv.org/abs/2506.00173
Files
| file | contents | parameters |
|---|---|---|
prior.ckpt |
stage 2, latent flow-matching prior (DiT, 8 blocks x 384): EMA weights + full training config | 18.8M |
codec.ckpt |
stage 1, shape-aware VAE codec: weights + full training config | 17.2M |
Both are fp32 PyTorch Lightning checkpoints and are the exact model evaluated in the paper. The prior finds its
codec as codec.ckpt in the same folder.
Usage
The checkpoints are loaded by the code repository; its programs download them automatically on first use.
git clone https://github.com/AIGAnimation/MotionPersona.git && cd MotionPersona
pip install -r requirements.txt
python generate.py --persona p02,p10 --style happy,drunk --body own,grid_h08_g04 --seconds 20 # -> outputs/*.bvh
python realtime_demo.py # interactive demo
To download them by hand:
from huggingface_hub import hf_hub_download
for f in ("prior.ckpt", "codec.ckpt"):
hf_hub_download("myshi/MotionPersona_V1", f, local_dir="checkpoints")
Model
- Codec. A KL-VAE that encodes a 45-frame future window into 9 latent tokens of 48 dimensions and decodes them on a target body given its 10 SMPL-X shape coefficients and the last 5 frames, with explicit forward-kinematics, contact and seam supervision.
- Prior. Flow matching over the codec's tokens with a DiT denoiser, sampled with two Euler steps per block. Conditions: persona (a learned performer-ID token plus role, affiliation and dominance attribute tokens), style, body (betas), trajectory (45 future positions and orientations) and motion history.
- Controls. Personas
p01–p44(the annotated participants of MotionPersona); nine styles (angry,depressed,fear,happy,neutral,bigstep,drunk,swimming,twofootjump); any SMPL-X body given by 10 betas, trained on 113 of the 128 MotionPersonaX bodies (captured and synthetic, 1.05–1.95 m; 15 held out). - Output. 30 fps motion on the 23-joint SMPL-X body skeleton (local joint rotations, root trajectory, foot contacts), exported as BVH by the code.
Training data
Trained on MotionPersonaX, the cross-body retargeting of MotionPersona onto 128 SMPL-X bodies (328,960 clips), with 15 bodies, 132 persona x body combinations and 744 source takes held out for evaluation. Codec: 60 epochs on 2 GPUs; prior: 100 epochs on 4 GPUs (RTX 4090). The training recipes are in the code repository.
Limitations
- Personas are the closed set of captured participants; a new persona needs new capture data.
- Bodies outside the SMPL-X shape range of the training bodies are extrapolation.
- Hands are not modelled (the 23-joint body skeleton has no fingers).
- The motion is kinematic; physical plausibility (contacts, balance) is learned from data, not enforced.
License
CC BY-NC 4.0, for non-commercial and research purposes only. For commercial use, contact myshi@cs.hku.hk and taku@hku.hk to obtain a commercial license. The body skeletons are SMPL-X skeletons; use of the SMPL-X model is subject to its own license (https://smpl-x.is.tue.mpg.de).
Citation
If you use the data, code, or any module from this repo, please cite the original paper:
@article{shi2025motionpersona,
title={MotionPersona: Characteristics-aware Locomotion Control},
author={Shi, Mingyi and Liu, Wei and Mei, Jidong and Tse, Wangpok and Chen, Rui and Chen, Xuelin and Komura, Taku},
journal={arXiv preprint arXiv:2506.00173},
year={2025}
}