LaCT-Motion SFT checkpoint (checkpoint-epoch6)

Curriculum supervised fine-tuning (SFT) checkpoint of LaCT-Motion (Latent Chain-of-Thought for Motion), the official implementation of the ECCV 2026 paper Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Motion Generation.

This is the checkpoint referenced by the GRPO configurations in the code repository: it initializes the Stage 2 GRPO run that refines the fully latent policy.

Built with Qwen. This model is a fine-tuned derivative of Qwen2.5-3B-Instruct and is released under the Qwen RESEARCH LICENSE AGREEMENT (non-commercial research and evaluation use only). See NOTICE.

Model details

Item Value
Base model Qwen/Qwen2.5-3B-Instruct (Qwen2ForCausalLM, 36 layers, hidden size 2048)
Vocabulary 152,184 tokens: the Qwen vocabulary extended with 512 motion codes (<Motion_0> … <Motion_511>), motion delimiters, and latent-reasoning special tokens
Latent settings c_thought: 2, max_latent_stage: 8, pad_latent_to_max: true (16 latent tokens per prompt)
Precision bfloat16 safetensors, two shards
Motion tokenizer Motion VQ-VAE with 512 codes (ckpt/vqvae.pth in the code repository, from the Motion-Agent release)
Training data HumanML3D, processed splits from data/ in the code repository

The checkpoint contains the model weights, configuration, and tokenizer only; the SFT optimizer state is not included.

Usage

Download it into the location expected by the code repository:

hf download RHYu2233/LaCT-Motion-SFT --local-dir checkpoints/sft/checkpoint-epoch6

Then follow the repository README to run GRPO from this checkpoint, evaluate it on HumanML3D, or generate motions with the demo:

python eval_t2m.py options/sft/t2m_coconut.yaml \
  --checkpoint checkpoints/sft/checkpoint-epoch6 \
  --batch-size 32 --repeat 1 --device cuda:0 \
  --output results/t2m_metrics.json

Generation requires the motion VQ-VAE, the T2M evaluator, and the GloVe files described in the repository README; the model outputs motion codes that the VQ-VAE decodes into HumanML3D joint features.

License

The weights are distributed under the Qwen RESEARCH LICENSE AGREEMENT inherited from the base model Qwen2.5-3B-Instruct: non-commercial research and evaluation use only; commercial use requires a license from Alibaba Cloud. The agreement is included as LICENSE, and the required attribution notice is in NOTICE. The training data derives from HumanML3D (AMASS and HumanAct12), which are likewise restricted to non-commercial research use.

Citation

@inproceedings{yu2026lactmotion,
  title     = {Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Motion Generation},
  author    = {Yu, Ronghao and Liu, Yang and Wang, Juncheng and Xu, Chao and Shao, Yimo and Sun, Baigui and Liu, Yong and Luo, Shan},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}
Downloads last month
229
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RHYu2233/LaCT-Motion-SFT

Base model

Qwen/Qwen2.5-3B
Finetuned
(1595)
this model