Instructions to use RHYu2233/LaCT-Motion-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use RHYu2233/LaCT-Motion-SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="RHYu2233/LaCT-Motion-SFT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("RHYu2233/LaCT-Motion-SFT") model = AutoModelForCausalLM.from_pretrained("RHYu2233/LaCT-Motion-SFT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use RHYu2233/LaCT-Motion-SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "RHYu2233/LaCT-Motion-SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RHYu2233/LaCT-Motion-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/RHYu2233/LaCT-Motion-SFT
- SGLang
How to use RHYu2233/LaCT-Motion-SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "RHYu2233/LaCT-Motion-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RHYu2233/LaCT-Motion-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "RHYu2233/LaCT-Motion-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RHYu2233/LaCT-Motion-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use RHYu2233/LaCT-Motion-SFT with Docker Model Runner:
docker model run hf.co/RHYu2233/LaCT-Motion-SFT
LaCT-Motion SFT checkpoint (checkpoint-epoch6)
Curriculum supervised fine-tuning (SFT) checkpoint of LaCT-Motion (Latent Chain-of-Thought for Motion), the official implementation of the ECCV 2026 paper Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Motion Generation.
- Code: https://github.com/Erwin2233/LaCT-Motion
- Paper (PDF): https://media.eventhosts.cc/Conferences/ECCV2026/pdfs/13122.pdf
This is the checkpoint referenced by the GRPO configurations in the code repository: it initializes the Stage 2 GRPO run that refines the fully latent policy.
Built with Qwen. This model is a fine-tuned derivative of Qwen2.5-3B-Instruct and is released under the Qwen RESEARCH LICENSE AGREEMENT (non-commercial research and evaluation use only). See NOTICE.
Model details
| Item | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct (Qwen2ForCausalLM, 36 layers, hidden size 2048) |
| Vocabulary | 152,184 tokens: the Qwen vocabulary extended with 512 motion codes (<Motion_0> … <Motion_511>), motion delimiters, and latent-reasoning special tokens |
| Latent settings | c_thought: 2, max_latent_stage: 8, pad_latent_to_max: true (16 latent tokens per prompt) |
| Precision | bfloat16 safetensors, two shards |
| Motion tokenizer | Motion VQ-VAE with 512 codes (ckpt/vqvae.pth in the code repository, from the Motion-Agent release) |
| Training data | HumanML3D, processed splits from data/ in the code repository |
The checkpoint contains the model weights, configuration, and tokenizer only; the SFT optimizer state is not included.
Usage
Download it into the location expected by the code repository:
hf download RHYu2233/LaCT-Motion-SFT --local-dir checkpoints/sft/checkpoint-epoch6
Then follow the repository README to run GRPO from this checkpoint, evaluate it on HumanML3D, or generate motions with the demo:
python eval_t2m.py options/sft/t2m_coconut.yaml \
--checkpoint checkpoints/sft/checkpoint-epoch6 \
--batch-size 32 --repeat 1 --device cuda:0 \
--output results/t2m_metrics.json
Generation requires the motion VQ-VAE, the T2M evaluator, and the GloVe files described in the repository README; the model outputs motion codes that the VQ-VAE decodes into HumanML3D joint features.
License
The weights are distributed under the Qwen RESEARCH LICENSE AGREEMENT inherited from the base model Qwen2.5-3B-Instruct: non-commercial research and evaluation use only; commercial use requires a license from Alibaba Cloud. The agreement is included as LICENSE, and the required attribution notice is in NOTICE. The training data derives from HumanML3D (AMASS and HumanAct12), which are likewise restricted to non-commercial research use.
Citation
@inproceedings{yu2026lactmotion,
title = {Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Motion Generation},
author = {Yu, Ronghao and Liu, Yang and Wang, Juncheng and Xu, Chao and Shao, Yimo and Sun, Baigui and Liu, Yong and Luo, Shan},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}
- Downloads last month
- 229