# Ego3DLM: Ego-Human Motion Prediction with 3D-Aware LLM
**ECCV 2026**
Ego3DLM forecasts human motion from an egocentric perspective by grounding a language
model in the **3D spatial and semantic context** of the surrounding environment. Given
three-point motion tracking (head + hands), 3D scene features, and egocentric video,
Ego3DLM **simultaneously decodes past pose, future pose, past narration, and future
narration in a single autoregressive pass**, so that predicted poses and descriptions
are grounded in one another for cross-modal and temporal consistency.
The model is trained in three stages on top of a frozen motion tokenizer:
1. **Stage I — Spatial-Semantic Scene Awareness Pretraining.** Encode the 3D scene and
inject it into the LM, training spatial and semantic scene understanding (scene QA +
obstacle/free-space QA).
2. **Stage II — Multi-Modal Multi-Task Instruction Tuning.** Holistically train past/future
pose tracking and prediction alongside past/future narration in a single pass.
3. **Stage III — Multi-Modal Reward GRPO.** Reinforcement finetuning with intra- and
inter-modal rewards that directly optimize pose–language fidelity.
Experiments on the **Nymeria** benchmark show state-of-the-art performance on future pose
prediction, past motion tracking, and language description.
## 0. News
- **2026** — Ego3DLM is accepted to **ECCV 2026**. Code released for reproducibility.
## 1. Environment Setup
Tested with Python 3.11, CUDA 11.8, PyTorch 2.0.0, Transformers 4.46.3, PyTorch-Lightning 2.0.0.
```bash
# create env
conda create -n ego3dlm python=3.11 -y
conda activate ego3dlm
# install pinned dependencies
pip install -r requirements-freeze.txt
```
The repository expects three top-level symlinks pointing to your data / dependency /
checkpoint roots (see [Section 4](#4-dependencies--pretrained-models)):
```bash
ln -s /path/to/checkpoints ./checkpoints
ln -s /path/to/datasets ./datasets
ln -s /path/to/deps ./deps
```
The obstacle / free-space QA labels (`obstacle_labels/`, required by Stage-I pretraining and
CoT) are distributed separately — download them and place the directory at `./obstacle_labels`.
## 2. Dataset
*(To be added.)*
## 3. Data Preprocessing
*(To be added.)*
## 4. Dependencies & Pretrained Models
All paths below are relative to the repository root and resolved through the
`checkpoints/`, `datasets/`, and `deps/` symlinks.
### Required for every stage
| Artifact | Path | Notes |
|---|---|---|
| Motion PQ-VAE tokenizer | `checkpoints/VQVAE_pq_full_4096_64/min-MPJPE-epoch=13039.ckpt` | frozen in Stages I–III |
| SMPL body models | `deps/smpl_models/` | motion ↔ joint recovery |
| Instruction templates | `deps/mGPT_instructions/` | prompt templates (`DATASET.TASK_ROOT`) |
| Obstacle / free-space labels | `obstacle_labels/` | used by the QA pretraining + CoT |
| GPT-2 medium backbone | (auto-downloaded from Hugging Face) | LM backbone |
### Required for evaluation
| Artifact | Path | Notes |
|---|---|---|
| Text–motion evaluator | `checkpoints/evaluator/text_mot_match_trainset_only_len20/finest.tar` | matching / R-precision / FID |
| GloVe + mean/std | `deps/t2m/` | evaluator inputs |
### Stage checkpoints (released)
| Stage | Path | Consumed by |
|---|---|---|
| Stage I (pretrain) | `checkpoints/Pretrain_egovlm_egolmrep_pq_gpt2_medium_s2t_obstacle_with_sceneqa_grid64_s4o1/epoch=9.ckpt` | Stage II `PRETRAINED` |
| Stage II (instruction-tuned) | `checkpoints/Instruct_egovlm_stp2mt_reverse_from_s2t_obs_scene_s4o1_cot_grid64/min-ADE_head-epoch=18-step=21527.ckpt` | Stage III `PRETRAINED` + `GRPO_REF_MODEL_PATH` |
## 5. Training & Evaluation
All stages share a single entry point, `train_egovlm_stage2.py`; the stage is selected by
the config (`TRAIN.STAGE`). The motion tokenizer (Stage 0) is a prerequisite that produces
the frozen PQ-VAE used by Stages I–III.
### Stage 0 — Motion PQ-VAE tokenizer
```bash
CUDA_VISIBLE_DEVICES=0 python -m train_egovlm_stage2 \
--cfg configs/config_h3d_stage1_pq_4096_64.yaml \
--nodebug
```
### Stage I — Spatial-Semantic Scene Awareness Pretraining
```bash
CUDA_VISIBLE_DEVICES=0 python -m train_egovlm_stage2 \
--cfg configs/egovlm_gpt2_medium_stage2_pretrain/s2t_obstacle_with_sceneqa_s4o1.yaml \
--nodebug
```
### Stage II — Multi-Modal Multi-Task Instruction Tuning
```bash
CUDA_VISIBLE_DEVICES=0 python -m train_egovlm_stage2 \
--cfg configs/egovlm_gpt2_medium_stage3_stp2mt/from_s2t_obstacle_scene_s4o1_cot.yaml \
--nodebug
```
### Stage III — Multi-Modal Reward GRPO
```bash
CUDA_VISIBLE_DEVICES=0,1 python -m train_egovlm_stage2 \
--cfg configs/egovlm_gpt2_medium_grpo_stp2mt/from_obs_scene_cot_bs4_4.yaml \
--nodebug
```
### Evaluation
Evaluation reuses the same per-stage config and entry point `validate_egovlm_stage2.py`.
The checkpoint to evaluate is the one set in `TRAIN.PRETRAINED` of the stage config
(`validate_egovlm_stage2.py` loads `TRAIN.PRETRAINED` and runs `trainer.validate`).
So, to evaluate a trained model, point `TRAIN.PRETRAINED` at that checkpoint and run:
```bash
# Stage I (pretrain) — scene / obstacle QA metrics
CUDA_VISIBLE_DEVICES=0 python -m validate_egovlm_stage2 \
--cfg configs/egovlm_gpt2_medium_stage2_pretrain/s2t_obstacle_with_sceneqa_s4o1.yaml --nodebug
# Stage II (instruction-tuned) — 4-task pose / language metrics
CUDA_VISIBLE_DEVICES=0 python -m validate_egovlm_stage2 \
--cfg configs/egovlm_gpt2_medium_stage3_stp2mt/from_s2t_obstacle_scene_s4o1_cot.yaml --nodebug
# Stage III (GRPO) — set GRPO_VAL: True in the config; evaluates the GRPO policy
CUDA_VISIBLE_DEVICES=0 python -m validate_egovlm_stage2 \
--cfg configs/egovlm_gpt2_medium_grpo_stp2mt/from_obs_scene_cot_bs4_4.yaml --nodebug
```
`--device ` and `--batch_size ` may be passed to override the config. The evaluated
checkpoint is selected purely by `TRAIN.PRETRAINED`; for Stage II/III set it to the
trained Stage-II / Stage-III checkpoint rather than the pretrain checkpoint used for init.
## Citation
```bibtex
@inproceedings{ego3dlm2026,
title = {Ego-Human Motion Prediction with 3D-Aware LLM},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}
```