# Ego3DLM: Ego-Human Motion Prediction with 3D-Aware LLM **ECCV 2026**
Ego3DLM forecasts human motion from an egocentric perspective by grounding a language model in the **3D spatial and semantic context** of the surrounding environment. Given three-point motion tracking (head + hands), 3D scene features, and egocentric video, Ego3DLM **simultaneously decodes past pose, future pose, past narration, and future narration in a single autoregressive pass**, so that predicted poses and descriptions are grounded in one another for cross-modal and temporal consistency. The model is trained in three stages on top of a frozen motion tokenizer: 1. **Stage I — Spatial-Semantic Scene Awareness Pretraining.** Encode the 3D scene and inject it into the LM, training spatial and semantic scene understanding (scene QA + obstacle/free-space QA). 2. **Stage II — Multi-Modal Multi-Task Instruction Tuning.** Holistically train past/future pose tracking and prediction alongside past/future narration in a single pass. 3. **Stage III — Multi-Modal Reward GRPO.** Reinforcement finetuning with intra- and inter-modal rewards that directly optimize pose–language fidelity. Experiments on the **Nymeria** benchmark show state-of-the-art performance on future pose prediction, past motion tracking, and language description. ## 0. News - **2026** — Ego3DLM is accepted to **ECCV 2026**. Code released for reproducibility. ## 1. Environment Setup Tested with Python 3.11, CUDA 11.8, PyTorch 2.0.0, Transformers 4.46.3, PyTorch-Lightning 2.0.0. ```bash # create env conda create -n ego3dlm python=3.11 -y conda activate ego3dlm # install pinned dependencies pip install -r requirements-freeze.txt ``` The repository expects three top-level symlinks pointing to your data / dependency / checkpoint roots (see [Section 4](#4-dependencies--pretrained-models)): ```bash ln -s /path/to/checkpoints ./checkpoints ln -s /path/to/datasets ./datasets ln -s /path/to/deps ./deps ``` The obstacle / free-space QA labels (`obstacle_labels/`, required by Stage-I pretraining and CoT) are distributed separately — download them and place the directory at `./obstacle_labels`. ## 2. Dataset *(To be added.)* ## 3. Data Preprocessing *(To be added.)* ## 4. Dependencies & Pretrained Models All paths below are relative to the repository root and resolved through the `checkpoints/`, `datasets/`, and `deps/` symlinks. ### Required for every stage | Artifact | Path | Notes | |---|---|---| | Motion PQ-VAE tokenizer | `checkpoints/VQVAE_pq_full_4096_64/min-MPJPE-epoch=13039.ckpt` | frozen in Stages I–III | | SMPL body models | `deps/smpl_models/` | motion ↔ joint recovery | | Instruction templates | `deps/mGPT_instructions/` | prompt templates (`DATASET.TASK_ROOT`) | | Obstacle / free-space labels | `obstacle_labels/` | used by the QA pretraining + CoT | | GPT-2 medium backbone | (auto-downloaded from Hugging Face) | LM backbone | ### Required for evaluation | Artifact | Path | Notes | |---|---|---| | Text–motion evaluator | `checkpoints/evaluator/text_mot_match_trainset_only_len20/finest.tar` | matching / R-precision / FID | | GloVe + mean/std | `deps/t2m/` | evaluator inputs | ### Stage checkpoints (released) | Stage | Path | Consumed by | |---|---|---| | Stage I (pretrain) | `checkpoints/Pretrain_egovlm_egolmrep_pq_gpt2_medium_s2t_obstacle_with_sceneqa_grid64_s4o1/epoch=9.ckpt` | Stage II `PRETRAINED` | | Stage II (instruction-tuned) | `checkpoints/Instruct_egovlm_stp2mt_reverse_from_s2t_obs_scene_s4o1_cot_grid64/min-ADE_head-epoch=18-step=21527.ckpt` | Stage III `PRETRAINED` + `GRPO_REF_MODEL_PATH` | ## 5. Training & Evaluation All stages share a single entry point, `train_egovlm_stage2.py`; the stage is selected by the config (`TRAIN.STAGE`). The motion tokenizer (Stage 0) is a prerequisite that produces the frozen PQ-VAE used by Stages I–III. ### Stage 0 — Motion PQ-VAE tokenizer ```bash CUDA_VISIBLE_DEVICES=0 python -m train_egovlm_stage2 \ --cfg configs/config_h3d_stage1_pq_4096_64.yaml \ --nodebug ``` ### Stage I — Spatial-Semantic Scene Awareness Pretraining ```bash CUDA_VISIBLE_DEVICES=0 python -m train_egovlm_stage2 \ --cfg configs/egovlm_gpt2_medium_stage2_pretrain/s2t_obstacle_with_sceneqa_s4o1.yaml \ --nodebug ``` ### Stage II — Multi-Modal Multi-Task Instruction Tuning ```bash CUDA_VISIBLE_DEVICES=0 python -m train_egovlm_stage2 \ --cfg configs/egovlm_gpt2_medium_stage3_stp2mt/from_s2t_obstacle_scene_s4o1_cot.yaml \ --nodebug ``` ### Stage III — Multi-Modal Reward GRPO ```bash CUDA_VISIBLE_DEVICES=0,1 python -m train_egovlm_stage2 \ --cfg configs/egovlm_gpt2_medium_grpo_stp2mt/from_obs_scene_cot_bs4_4.yaml \ --nodebug ``` ### Evaluation Evaluation reuses the same per-stage config and entry point `validate_egovlm_stage2.py`. The checkpoint to evaluate is the one set in `TRAIN.PRETRAINED` of the stage config (`validate_egovlm_stage2.py` loads `TRAIN.PRETRAINED` and runs `trainer.validate`). So, to evaluate a trained model, point `TRAIN.PRETRAINED` at that checkpoint and run: ```bash # Stage I (pretrain) — scene / obstacle QA metrics CUDA_VISIBLE_DEVICES=0 python -m validate_egovlm_stage2 \ --cfg configs/egovlm_gpt2_medium_stage2_pretrain/s2t_obstacle_with_sceneqa_s4o1.yaml --nodebug # Stage II (instruction-tuned) — 4-task pose / language metrics CUDA_VISIBLE_DEVICES=0 python -m validate_egovlm_stage2 \ --cfg configs/egovlm_gpt2_medium_stage3_stp2mt/from_s2t_obstacle_scene_s4o1_cot.yaml --nodebug # Stage III (GRPO) — set GRPO_VAL: True in the config; evaluates the GRPO policy CUDA_VISIBLE_DEVICES=0 python -m validate_egovlm_stage2 \ --cfg configs/egovlm_gpt2_medium_grpo_stp2mt/from_obs_scene_cot_bs4_4.yaml --nodebug ``` `--device ` and `--batch_size ` may be passed to override the config. The evaluated checkpoint is selected purely by `TRAIN.PRETRAINED`; for Stage II/III set it to the trained Stage-II / Stage-III checkpoint rather than the pretrain checkpoint used for init. ## Citation ```bibtex @inproceedings{ego3dlm2026, title = {Ego-Human Motion Prediction with 3D-Aware LLM}, booktitle = {European Conference on Computer Vision (ECCV)}, year = {2026} } ```