CG-MLLM (v0.1)

CG-MLLM: Captioning and Generating 3D Content via Multi-modal Large Language Models (ICML 2026)

CG-MLLM is a 3D multimodal large language model (3D MLLM) built upon Qwen3-VL and Hunyuan3D-2.1 VAE for unified 3D understanding, 3D captioning, and high-resolution 3D content generation.

It uses a Mixture-of-Transformer design: a TokenAR Transformer for token-level content and a BlockAR Transformer for block-level 3D latents, enabling long-context interaction between standard tokens and spatial blocks in one architecture.

Links

Model Details

Item Value
Version v0.1
Architecture Mixture-of-Transformer (TokenAR + BlockAR)
Vision-language backbone Qwen3-VL-2B-Instruct
3D latent tokenizer Hunyuan3D-2.1 VAE
Tasks Image-to-3D, text-to-3D, image understanding, 3D understanding
Venue ICML 2026
License Apache-2.0
Authors Junming Huang, Chi Wang, Letian Li, Guangkai Xu, Donglin Huang, Hao Chen, Qiang Dai, Weiwei Xu
Affiliations Zhejiang University; LIGHTSPEED

Repository Contents

File Description
ema.safetensors EMA checkpoint weights for CG-MLLM v0.1

How to Use

Install and run inference from the official code repository:

git clone https://github.com/dreaming-huang/CG-MLLM.git
cd CG-MLLM
# follow Installation in the GitHub README, then:
hf download JreamH/CGMLLM ema.safetensors --local-dir models/CGMLLM

Example (image-to-3D):

python inference.py \
  --llm_base_path Qwen/Qwen3-VL-2B-Instruct \
  --checkpoint models/CGMLLM \
  --obj_vae_path tencent/Hunyuan3D-2.1 \
  --obj_vae_len 4096 \
  --use_qwen_vit --use_qwen_vl --qk_norm \
  --mode i2obj --image examples/chairo.png

See the GitHub README for text-to-3D, image understanding, and 3D understanding examples.

Citation

If you find this work useful, please cite:

@article{huang2026cg,
  title={CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models},
  author={Huang, Junming and Wang, Chi and Li, Letian and Xu, Guangkai and Huang, Donglin and Chen, Hao and Dai, Qiang and Xu, Weiwei},
  journal={arXiv preprint arXiv:2601.21798},
  year={2026}
}

Acknowledgments

Inference code builds on BAGEL and the Hunyuan3D-2.1 shape VAE. Please follow their licenses when using those components.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for JreamH/CGMLLM