Instructions to use JreamH/CGMLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JreamH/CGMLLM with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("JreamH/CGMLLM", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CG-MLLM (v0.1)
CG-MLLM: Captioning and Generating 3D Content via Multi-modal Large Language Models (ICML 2026)
CG-MLLM is a 3D multimodal large language model (3D MLLM) built upon Qwen3-VL and Hunyuan3D-2.1 VAE for unified 3D understanding, 3D captioning, and high-resolution 3D content generation.
It uses a Mixture-of-Transformer design: a TokenAR Transformer for token-level content and a BlockAR Transformer for block-level 3D latents, enabling long-context interaction between standard tokens and spatial blocks in one architecture.
Links
- Paper (arXiv): https://arxiv.org/abs/2601.21798
- Hugging Face Papers: https://huggingface.co/papers/2601.21798
- Project page: https://cv.jream.top/CG-MLLM-page/
- Code: https://github.com/dreaming-huang/CG-MLLM
- ICML poster: https://icml.cc/virtual/2026/poster/63909
Model Details
| Item | Value |
|---|---|
| Version | v0.1 |
| Architecture | Mixture-of-Transformer (TokenAR + BlockAR) |
| Vision-language backbone | Qwen3-VL-2B-Instruct |
| 3D latent tokenizer | Hunyuan3D-2.1 VAE |
| Tasks | Image-to-3D, text-to-3D, image understanding, 3D understanding |
| Venue | ICML 2026 |
| License | Apache-2.0 |
| Authors | Junming Huang, Chi Wang, Letian Li, Guangkai Xu, Donglin Huang, Hao Chen, Qiang Dai, Weiwei Xu |
| Affiliations | Zhejiang University; LIGHTSPEED |
Repository Contents
| File | Description |
|---|---|
ema.safetensors |
EMA checkpoint weights for CG-MLLM v0.1 |
How to Use
Install and run inference from the official code repository:
git clone https://github.com/dreaming-huang/CG-MLLM.git
cd CG-MLLM
# follow Installation in the GitHub README, then:
hf download JreamH/CGMLLM ema.safetensors --local-dir models/CGMLLM
Example (image-to-3D):
python inference.py \
--llm_base_path Qwen/Qwen3-VL-2B-Instruct \
--checkpoint models/CGMLLM \
--obj_vae_path tencent/Hunyuan3D-2.1 \
--obj_vae_len 4096 \
--use_qwen_vit --use_qwen_vl --qk_norm \
--mode i2obj --image examples/chairo.png
See the GitHub README for text-to-3D, image understanding, and 3D understanding examples.
Citation
If you find this work useful, please cite:
@article{huang2026cg,
title={CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models},
author={Huang, Junming and Wang, Chi and Li, Letian and Xu, Guangkai and Huang, Donglin and Chen, Hao and Dai, Qiang and Xu, Weiwei},
journal={arXiv preprint arXiv:2601.21798},
year={2026}
}
Acknowledgments
Inference code builds on BAGEL and the Hunyuan3D-2.1 shape VAE. Please follow their licenses when using those components.