Instructions to use Plurato123/OctLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Plurato123/OctLLM with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Plurato123/OctLLM") model = AutoModelForMultimodalLM.from_pretrained("Plurato123/OctLLM", device_map="auto") - Notebooks
- Google Colab
- Kaggle
OctLLM
Octrees as an Explicit 3D Language
Ran Dan1, Si-Tong Wei1, Pengfei Xiong2, Wei Zhang2, Yadong Mu1, Peng-Shuai Wang1,â€
1Peking University 2Independent Researcher
†Corresponding author
Official model weights for Octrees as an Explicit 3D Language.
OctLLM unifies text-to-3D, image-to-3D, and 3D understanding through an explicit 3D language. It represents shapes as Sparse Octree (S-Octree) sequences: occupancy tokens anchored to 3D coordinates and octree depth. This structured representation connects geometry generation and understanding within a single model.
To learn the 3D modality, OctLLM adds trainable branches alongside a frozen Qwen2.5-VL backbone. Mesh tokens and text/image tokens follow separate processing pathways and interact through shared self-attention. Generated octrees are completed by a 3D U-Net and decoded into textured meshes with pretrained TRELLIS decoders. The video below demonstrates generation, 3D understanding, and language interaction.
This repository contains the OctLLM checkpoint and occupancy completion weights. Use the official implementation for installation and inference; its loading code is required for the custom 3D components. More results are available on the project page.
- Downloads last month
- 41
Model tree for Plurato123/OctLLM
Base model
Qwen/Qwen2.5-VL-7B-Instruct