Alchemy3D / README.md
libadi's picture
Update README.md
25a41f0 verified
|
Raw History Blame Contribute Delete
6.41 kB
metadata
license: apache-2.0
tags:
  - 3d
  - 3d-editing
  - asset-editing
  - image-to-3d
  - diffusion
pipeline_tag: image-to-3d

Alchemy3D Alchemy3D

Scaling Versatile 3D Assets Editing with a Million-Scale Dataset

Badi Li1,2,4, Tianxin Huang1, Yu Zhou3, Wei-Shi Zheng2,4, Yi Ma1,2, Shenghua Gao1,2†

1 The University of Hong Kong    2 Shenzhen Loop Area Institute
3 Shanghai Innovation Institute    4 Sun Yat-Sen University

† Corresponding author

Project Page arXiv PDF GitHub Hugging Face Model ModelScope Model Dataset Benchmark Evaluation

Alchemy3D teaser: addition, removal, replacement, local/global appearance, and animation.

Alchemy3D is the primary checkpoint of an open-sourced foundation model for editing existing 3D assets while preserving identity and structure. Given a source mesh (e.g. .glb) and a target reference image, it performs versatile edits including Addition, Removal, Replacement, Local/Global Appearance, and Animation.

The model is trained on Alchemy3D-1M (1.25M unique assets, 1.38M edit pairs, 7 edit types) and builds on TRELLIS.2 structured latents.

Model Variants

Weights are mirrored on Hugging Face and ModelScope. If Hugging Face is unreachable, load the ModelScope IDs below — the official code falls back automatically.

Model Description Hugging Face ModelScope
Alchemy3D Primary image-conditioned editing model libadi/Alchemy3D libd55/Alchemy3D
Alchemy3D-Turbo Step-distilled variant for faster inference with competitive quality libadi/Alchemy3D-Turbo libd55/Alchemy3D-Turbo
Alchemy3D-Flux Replaces DINOv3 with a Flux2 encoder for stronger PBR / appearance editing libadi/Alchemy3D-Flux libd55/Alchemy3D-Flux
Alchemy3D-Instruct Instruction-driven editing from natural-language text instead of a target image libadi/Alchemy3D-Instruct libd55/Alchemy3D-Instruct
Alchemy3D-Segment Downstream 3D part segmentation from 1–8 multi-view 2D segmentation maps libadi/Alchemy3D-Segment libd55/Alchemy3D-Segment

Quick Start

Install and run from the official repository.

from alchemy3d.pipelines import Pipeline
import o_voxel

pipe = Pipeline.from_pretrained("libadi/Alchemy3D")
pipe.cuda()

output = pipe.run(
    source="./assets/examples/edits/01/source.glb",
    image="./assets/examples/edits/01/target_image.png",
    comparison_video="example.mp4",
)[0]

glb = o_voxel.postprocess.to_glb(
    vertices=output.vertices,
    faces=output.faces,
    attr_volume=output.attrs,
    coords=output.coords,
    attr_layout=output.layout,
    voxel_size=output.voxel_size,
    aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
    decimation_target=1_000_000,
    texture_size=4096,
    remesh=True,
    remesh_band=1,
    remesh_project=0,
    verbose=False,
)
glb.export("example.glb")

Hardware: NVIDIA GPU recommended; ~24GB VRAM is a practical minimum. See the GitHub README for full environment setup (setup.sh).

Citation

If you use Alchemy3D, please cite:

@misc{li2026scalingversatile3dassets,
      title={Scaling Versatile 3D Assets Editing with a Million-Scale Dataset}, 
      author={Badi Li and Tianxin Huang and Yu Zhou and Wei-Shi Zheng and Yi Ma and Shenghua Gao},
      year={2026},
      eprint={2609.34271},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.34271}, 
}

License

Apache License 2.0. See the GitHub repository for dependency licenses (O-Voxel / TRELLIS.2, nvdiffrast, nvdiffrec, etc.).