Ditto talking head, ONNX Runtime repack
The ONNX models of Ditto.
Only the files the ONNX Runtime pipeline loads are included. Upstream's TensorRT engines, PyTorch
checkpoints, TensorRT plugin, configuration pickles, warp_network.onnx (which needs a custom TensorRT
operator) and the fp32 decoder.onnx are left out.
Attribution
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
Tianqi Li, Ruobing Zheng, Minghui Yang, Jingdong Chen, Ming Yang. Ant Group. ACM MM 2025.
- Paper: https://arxiv.org/abs/2411.19509
- Project page: https://digital-avatar.github.io/ai/Ditto/
- Code: https://github.com/antgroup/ditto-talkinghead
- Original models: https://huggingface.co/digital-avatar/ditto-talkinghead
Ditto's implementation is based on LivePortrait and S2G-MDDiffusion.
@article{li2024ditto,
title={Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis},
author={Li, Tianqi and Zheng, Ruobing and Yang, Minghui and Chen, Jingdong and Yang, Ming},
journal={arXiv preprint arXiv:2411.19509},
year={2024}
}
Source
Taken from digital-avatar/ditto-talkinghead at commit
e4a2f60328ee7c32af585ac4b3cce299e4c8e254,
folder ditto_onnx/.
Modifications
warp_network_opset20.onnx, fromwarp_network_ori.onnx. ItsGridSamplenodes are 5-D, which ONNX only allows from opset 20, but the graph is declared at opset 17, so ONNX Runtime refuses to load it. The default-domain opset import is set to 20, the IR version to 9, and eachGridSamplemodeattribute is renamed to its opset 20 spelling (bilineartolinear,bicubictocubic). No weights or other nodes change.decoder_fp16.onnx, fromdecoder.onnx. Converted to float16 withonnxconverter-common1.16.0 (float16.convert_float_to_float16,keep_io_types=True,disable_shape_infer=True), saved withonnx1.22.0. Inputs and outputs stay float32.
Every other file is byte-for-byte the upstream file.
License
Ditto's models are released by Ant Group under the Apache License, Version 2.0; see LICENSE and NOTICE. The modifications above are distributed under the same license.
Model tree for voxta/ditto-talkinghead-onnx
Base model
digital-avatar/ditto-talkinghead