File size: 7,904 Bytes
327bc1a
2fc3ec1
 
 
ad43b1b
 
 
2081f54
ad43b1b
 
 
 
 
 
 
89b8b1b
 
2081f54
 
 
 
 
 
2fc3ec1
2081f54
122ce0c
 
9d2f257
122ce0c
 
 
 
 
 
 
 
 
1e56214
9e4105a
122ce0c
 
 
9e4105a
2081f54
2fc3ec1
 
 
 
2081f54
 
122ce0c
 
9d2f257
122ce0c
9d2f257
122ce0c
 
 
 
2081f54
122ce0c
 
 
 
 
 
 
 
 
 
2081f54
 
 
122ce0c
2081f54
122ce0c
89b8b1b
122ce0c
9e4105a
122ce0c
 
 
 
 
 
 
bdaf273
2081f54
89b8b1b
 
122ce0c
89b8b1b
122ce0c
89b8b1b
122ce0c
9e4105a
122ce0c
 
9e4105a
 
 
 
 
 
 
 
89b8b1b
122ce0c
89b8b1b
122ce0c
89b8b1b
122ce0c
 
2081f54
122ce0c
89b8b1b
122ce0c
89b8b1b
122ce0c
2081f54
 
1e56214
2081f54
 
 
 
 
 
1e56214
122ce0c
2081f54
 
 
1e56214
2081f54
9d2f257
2081f54
 
 
 
 
122ce0c
 
 
 
 
 
 
2081f54
 
 
 
 
122ce0c
 
 
 
 
 
2fc3ec1
122ce0c
2fc3ec1
 
 
9e4105a
2fc3ec1
 
 
 
122ce0c
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
---
license: other
license_name: vtm-spark-mixed
license_link: https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md
language:
- en
pipeline_tag: image-to-image
library_name: pytorch
tags:
- image-to-image
- pose-transfer
- pose-guided
- human-pose
- keypoints
- diffusion
- dit
- rectified-flow
- vtuber
- anime
- windows
- cuda
---

# VTM-1.5.1

**Turn one anime picture into a live VTuber.** A keypoint-driven DiT draws your character in whatever pose and expression your face gives it, frame by frame.

These are the weights behind **[VTM Spark](https://vtmstudio.dev/download)**, the free Windows app that tracks you with a webcam or iPhone and sends the character to OBS, Discord or Zoom as a webcam.

**Status:** beta · **Platform:** Windows + NVIDIA CUDA

<table>
  <tr>
    <th>One still in</th>
    <th>Animated by this model</th>
  </tr>
  <tr>
    <td align="center"><img src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/character-blueprint.png" alt="Input still: chest-up anime character on a green background" width="340"></td>
    <td align="center"><video src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/demo-60fps.mp4" controls autoplay loop muted playsinline width="340"></video></td>
  </tr>
</table>

<sub>Every frame on the right was drawn by `VTM-1.5.1.pt` from the one picture on the left, recorded straight from the VTM Spark desk with Max FPS at 100 (no cap in practice) and the Ultra decoder on an RTX 5060 Ti: about 60 frames a second shown. The head motion is a recorded iPhone (iFacialMocap) session replayed through the app instead of a live face.</sub>

## How it works

<img src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/pipeline.png" alt="Webcam or iPhone, then Track Lab, then 37 keypoints, then the VTM-1.5.1 DiT (fed with your character picture), then the SD VAE, then the VTM Spark camera" width="100%">

---

## Use it

The easy way is **[VTM Spark](https://vtmstudio.dev/download)**. It downloads everything in this repo for you:

1. Download VTM Spark from [vtmstudio.dev/download](https://vtmstudio.dev/download).
2. Double-click `install.bat`.
3. Double-click `run.exe`, add your character picture, start tracking, and pick the **VTM Spark** camera in OBS / Discord / Zoom.

You'll need Windows 10/11 and an NVIDIA RTX 30, 40 or 50 series GPU.

### Your character picture

The model expects a still laid out like the one above:

- square, **768 × 768**, solid green `#00FF00` background
- chest-up, centred, facing the camera, chin level
- both eyes open, small closed-mouth smile
- no hands, props, text or extra people

VTM Spark ships this blueprint and a ready-to-paste image-AI prompt in `character-blueprint/`.

---

## What's in this repo

| File | Size | What it is |
|---|---|---|
| `VTM-1.5.1.pt` | 360 MB | The generator: keypoint-conditioned DiT (this model) |
| `decoder/vtm-fast-decoder.pt` | 4 MB | Ultra fast decoder: turns the DiT's latents into the 768 × 768 frame for live streaming |
| `trackers/iris_pose.pt` | 6 MB | Iris / pupil tracker on the character still (YOLO-pose fine-tune) |
| `trackers/dwpose_v2.pt` | 23 MB | Upper-body keypoints on the still (YOLO-pose fine-tune) |
| `trackers/animeseg_hair3.pt` | 432 MB | Hair-part segmentation (Mask2Former fine-tune), so hair follows the head |
| `trackers/pose_landmarker_lite.task` | 6 MB | MediaPipe pose landmarker for body tracking |
| `openseeface/*` | 21 MB | OpenSeeFace webcam face tracking models |
| `media/*` | | Demo images for this card |

VTM Spark places them under `models/dit/`, `models/decoder/`, `models/trackers/` and `vendor/tools/openseeface/models/`. On first setup it also fetches a few models straight from their publishers (SD VAE, a tiny VAE and two anime-face detectors), about 1.24 GB in all.

---

## Model

`VTM-1.5.1.pt`

- **Architecture:** DiT, **30.3M parameters** (hidden 320, depth 10, 5 heads, patch 4 on SD-VAE latents), with RoPE on keypoints, QK-norm, SwiGLU and RMSNorm
- **Output:** 768 × 768 through the SD VAE, or through the Ultra fast decoder while streaming
- **Pose input:** 37 keypoints covering the face outline, brows, eyes, irises, nose, mouth and upper body
- **Identity input:** one reference image of the character, as image tokens plus face tokens
- **Sampling:** rectified flow, **distilled to run in 1 step** from a 30-step teacher (`VTM-1.5.1-SFT`, a fine-tune of the 1.5.1 base). Guidance is baked into the weights, so run it at **1 step, pose CFG 1.0, identity CFG 1.0** (VTM Spark's defaults)
- **Speed:** built for live use in VTM Spark. On an RTX 5060 Ti the model plus Ultra decoder draws about 117 frames a second on the GPU; the desk shows about 60–65, limited by the app's CPU-side work. Frame rate depends on the GPU and the app's Batch, Inbetweens and Max FPS settings.

The previous `VTM-1.5.1.pt` (distilled from a 10-step teacher, 1 or 2 steps, CFG not baked in) is still in this repo's commit history.

### Ultra fast decoder

`decoder/vtm-fast-decoder.pt` is a small PixelShuffle decoder distilled from the Hybrid TinyVAE. It decodes the same SD latents in one pass (a whole live frame takes about 11 ms instead of 17 on an RTX 5060 Ti) and is what VTM Spark's "Ultra fast" stream mode runs. Without it the app falls back to the TinyVAE.

### Data

Training used a small private character set:

- ~**1,500** characters
- ~**16–30** images per character

Generation quality in this release is limited mainly by **model capacity (DiT-30M)** and **dataset scale**.

---

## Download

```bash
hf download sinBoo1/VTM-Spark VTM-1.5.1.pt --local-dir ./VTM-Spark
```

```python
from huggingface_hub import hf_hub_download

ckpt = hf_hub_download(
    repo_id="sinBoo1/VTM-Spark",
    filename="VTM-1.5.1.pt",
)
```

Everything at once: `hf download sinBoo1/VTM-Spark --local-dir ./VTM-Spark`

The runtime (live camera → keypoints → this model → virtual camera) is [VTM Spark](https://vtmstudio.dev/download).

---

## Limitations

- Beta: soft detail, identity drift and pose errors happen
- Small data and DiT-30M capacity limit image quality
- Windows + NVIDIA CUDA only (no AMD, macOS, Linux or CPU)
- **Framing:** torso-up only (roughly head to mid-torso). Legs and most of the waist are not supported
- **Hands:** not supported
- **Character types not supported:** realistic humans; non-humanoid / furries
- **Accessories:** glasses and hats generally work; most other accessories are not supported

---

## Intended use

Live VTubing and research on pose → image pipelines (live drive, pose retargeting). Not a finished production renderer.

---

## Licences

No single licence covers every file here, so this repo is tagged `license: other`. Apache License 2.0 covers **our** training work: `VTM-1.5.1.pt` and our tracker fine-tunes. It does **not** re-license anyone else's weights, and some files here start from third-party weights with their own terms:

| File | Licence | Commercial use |
|---|---|---|
| `VTM-1.5.1.pt` | Apache-2.0 (ours) | Yes |
| `decoder/vtm-fast-decoder.pt` | Our training is Apache-2.0, but it is distilled from and warm-started on `cqyan/hybrid-sd-tinyvae`, whose weight licence is **not declared** | **Open question** until that publisher states a licence |
| `trackers/animeseg_hair3.pt` | Our fine-tune of Mask2Former ADE20k weights, which Meta licenses **CC BY-NC 4.0** | **No** |
| `trackers/iris_pose.pt`, `trackers/dwpose_v2.pt` | Our fine-tunes of Ultralytics YOLO-pose pretrained weights, which Ultralytics licenses **AGPL-3.0** | Under AGPL terms |
| `trackers/pose_landmarker_lite.task` | Apache-2.0 (Google / MediaPipe) | Yes |
| `openseeface/*` | BSD 2-Clause ([emilianavt/OpenSeeFace](https://github.com/emilianavt/OpenSeeFace)) | Yes |

Full inventory and sources: [THIRD_PARTY_NOTICES.md](https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md) in the VTM Spark repo.