VTM-Elf 0.01

Turn one anime picture into a live VTuber. A keypoint-driven DiT draws your character in whatever pose and expression your face gives it, frame by frame.

These are the weights behind VTM Spark, the free Windows app that tracks you with a webcam or iPhone and sends the character to OBS, Discord or Zoom as a webcam.

Status: beta ยท Platform: Windows + NVIDIA CUDA

One still in Animated by this model
Input still: chest-up anime character on a green background The same character turning, nodding, blinking and talking

Every frame on the right was drawn by VTM-1.5.1.pt from the one picture on the left. It was driven through VTM Spark's live tracking pipeline (head turn, nod and tilt, blinks, eye direction, mouth shapes), fed with scripted iPhone-style motion instead of a real face.


Use it

The easy way is VTM Spark. It downloads everything in this repo for you:

  1. Download or clone VTM Spark.
  2. Double-click install.bat.
  3. Double-click run.exe, add your character picture, start tracking, and pick the VTM Spark camera in OBS / Discord / Zoom.

You'll need Windows 10/11 and an NVIDIA RTX 30, 40 or 50 series GPU.

Your character picture

The model expects a still laid out like the one above:

  • square, 768 ร— 768, solid green #00FF00 background
  • chest-up, centred, facing the camera, chin level
  • both eyes open, small closed-mouth smile
  • no hands, props, text or extra people

VTM Spark ships this blueprint and a ready-to-paste image-AI prompt in character-blueprint/.


What's in this repo

File Size What it is
VTM-1.5.1.pt 360 MB The generator: keypoint-conditioned DiT (this model)
trackers/iris_pose.pt 6 MB Iris / pupil tracker on the character still (YOLO-pose fine-tune)
trackers/dwpose_v2.pt 23 MB Upper-body keypoints on the still (YOLO-pose fine-tune)
trackers/animeseg_hair3.pt 432 MB Hair-part segmentation (Mask2Former fine-tune), so hair follows the head
trackers/pose_landmarker_lite.task 6 MB MediaPipe pose landmarker for body tracking
openseeface/* 21 MB OpenSeeFace webcam face tracking models
media/* Demo images for this card

VTM Spark places them under models/dit/, models/trackers/ and vendor/tools/openseeface/models/.


Model

VTM-1.5.1.pt

  • Architecture: DiT, 30.3M parameters (hidden 320, depth 10, 5 heads, patch 4 on SD-VAE latents), with RoPE on keypoints, QK-norm, SwiGLU and RMSNorm
  • Output: 768 ร— 768 through the SD VAE
  • Pose input: 37 keypoints covering the face outline, brows, eyes, irises, nose, mouth and upper body
  • Identity input: one reference image of the character, as image tokens plus face tokens
  • Sampling: rectified flow, distilled from a 10-step teacher to run in 1 step (2 steps also supported)
  • Speed: built for live use in VTM Spark. Frame rate depends on the GPU and the app's Batch, Inbetweens and Max FPS settings.

Data

Training used a small private character set:

  • ~1,500 characters
  • ~16โ€“30 images per character

Generation quality in this release is limited mainly by model capacity (DiT-30M) and dataset scale.


Download

hf download sinBoo1/VTM-Spark VTM-1.5.1.pt --local-dir ./VTM-Spark
from huggingface_hub import hf_hub_download

ckpt = hf_hub_download(
    repo_id="sinBoo1/VTM-Spark",
    filename="VTM-1.5.1.pt",
)

Everything at once: hf download sinBoo1/VTM-Spark --local-dir ./VTM-Spark

The runtime (live camera โ†’ keypoints โ†’ this model โ†’ virtual camera) is VTM Spark.


Limitations

  • Beta: soft detail, identity drift and pose errors happen
  • Small data and DiT-30M capacity limit image quality
  • Windows + NVIDIA CUDA only (no AMD, macOS, Linux or CPU)
  • Framing: torso-up only (roughly head to mid-torso). Legs and most of the waist are not supported
  • Hands: not supported
  • Character types not supported: realistic humans; non-humanoid / furries
  • Accessories: glasses and hats generally work; most other accessories are not supported

Intended use

Live VTubing and research on pose โ†’ image pipelines (live drive, pose retargeting). Not a finished production renderer.


Licences

Apache License 2.0 covers our training work: VTM-1.5.1.pt and our tracker fine-tunes. It does not re-license anyone else's weights, and some files here start from third-party weights with their own terms:

File Licence
VTM-1.5.1.pt Apache-2.0 (ours)
trackers/animeseg_hair3.pt Our fine-tune of Mask2Former ADE20k weights, which Meta licenses CC BY-NC 4.0: non-commercial use only
trackers/iris_pose.pt, trackers/dwpose_v2.pt Our fine-tunes of Ultralytics YOLO-pose pretrained weights, which Ultralytics licenses AGPL-3.0
trackers/pose_landmarker_lite.task Apache-2.0 (Google / MediaPipe)
openseeface/* BSD 2-Clause (emilianavt/OpenSeeFace)

Full inventory and sources: THIRD_PARTY_NOTICES.md in the VTM Spark repo.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support