Instructions to use episod/tt-animatediff with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use episod/tt-animatediff with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("episod/tt-animatediff", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("episod/tt-animatediff", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]tt-animatediff
This is a port, not a checkpoint. This repository ships no model weights. It is an implementation of
AnimateDiff that runs the Stable Diffusion 1.4 UNet and VAE on a single Tenstorrent Blackhole chip
through TTNN. Motion coherence comes from cross-frame temporal attention (temporal_alpha), which
approximates the AnimateDiff MotionAdapter rather than running it. The package produces 512×512 text-to-video
GIFs over a small HTTP API, and it also works as a diffusers custom pipeline with a CPU fallback. The weights
come from their upstream repositories when you first pull or generate. This is a community bring-up and is
not a Tenstorrent-supported product.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 6, "thin").
GitHub · Docs · How it was built
At a glance
| Architecture | SD 1.4 UNet + VAE (TTNN) with cross-frame temporal attention on noise predictions (Phase 2.5) |
| Hardware | 1 Blackhole chip: p150 profile (mesh 1x1). Measured on one P300C chip. |
| Input / output limits | 512×512 only; 1-64 frames; 1-100 denoising steps; output is an animated GIF (base64) |
| License | Code Apache-2.0; SD 1.4 weights CreativeML Open RAIL-M (see Licensing) |
| Status | Experimental community bring-up. The bundle was verified by a fresh install and serve on 2026-09-27; there is no accuracy-vs-reference measurement yet |
| Model CI v0 | not yet run |
Intended use
Direct use: Short text-to-video clips (8-16 frames, 512×512) on a single Blackhole chip. Use it for demos and experiments, or as a reference for bringing a diffusion model up on TTNN.
Out-of-scope use:
- Resolutions other than 512×512. The UNet is compiled once at a 64×64 latent.
- Multi-chip serving. The server refuses a multi-chip mesh.
- Concurrent throughput. Requests are serialized on one device.
- Anything prohibited by the RAIL-M use restrictions on the weights.
- The full AnimateDiff MotionAdapter is not on the served path. It is only available through the GitHub
repo's CLI (
--motion-adapter, Phase 3).
Quickstart
uv tool install tenstorrent # once — the Tenstorrent CLI, `tt`
tt model pull episod/tt-animatediff
tt serve episod/tt-animatediff
tt model pull downloads the bundle and the SD 1.4 weights (CompVis/stable-diffusion-v1-4, pinned to
revision 133a221) into the bundle's own HF cache (<install>/.hf). Weights come down by default for a
bundle; tt has no --with-weights flag. The server needs only about 4 GB of SD 1.4 (unet/, vae/,
text_encoder/, tokenizer/), but the prefetch has no file filter and takes the whole repo, about 22 GB.
If disk is tight, run tt-model pull without --with-weights: the server then downloads just the 4 GB it
reads, at the same pinned revision, on first start.
Install builds a per-model venv from tt-model-requirements.txt: ttnn==0.77.0 from PyPI plus the two
wheels in wheels/. tt-model serve listens on port 20000, or the next free port. (run.sh run by hand
defaults to 8000.) The device is opened and the pipeline warmed inside the ASGI lifespan, before readiness is
reported. The server is ready when uvicorn logs Application startup complete.
Measured on a fresh install with an empty HF cache (one P300c chip, 2026-09-27):
- Launch to ready: 153 s. This includes the 4 GB weight download and the TTNN UNet build.
- First request: about 155 s (8 frames x 8 steps). Most of it is one-time kernel compile.
- After a warm restart: ready in about 10 s, and the first request takes about 13 s.
Without tt-cli — tt-model alone does the whole job:
tt-model pull episod/tt-animatediff --with-weights
tt-model serve episod/tt-animatediff
Serve profiles
| profile | hardware | mesh | frames / resolution |
|---|---|---|---|
p150 |
1 Blackhole chip (P150 declared; measured on one P300C chip) | P150, ANIMATEDIFF_MESH_SHAPE=1x1 |
1-64 frames, 512×512 only |
One chip is the model's limit, not a conservative default. A four-chip serve failed in the lifespan with
Can't convert a tensor distributed on MeshShape([1, 4]) mesh to row-major logical tensor, because the SD
demo's tensors are not sharded. The server now refuses a multi-chip mesh up front.
Using it
This is not an OpenAI-compatible chat API. /v1/models exists for tooling. The real endpoint is
OpenAI-shaped /v1/videos/generations, with the same field names tt-inference-server's media server uses.
| endpoint | purpose |
|---|---|
POST /v1/videos/generations |
generate a clip; returns base64 GIF |
GET /v1/models |
reports the weights repo (CompVis/stable-diffusion-v1-4) |
GET /health |
{} when ready, 503 while starting |
GET /tt-liveness |
liveness; answers while a generation is running |
curl -s http://localhost:20000/v1/videos/generations \
-H 'Content-Type: application/json' \
-d '{"prompt": "a swirling nebula, teal and gold, cinematic",
"num_frames": 8, "num_inference_steps": 25, "seed": 42}' \
| jq -r '.data[0].b64_json' | base64 -d > out.gif
Request fields, with server defaults:
| field | default | notes |
|---|---|---|
prompt |
required | |
negative_prompt |
"" |
|
num_frames |
16 | 1-64 (the diffusers pipeline defaults to 8) |
num_inference_steps |
25 | 1-100 |
guidance_scale |
7.5 | 0-20 |
seed |
42 | |
temporal_alpha |
0.35 | 0 gives Phase 2 shared noise; 1 gives full cross-frame attention |
height, width |
512 | only 512 is accepted |
response_format |
b64_json |
the only supported value |
The response is {"created": <unix time>, "data": [{"b64_json": "<GIF>"}]}. There is one denoise loop at a
time. Concurrent requests queue behind a device lock.
As a diffusers pipeline
The same repo also loads as a custom diffusers pipeline. trust_remote_code=True is required, which means
loading executes pipeline.py from this repo:
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"episod/tt-animatediff",
custom_pipeline="episod/tt-animatediff",
trust_remote_code=True,
)
# mode="auto": Blackhole if the ttnn runtime is importable, CPU otherwise.
frames = pipe("a swirling nebula, teal and gold, cinematic").frames
frames[0].save("out.gif", save_all=True, append_images=frames[1:], duration=125, loop=0)
print(pipe.resolved_mode) # "blackhole" or "cpu" — what actually ran
Loading is offline-safe and opens no device. Nothing is fetched until you call the pipeline.
model_index.json's base_model, motion_adapter and lightning_repo fields are declarative metadata, not
configurable inputs. They record the upstream weights the pipeline resolves, but __call__ never reads them.
Passing a different value at from_pretrained() time is accepted and changes nothing.
On a Blackhole box, mode="blackhole" requires the hardware instead of falling back. A missing runtime is
then an error rather than a silent 100× slowdown:
frames = pipe("a swirling nebula", mode="blackhole", num_frames=8, num_steps=25).frames
On any machine without hardware, the distilled 4-step Lightning weights make CPU tolerable:
frames = pipe(
"a swirling nebula", mode="cpu", use_lightning=True, lightning_steps=4,
num_frames=4, guidance_scale=1.0,
).frames
The CPU path uses the full AnimateDiff MotionAdapter (guoyww/animatediff-motion-adapter-v1-5-2), which is
why it is listed in base_model. The tt-model bundle never fetches it. The delegate package is imported if
installed, and otherwise fetched from this repo. To install it explicitly:
pip install 'animatediff-ttnn @ git+https://github.com/tenstorrent/tt-animatediff'
Implementation phases
| Phase | What runs where | Motion mechanism |
|---|---|---|
| 1 | CPU, diffusers AnimateDiffPipeline |
Full MotionAdapter |
| 2.5 | TTNN UNet on Blackhole (the served path, and this pipeline's default) | Cross-frame temporal attention (temporal_alpha) |
| 3 | TTNN UNet + MotionAdapter injected into the denoising loop | Full MotionAdapter, 7 injection points |
The HTTP server and mode="blackhole" both run Phase 2.5. mode="cpu" runs Phase 1. Phase 3 is reachable
only through the GitHub repo's CLI (--motion-adapter).
Expected performance
This bundle, served by tt-model serve:
| Configuration | Hardware | Accuracy vs reference | Latency, median (min-max) | Provenance |
|---|---|---|---|---|
| HTTP, 16 frames × 25 steps (server default) | 1× P300C chip | not yet measured | 30.72 s (30.54-30.87), 1.92 s/frame | published bundle 6ccfdce, N=3, 2026-09-27 |
| same, this bundle revision | 1× P300C chip | output GIFs byte-identical to the row above at seeds 1-3 | 30.68 s (30.46-30.86) | re-verified, N=3, 2026-09-27 |
| HTTP, 8 frames × 8 steps | 1× P300C chip | not yet measured | 6.78 s (6.73-6.91), 0.85 s/frame | published bundle 6ccfdce, N=3, 2026-09-27 |
| same, this bundle revision | 1× P300C chip | byte-identical to the row above | 6.72 s (6.69-6.91) | re-verified, N=3, 2026-09-27 |
Methodology:
- Setup: one Blackhole chip of a P300c (mesh 1x1,
ttnn==0.77.0from PyPI), prompt "a lighthouse in a storm, cinematic", 512×512, guidance 7.5,temporal_alpha0.35, seeds 1-3. - Timing: end-to-end HTTP wall time of
POST /v1/videos/generations, including base64 GIF encoding. One untimed 8×8 warmup came first, and N=3 requests were run per configuration. - Runs: the published bundle
6ccfdceand this bundle revision (weights133a221), both on 2026-09-27.
Earlier development measurements (tt-metal v0.77.0 host environment, not this bundle):
| Configuration | Latency | Provenance |
|---|---|---|
| HTTP, 8 frames × 4 / 8 / 16 steps | 5.149 / 7.363 / 11.983 s median (n=3) | docs/measurements/serving-benchmark.json @ 84045fb |
| HTTP, 16 frames × 8 steps | 14.675 s median (linear in frames, 1.99×) | same |
in-process generate_animation(), 8 frames × 25 steps / × 8 steps |
docs/model-card.md @ 3563466, measured 2026-09-07. Prose only; no raw log is committed |
|
mode="cpu" / CPU Lightning 4-step |
~2 min/frame / ~20 s/frame | earlier estimate, not re-verified |
- Latency model: it fits
2.84 s + 0.571 s × stepsat 8 frames. - Concurrency: three concurrent requests took 1.016× the serial time. Requests are serialized by design.
- Supersedes: these numbers replace an older claim of ~12.5 s/frame, which was 6.4× too slow.
Accuracy vs reference: not yet measured. No PCC, CLIP or SSIM comparison exists between the served Phase 2.5 output and a diffusers SD 1.4 reference. The only PCC figures on record cover kernel-level TT-Lang tests and a 4-chip determinism check, not the served output. The byte-identical result above is a regression check between two bundle revisions, not an accuracy measurement.
Limitations
- One chip, 512×512 only. Other sizes are rejected at the request edge. A multi-chip mesh is refused at startup.
- P150 is declared but not measured. Every Blackhole number above comes from one P300C chip.
pull --with-weightsover-fetches. The whole SD 1.4 repo (~22 GB) comes down, although serving reads ~4 GB.- Requests are serialized. No multi-user throughput exists, and none should be expected from this design.
- Phase 2.5 approximates AnimateDiff. Temporal attention runs on 4-channel noise predictions on CPU, not on the 320-dimensional UNet features the MotionAdapter uses. The full adapter (Phase 3) is CLI-only and much slower (~52 s/frame, not re-measured).
- The thin bundle installs
pytest. tt-metal v0.77.0'smodels/common/utility_functions.pyimports it at module scope. The fix (tt-metalfdcb5cf) is in v0.79.0 but not in v0.77.0. - The diffusers TTNN path needs a local build. It requires a Blackhole board and a local tt-metal build. Most users of the diffusers pipeline will only exercise the CPU path.
- Multi-chip frame counts. On multi-chip boards (CLI mesh frame sharding), the frame count must be a multiple of the chip count. 8 works on 1, 2 and 4 chips.
- No distilled LCM weights. The LCM distillation track is closed: all four runs failed, and no distilled
weights ship here.
use_lightning=Trueon CPU uses ByteDance's published checkpoint. mode="sim"is slow. The ttsim virtual Blackhole is bit-exact but 10-100× slower per op. It needs the simulator binary: passsim_so="/path/to/libttsim_bh.so", or leave it unset only if the binary is at~/sim/libttsim_bh.so.
Risks and safety considerations
- Output comes from SD 1.4 and inherits its known biases and failure modes. The RAIL-M use restrictions apply to anything you generate.
- Motion is an approximation. Frames can drift or flicker in ways a full AnimateDiff run would not.
- The numerical deviation from the diffusers reference on Blackhole (bf16 TTNN UNet) has not been measured.
- The diffusers pipeline needs
trust_remote_code=True, which executes this repo'spipeline.py.
Licensing
Read this before redistributing output. The code in this repository is Apache-2.0. The weights it downloads at runtime are not, and this repository cannot grant you their terms:
| Artifact | License |
|---|---|
| This repo's code | Apache-2.0 |
CompVis/stable-diffusion-v1-4 |
CreativeML Open RAIL-M |
ByteDance/AnimateDiff-Lightning |
CreativeML Open RAIL-M |
guoyww/animatediff-motion-adapter-v1-5-2 |
Undeclared upstream |
Using this package means accepting the RAIL-M use restrictions on the weights it fetches, even though this repository does not carry them. The MotionAdapter's publisher does not state its terms. If that matters to your use, resolve it with the upstream author rather than inferring permission from this repo's Apache-2.0 header.
A note on the model tree
The base_model entries make the Hub show stable-diffusion-v1-4 and the AnimateDiff MotionAdapter in this
repo's model tree. Read that as "this code resolves those weights at runtime". It does not mean the weights
were fine-tuned into a new checkpoint. Nothing here is trained, and the repo ships no weights of its own. The
Hub has no relation type for a reimplementation on different hardware, which is what this is.
Changelog
| date | change |
|---|---|
| 2026-09-28 | Repackaged with animatediff-ttnn 0.11.2 (tenstorrent/tt-animatediff main @ 5ca4ef9). The Phase 3 CLI's motion-adapter loader now defaults to the pinned adapter repo instead of a copied id. Served output is unchanged: re-verified on one Blackhole chip, byte-identical GIFs to 0.11.1. |
| 2026-09-27 | Repackaged with animatediff-ttnn 0.11.1 (GitHub PR #12). Every Hub load now uses a pinned revision: SD 1.4 at 133a221, overridable with TT_MODEL_WEIGHTS_REVISION; the CPU-path MotionAdapter and Lightning weights are pinned as well. The manifest records the weights revision and the correct tt_metal_version (0.77.0, previously a stale 0.65.1rc17.dev6200). Restored the Gradio app.py, which the 2026-09-15 bundle push had overwritten with the ASGI server; the server's runner copy is now tt_model_server.py. run.sh picks chips from the bundle's device count. Re-verified on hardware from a fresh install with an empty HF cache: 30.68 s for 16 frames × 25 steps, output identical to the previous bundle. |
| 2026-09-15 | Added the v6 thin tt-model bundle (tt-dit-server, 1 chip, ttnn 0.77.0). Moved the tt-model pins to tt-model-requirements.txt so the diffusers requirements.txt is untouched. |
| 2026-09-15 | Discovery tags added; manifest repo target corrected (GitHub #11). |
| 2026-09-08 | Post-merge follow-ups (GitHub #10): declared weights corrected to SD 1.4, a malformed mesh shape is refused, device rollback, and pytest shipped for the tt-metal v0.77.0 import. |
| 2026-09-07 | Blackhole performance re-measured (~1.94 s/frame at 25 steps, replacing ~12.5 s/frame); scheduler description corrected to Euler. |
| 2026-09-04 | Repo made public. |
| 2026-08-19 | First published as a weights-free diffusers custom pipeline. |
Related packages
episod/tt-animatediff-demois the public demo Space (CPU).- The GitHub repo also carries a v5.1 container manifest (
tt_model_package.yaml). It has not been published. Itsrepo:field currently names this repo; publishing it there would replace this bundle.
Feedback
Questions or problems with this package: open a discussion at
https://huggingface.co/episod/tt-animatediff/discussions. That is the one channel that reaches the bundle's
author. Code issues can also go to tenstorrent/tt-animatediff.
For a problem with the tt tooling itself, use tt report issue. It collects your environment and opens a
prefilled issue against tenstorrent/tt-cli, so it does not reach this package's author. Send product feedback
to support@tenstorrent.com.
Provenance
| component | built from |
|---|---|
| tt-metal / ttnn | ttnn==0.77.0 (PyPI), plus tt_animatediff_models_closure-0.77.0: the vendored models/common and the SD wormhole demo from tt-metal v0.77.0. The manifest tt_metal_version is 0.77.0 |
| model code | animatediff_ttnn-0.11.2 wheel and the root animatediff_ttnn/ tree, both from tenstorrent/tt-animatediff @ 5ca4ef9 (main). Wheel sha256 2bb79d9eb3b86c9b…. On the serve path, run.sh's PYTHONPATH loads the root tree; the root tree and the wheel carry the same code |
| closure wheel | sha256 daa6ccd28440c555… |
| weights | CompVis/stable-diffusion-v1-4 @ 133a221b8aa7292a167afc5127cb63fb5005638b, pinned in the manifest and the code. CPU path only: guoyww/animatediff-motion-adapter-v1-5-2 @ 6167b88ffe39b4441fdf2113e77b99a6f56b7906 and ByteDance/AnimateDiff-Lightning @ 027c893eec01df7330f5d4b733bc9485ee02e8b2 |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager (producer tt_kernel_version 0.1.0) · package-thin --kind tt-dit-server --app animatediff_ttnn.server.app:app |
- Downloads last month
- 57
Model tree for episod/tt-animatediff
Base model
CompVis/stable-diffusion-v1-4