tt-skyreels

SkyReels-V2-DF-1.3B-540P is a text-to-video diffusion transformer, and this package runs it on Tenstorrent Blackhole through TTNN. It uses one P300x2 board (4 chips) as a 2x2 QB2 mesh. It reuses tt-metal's WanTransformer3DModel wholesale, because SkyReelsV2's transformer is weight-compatible with it. A small ASGI server wraps the model, following the tt-dit-server pattern that tt-animatediff established. The package produces short 480x272 clips at 24 fps from a text prompt. This is an experimental community bring-up. It is verified on hardware to generate valid, deterministic video, and its latency has been measured. Accuracy against a reference has not been measured yet.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 6, v6 thin bundle). It installs with pip into a venv; it is not a container image.

At a glance

Architecture SkyReels-V2 DF 1.3B diffusion transformer (Wan-compatible), UMT5 text encoder, Wan VAE
Hardware p300x2: 4 Blackhole chips, 2x2 QB2 mesh
Context / input limits Text prompt in. Output is fixed at 480x272, 24 fps, 1-97 frames with (N-1) % 4 == 0. The default is 33 frames (about 1.4 s).
License Weights: Skywork Community License (other). Code: Apache-2.0. See Licensing below.
Status Experimental. Generates valid video with measured latency; accuracy vs a reference is not yet measured
Model CI v0 not yet run

Intended use

Direct use: Generating short text-to-video clips on a local Tenstorrent P300x2 (QB2) box through a simple HTTP endpoint. The package also serves as a reference example of a tt-dit-server v6 thin bundle.

Out-of-scope use:

  • Image-to-video. The SkyReels-V2-I2V-14B sibling is not packaged.
  • Resolutions other than 480x272.
  • Diffusion-forcing / long-video autoregressive extension.
  • Concurrent multi-user serving.
  • Hardware other than a 4-chip 2x2 mesh.
  • Anything the Skywork Community License disallows. See Licensing below.

Quickstart

uv tool install tenstorrent   # once: the Tenstorrent CLI, `tt`
tt model pull episod/tt-skyreels
tt serve episod/tt-skyreels

tt model pull downloads two things into the bundle's own HF cache (<install>/.hf):

  • This bundle: a small venv-install recipe (install.sh/run.sh) plus two wheels. ttnn, torch, the diffusers stack and the base HTTP stack resolve from an index at install time.
  • The Skywork/SkyReels-V2-DF-1.3B-540P-Diffusers weights, pinned to revision 958acd6, about 29 GB. Most of that is the UMT5 text encoder. For a bundle, weights come down by default, so tt has no --with-weights flag.

The server listens on port 20000, or the next free port if that one is busy. It is ready when it logs Application startup complete.

Measured on a fresh install with an empty HF cache (QuietBox 2, 2026-09-27):

step time
pull --with-weights (venv plus 29 GB of weights) 8 min 13 s
Launch to ready 30 s
First request 89 s, which includes about 30 s of one-time kernel compile
Warm restart to ready 19 s or less

Without tt-cli, tt-model alone does the whole job:

tt-model pull  episod/tt-skyreels --with-weights
tt-model serve episod/tt-skyreels

Serve profiles

profile hardware mesh
default p300x2 (4 Blackhole chips) QB2, SKYREELS_MESH_SHAPE=2x2

The server code also accepts a 1x4 mesh (for example P150x4), but no bundle ships it and this package has not validated it.

Using it

This is not an OpenAI chat-completions API. It is a video endpoint whose field names match tt-inference-server's and tt-animatediff's /v1/videos/generations.

curl localhost:20000/v1/videos/generations \
  -H 'Content-Type: application/json' \
  -d '{"prompt": "a red bicycle leaning against a brick wall in soft morning light",
       "num_frames": 33, "num_inference_steps": 20, "seed": 0}' \
  | python3 -c 'import sys,json,base64; open("out.mp4","wb").write(base64.b64decode(json.load(sys.stdin)["data"][0]["b64_json"]))'

Request fields for POST /v1/videos/generations:

field type default notes
prompt string required non-empty
negative_prompt string ""
num_frames int 33 1-97. Must satisfy (N-1) % 4 == 0: 9, 13, 17, ..., 33, 65, 97.
num_inference_steps int 20 1-100
guidance_scale float 6.0 0-20. SkyReels recommends 5-7.
seed int 0
response_format string b64_json This is the only format supported; any other value returns 400.
model string none accepted and ignored

The response is {"created": <unix time>, "data": [{"b64_json": "<base64 H.264 MP4>"}]}.

Other routes:

  • GET /health returns 200 when ready and 503 while loading.
  • GET /tt-liveness returns 200 whenever the process responds.
  • GET /v1/models reports the weights repo id.

The server takes one request at a time. The pipeline owns the mesh, so it serialises concurrent requests on a lock, and /health still answers during a generation.

A local Gradio UI (port 7861) and the full source live in tsingletaryTT/tt-skyreels.

Expected performance

profile frames steps resolution latency per video, median (min-max) on-device denoise accuracy vs reference
p300x2 / QB2 2x2 33 20 (server default) 480x272 @ 24 fps 57.64 s (57.50-58.18) on the published bundle; 60.02 s (59.99-60.29) on this revision about 4.2 s (4.7 it/s, about 0.21 s/step) not yet measured
p300x2 / QB2 2x2 33 8 480x272 @ 24 fps 54.30 s (53.92-54.31) on the published bundle; 58.46 s (58.24-58.98) on this revision about 1.7 s not yet measured

Methodology: on a QuietBox 2 (2 x p300c, 4 Blackhole chips, 2x2 QB2 mesh, FABRIC_1D, ttnn==0.78.0 from PyPI), the prompt "A red bicycle leaning against a brick wall on a sunny street, cinematic" was sent at the server defaults (480x272, 33 frames, guidance 6.0) with seeds 1000-1002, one request at a time, after one untimed warmup, timing client-side from POST to the complete base64 MP4 response: N=3 per configuration on the published bundle 69473ad and on this bundle revision (weights 958acd6), both on 2026-09-27.

How to read these numbers:

  • Where the time goes. The TT transformer denoise loop is about 4 s of a roughly 58 s request. The rest is the UMT5 text encoder and the VAE decoder on the host CPU; the split between the two was not measured. So 8 steps saves only a few seconds over 20, and a wall-clock "s/step" figure means nothing.
  • Why this revision reads slower. It is 4-8% slower than the published bundle, but its on-device loop ran at the same 4.7 it/s. The difference is host-side, and it was measured while other workloads shared the host CPU.
  • Determinism. Each seed produced a byte-identical MP4 across the warmup, the timed runs, a full server restart, and both bundle revisions.
  • Accuracy vs reference: no transformer-output PCC or perceptual comparison against the diffusers CPU pipeline has been made. Checked by eye, frame 16 at seed 1000 shows a coherent red bicycle against a cream wall on a sunlit street.

Limitations

  • Text-to-video only. SkyReels-V2-I2V-14B-540P (image-to-video) has not been started.
  • Fixed 480x272 output. No other resolution has been validated against the compiled TTNN graph.
  • Plain T2V pipeline, not diffusion-forcing. The port builds diffusers' SkyReelsV2Pipeline, not the SkyReelsV2DiffusionForcingPipeline the upstream repo advertises. That means no long-video autoregressive extension. The SkyReels-only fps_embedding and fps_projection weights are dropped at load time.
  • Only the transformer runs on Tenstorrent. The UMT5 text encoder (bf16) and the Wan VAE decoder (fp32) run on the host CPU.
  • One request at a time. There is no batching or concurrency.
  • Validated on one topology only: p300x2 as a 2x2 QB2 mesh.
  • The bundle ships kernel_patch/. The ttnn 0.78.0 PyPI wheel omits three fabric kernel sources that FABRIC_1D needs. The bundle vendors them byte-for-byte from tt-metal v0.78.0 and adds them to the search path with TT_METAL_KERNEL_PATH (set in run.sh). tt-metal's wheel packaging list (setup.py) still omits these files at v0.79.0 and on main, so the patch is expected to stay needed. See kernel_patch/README.md.
  • No accuracy measurement yet. See the note under Expected performance.
  • pull --with-weights downloads the whole weights repo, about 29 GB, including two small assets/ images the server never reads.

Risks and safety considerations

  • Kernel-path override. TT_METAL_KERNEL_PATH puts the vendored v0.78.0 kernels ahead of the installed ttnn's own tree. If you swap in a different ttnn version without removing kernel_patch/, the stale kernels could shadow newer ones. Keep ttnn==0.78.0 with this bundle.
  • Dropped weights. Dropping fps_embedding changes conditioning relative to the reference. No reference comparison exists yet, so visual-quality differences from upstream SkyReels have not been measured.
  • Generative-video content. Output can be inaccurate, and the model can produce inappropriate content. The Skywork license prohibits unlawful use and use that threatens national or societal security, and it asks for security review before any internet-facing deployment.

Licensing

component license
Weights (Skywork/SkyReels-V2-DF-1.3B-540P-Diffusers) Skywork Community License (license: other, skywork-license). The upstream terms allow commercial use subject to that license's conditions, and they disclaim liability. Read the upstream LICENSE and the linked Skywork Community License PDF before use. This bundle never embeds the weights; they download to your own HF cache.
Vendored tt-metal code (tt_skyreels_models_closure, kernel_patch/) Apache-2.0 (tenstorrent/tt-metal)
Port and serving code (skyreels_ttnn) Apache-2.0 (tsingletaryTT/tt-skyreels)

Changelog

date change
2026-09-27 Repackaged with skyreels-ttnn 0.1.1. All five from_pretrained calls now load the pinned weights revision 958acd6, overridable with TT_MODEL_WEIGHTS_REVISION. The manifest records that revision and tt_metal_version 0.78.0. run.sh picks chips from the bundle's device count. Corrected the card's weights license (Skywork Community License) and step default (20). ttnn==0.78.0, the closure wheel and kernel_patch/ are unchanged. Re-verified on hardware from a fresh install with an empty HF cache: 60.02 s for 33 frames x 20 steps, output byte-identical to the previous bundle.
2026-09-15 Repackaged as a v6 thin bundle (pip/venv) in place of the v5.1 container. Added kernel_patch/ for the ttnn 0.78.0 wheel's missing fabric kernels. Pinned tt-metal to v0.78.0. Re-verified on hardware with valid 480x272 MP4 output.
2026-09-09 First release as a v5.1 container. Fixed five bring-up bugs found only by serving on hardware: a module-scope ttnn import, missing pytest and ftfy dependencies, a missing .device shim, and a timestep dtype mismatch.

Related packages

  • tenstorrent/tt-animatediff: the reference tt-dit-server package this one follows.
  • SkyReels-V2-I2V-14B-540P (image-to-video): not yet packaged.

Feedback

  • Questions or problems with this package: open a discussion at https://huggingface.co/episod/tt-skyreels/discussions. It is the one channel that reaches the bundle's author.
  • A problem with the tt tooling itself: use tt report issue. It collects your environment and opens a prefilled issue against tenstorrent/tt-cli; it does not reach this package's author.
  • Product feedback: support@tenstorrent.com.

Provenance

component built from
tt-metal / ttnn v0.78.0: ttnn==0.78.0 from PyPI, with models/tt_dit vendored from the v0.78.0 tag. The manifest tt_metal_version is 0.78.0
skyreels-ttnn-0.1.1 wheel tsingletaryTT/tt-skyreels @ 28db465: session.py, the ASGI app and pipeline_skyreels.py, built by packaging/package-thin.sh from the repo's own setup.py. sha256 099ef8300b1e7a21…
tt_skyreels_models_closure-0.78.0 wheel tt-metal's models/tt_dit (minus tests/, experimental/, reference/) plus models/common/{utility_functions.py, device_utils.py, modules/tt_ccl.py}. This is a narrower local stand-in, not the upstream tt-metal-models package.
kernel_patch/ fabric_router_mux_extension.cpp, fabric_router_relay_extension.cpp and fabric_router_udm_mux_extension.cpp, byte-for-byte from tt-metal v0.78.0 tt_metal/fabric/impl/kernels/edm_fabric/.
weights Skywork/SkyReels-V2-DF-1.3B-540P-Diffusers @ 958acd63685c7e632e4b194549f2a703e34bd98b, pinned in the manifest and the code
Model CI v0 not yet run
build 2026-09-27 · tt-model-manager (producer tt_kernel_version 0.1.0), staged by packaging/package-thin.sh
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for episod/tt-skyreels

Finetuned
(1)
this model