wlsaidhi commited on
Commit
76878a0
·
verified ·
1 Parent(s): b598adf

Model card: public quickstart via FastVideo PR #1852 (basic_fasth3_8step.py); drop private/compatibility-hold text; record GPU verification

Browse files
Files changed (3) hide show
  1. INFERENCE.md +20 -28
  2. README.md +25 -11
  3. RELEASE_CHECKLIST.md +9 -8
INFERENCE.md CHANGED
@@ -2,18 +2,13 @@
2
 
3
  ## Compatibility status
4
 
5
- Public FastVideo `main` was checked at commit
6
- `a943220c115228ade5d57b3bab9a6a87fd600a10` on 2026-09-08.
7
- Its H3 pipeline rejects video scheduler shifts other than 12. This checkpoint
8
- requires 10. Its public `basic_fasth3.py` also lacks `--video-scheduler-shift`
9
- and `--audio-scheduler-shift`. Do not advertise that unchanged example as a
10
- working quickstart for this model.
11
-
12
- The existing internal FastVideo checkout at
13
- `0f2334f6b5d2ca2fafad15b95932df0e3d1a77c0` exposes both options and applies
14
- them to the scheduler instances. The command below is for reviewers who
15
- already have that compatible environment. It is not a public installation
16
- recipe, and no new end-to-end GPU run was performed during release preparation.
17
 
18
  ## Exact sampling contract
19
 
@@ -30,32 +25,29 @@ scheduler shifts must also be applied. `fastvideo_inference.json` records
30
  the same settings; verify that your launcher applies them rather than
31
  assuming its defaults match this file.
32
 
33
- ## Private-review command
34
 
35
- Run in an already configured compatible FastVideo checkout, on a GPU worker:
36
 
37
  ```bash
38
- export FASTVIDEO_DMD_DENOISING_STEPS=999,874,749,624,500,375,250,125
39
- unset FASTVIDEO_DMD_STOCHASTIC_RENOISE
40
-
41
- python examples/inference/basic/basic_fasth3.py \
42
- --model-path FastVideo/FastVideo-FastH3-8-Step-V2 \
43
  --prompt 'integrated_multimodal_description: A red fox runs through fresh snow at dawn. overall_soundscape: Fast pawsteps in snow, winter wind, and distant birds.' \
44
  --height 768 --width 1344 --num-frames 124 \
45
- --steps 9 \
46
- --video-scheduler-shift 10 --audio-scheduler-shift 3 \
47
- --num-gpus 4 \
48
- --vsa-sparsity 0.8 --vsa-tile-size 64 --vsa-kernel sm100a \
49
  --output outputs/fasth3-v2 \
50
  --repeats 1
51
  ```
52
 
 
 
 
53
  This example requests 1344×768, 124 frames at 24 fps (about 5.17 seconds),
54
- and writes `outputs/fasth3-v2/fasth3.mp4`. The `sm100a` route needs a compatible
55
  Blackwell GPU and kernel build. Triton is a different kernel option, not a
56
  license to change the trained sparsity, tile geometry, or shifts.
57
 
58
- Before public release, verify a clean public install, eight actual student
59
- forwards, video/audio shifts 10/3, VSA80 tile64, successful MP4 audio muxing,
60
- and representative high-motion output quality. Report any latency with the
61
- exact hardware, resolution, frame count, attention path and warmup policy.
 
2
 
3
  ## Compatibility status
4
 
5
+ FastVideo [PR #1852](https://github.com/hao-ai-lab/FastVideo/pull/1852) adds
6
+ the runtime support this checkpoint needs: scheduler shifts are read from the
7
+ checkpoint (video 10, audio 3 here; base H3 keeps 12/3), the eight-rung DMD
8
+ ladder is loaded from `fastvideo_inference.json`, and a new
9
+ `examples/inference/basic/basic_fasth3_8step.py` example pins the recipe.
10
+ Public `main` before that PR hard-codes video shift 12 and cannot run this
11
+ checkpoint unchanged.
 
 
 
 
 
12
 
13
  ## Exact sampling contract
14
 
 
25
  the same settings; verify that your launcher applies them rather than
26
  assuming its defaults match this file.
27
 
28
+ ## Command
29
 
30
+ On the PR #1852 branch (or `main` once merged), on a GPU worker:
31
 
32
  ```bash
33
+ python examples/inference/basic/basic_fasth3_8step.py \
 
 
 
 
34
  --prompt 'integrated_multimodal_description: A red fox runs through fresh snow at dawn. overall_soundscape: Fast pawsteps in snow, winter wind, and distant birds.' \
35
  --height 768 --width 1344 --num-frames 124 \
36
+ --num-gpus 4 --vsa-kernel sm100a \
37
+ --profile strict --no-inference-torch-compile --no-compile-vae \
 
 
38
  --output outputs/fasth3-v2 \
39
  --repeats 1
40
  ```
41
 
42
+ No environment variables are needed: the example reads the ladder and shifts
43
+ from the checkpoint and rejects any `--steps` other than 9.
44
+
45
  This example requests 1344×768, 124 frames at 24 fps (about 5.17 seconds),
46
+ and writes MP4s with a stereo 32 kHz audio track under `outputs/fasth3-v2/`. The `sm100a` route needs a compatible
47
  Blackwell GPU and kernel build. Triton is a different kernel option, not a
48
  license to change the trained sparsity, tile geometry, or shifts.
49
 
50
+ Verified on the PR #1852 head on 4x GB200: eight student forwards,
51
+ video/audio shifts 10/3, VSA 0.8 tile 64, and a playable MP4 with audio
52
+ (832x480, 124 frames). Report any latency with the exact hardware,
53
+ resolution, frame count, attention path and warmup policy.
README.md CHANGED
@@ -38,8 +38,6 @@ Video Sparse Attention (VSA).
38
 
39
  Powered by MiniMax H3.
40
 
41
- > Private release candidate. Public release is pending the compatibility and
42
- > review checks in [RELEASE_CHECKLIST.md](RELEASE_CHECKLIST.md).
43
 
44
  ## Model
45
 
@@ -77,17 +75,30 @@ UV_TORCH_BACKEND=cu130 uv pip install \
77
  -e ".[fasth3]"
78
  ```
79
 
80
- **Compatibility hold:** the public H3 pipeline currently enforces video shift
81
- 12 and its `basic_fasth3.py` example does not expose scheduler-shift overrides.
82
- It cannot yet run this shift-10 checkpoint unchanged. Installing FastVideo is
83
- not sufficient until that support lands; do not substitute shift 12 or the
84
- four-step defaults. See [INFERENCE.md](INFERENCE.md) for the exact contract
85
- and the private-review command for a compatible checkout.
86
 
87
- While the repo is private, reviewers need an account with access:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
88
 
89
  ```bash
90
- hf auth login
91
  hf download FastVideo/FastVideo-FastH3-8-Step-V2 \
92
  --local-dir ./FastH3-8-Step-V2
93
  ```
@@ -103,7 +114,10 @@ FastVideo with VSA-H3, not an unmodified dense Diffusers pipeline.
103
  - Difficult motion, small details and some audio can remain below Base H3.
104
  - Sampling outside the trained schedule or attention policy is an ablation.
105
  - The four-step blog's latency and evaluation numbers do not establish this
106
- eight-step model's performance. Release-specific measurements are pending.
 
 
 
107
 
108
  This model inherits the [MiniMax H3 Community License](LICENSE), including
109
  its territorial, use and commercial restrictions. Review it before use or
 
38
 
39
  Powered by MiniMax H3.
40
 
 
 
41
 
42
  ## Model
43
 
 
75
  -e ".[fasth3]"
76
  ```
77
 
78
+ Inference support for this checkpoint (video shift 10 and the trained
79
+ eight-rung ladder) is in FastVideo
80
+ [PR #1852](https://github.com/hao-ai-lab/FastVideo/pull/1852). Until it is
81
+ merged, check out that branch; afterwards, `main` works unchanged.
 
 
82
 
83
+ ```bash
84
+ python examples/inference/basic/basic_fasth3_8step.py \
85
+ --prompt 'A slow cinematic drone shot glides over a coastal town at golden hour; gulls call over the harbor as a church bell rings twice.' \
86
+ --num-gpus 4 --vsa-kernel sm100a \
87
+ --profile strict --no-inference-torch-compile --no-compile-vae \
88
+ --height 768 --width 1344 --num-frames 124 \
89
+ --output outputs/fasth3-8step
90
+ ```
91
+
92
+ `basic_fasth3_8step.py` pins this checkpoint and its trained recipe (nine
93
+ sigma-grid points = eight forwards, VSA 0.8, 64-token tiles) and reads the
94
+ shifts and rung ladder from `scheduler/`, `audio_scheduler/` and
95
+ [fastvideo_inference.json](fastvideo_inference.json). It shares the FastH3
96
+ preview example's CLI, so every other flag works unchanged; pass
97
+ `--model-path` to use a local snapshot.
98
+
99
+ To download ahead of time:
100
 
101
  ```bash
 
102
  hf download FastVideo/FastVideo-FastH3-8-Step-V2 \
103
  --local-dir ./FastH3-8-Step-V2
104
  ```
 
114
  - Difficult motion, small details and some audio can remain below Base H3.
115
  - Sampling outside the trained schedule or attention policy is an ablation.
116
  - The four-step blog's latency and evaluation numbers do not establish this
117
+ eight-step model's performance. On 4x GB200 (SP4, eager strict profile,
118
+ VSA 0.8 tile 64, 832x480, 124 frames + 32 kHz audio) one measured request
119
+ took 6.3 s end to end including saving (FastVideo PR #1852 smoke). Quality
120
+ has not been formally compared against Base H3 or the four-step preview.
121
 
122
  This model inherits the [MiniMax H3 Community License](LICENSE), including
123
  its territorial, use and commercial restrictions. Review it before use or
RELEASE_CHECKLIST.md CHANGED
@@ -1,6 +1,6 @@
1
  # FastH3 8-Step V2 release checklist
2
 
3
- Prepared 2026-09-08. Keep the repository private until the author approves release.
4
 
5
  ## Verified packaging
6
 
@@ -24,13 +24,15 @@ checkpoint weights, scheduler settings or original provenance hashes.
24
 
25
  ## Required before publishing
26
 
27
- - [ ] Land or provide a reviewed public FastVideo inference path for video shift
28
- 10 and the exact eight-rung ladder. Public `main` at
29
  `a943220c115228ade5d57b3bab9a6a87fd600a10` still hard-codes video shift 12;
30
  the old README's scheduler-shift CLI options are absent there.
31
- - [ ] Replace the compatibility hold with a verified public quickstart.
32
- - [ ] Run one end-to-end GPU generation using the final public installation;
33
- confirm eight forwards, shifts 10/3, VSA80/tile64 and playable video with audio.
 
 
34
  - [ ] Inspect representative high-motion samples. Do not copy the four-step
35
  models' evaluation results or speed claims onto this checkpoint.
36
  - [ ] Complete author/legal review of the MiniMax H3 territory, redistribution
@@ -40,8 +42,7 @@ checkpoint weights, scheduler settings or original provenance hashes.
40
  provenance files before publication. No credentials were found in the
41
  reviewed metadata, but provenance sanitization would require updating the
42
  corresponding integrity references deliberately.
43
- - [ ] Obtain explicit approval to change visibility. No public collection edit,
44
- announcement, or visibility change is part of this preparation.
45
 
46
  Qwen license source:
47
  `https://github.com/QwenLM/Qwen3-VL/blob/96588727e44c78b25ba03ea03b8e12f7e64fd0da/LICENSE`.
 
1
  # FastH3 8-Step V2 release checklist
2
 
3
+ Prepared 2026-09-08; updated 2026-09-15 after the repository was made public and FastVideo PR #1852 was opened.
4
 
5
  ## Verified packaging
6
 
 
24
 
25
  ## Required before publishing
26
 
27
+ - [x] Land or provide a reviewed public FastVideo inference path for video shift
28
+ 10 and the exact eight-rung ladder. Open as FastVideo PR #1852 (not yet merged). Public `main` at
29
  `a943220c115228ade5d57b3bab9a6a87fd600a10` still hard-codes video shift 12;
30
  the old README's scheduler-shift CLI options are absent there.
31
+ - [x] Replace the compatibility hold with a verified public quickstart
32
+ (`basic_fasth3_8step.py`, PR #1852).
33
+ - [x] Run one end-to-end GPU generation using the PR #1852 head; confirmed
34
+ eight forwards, shifts 10/3, VSA80/tile64 and playable video with audio
35
+ (4x GB200, 832x480x124f). Re-verify on `main` after merge.
36
  - [ ] Inspect representative high-motion samples. Do not copy the four-step
37
  models' evaluation results or speed claims onto this checkpoint.
38
  - [ ] Complete author/legal review of the MiniMax H3 territory, redistribution
 
42
  provenance files before publication. No credentials were found in the
43
  reviewed metadata, but provenance sanitization would require updating the
44
  corresponding integrity references deliberately.
45
+ - [x] Repository made public by the author on 2026-09-15.
 
46
 
47
  Qwen license source:
48
  `https://github.com/QwenLM/Qwen3-VL/blob/96588727e44c78b25ba03ea03b8e12f7e64fd0da/LICENSE`.