matlod's picture
Add recovered fight source and reproducible four-tile canvas examples
42cbd13 verified
|
Raw History Blame Contribute Delete
6.51 kB
# Reproduce the historical fight quad
The top-left tile is the supplied source video, not another processing pass.
Copy `seedhunt_20260963_00001_.mp4` into your ComfyUI `input` directory. Keep its
filename unchanged. It is the original 124-frame H3 generation (1152 × 640,
24 fps, 5.167 s), recovered from the local experiment outputs.
## Open it on the ComfyUI canvas
- **[Winner only — easy starting point](adapter.canvas.json)**: the verified
adapter recipe with setup notes and editable controls.
- **[All four tiles — complete comparison](fight_quad.canvas.json)**: shared
source/model loaders, color-coded processing lanes, four individual video
outputs, and an automatically assembled comparison with source audio.
Download the JSON, then drag it onto ComfyUI or use **Workflow → Open**. Put the
source MP4 in `ComfyUI/input`, select your model files, and click **Run**. The
amber lane is the full-clip pass, rose is the windowed control, and green is the
adapter winner. Advanced nodes are collapsed; double-click a title to expand.
The full comparison takes longer and uses more memory than the winner alone.
The normal Save Video nodes embed both the prompt and frontend workflow when
run through the canvas with metadata enabled.
## Requirements
Use a ComfyUI build with MiniMax-H3 support, [ComfyUI-MAINodes](https://github.com/matlowai/ComfyUI-MAINodes),
and ComfyUI-KJNodes with SageAttention available. Python 3 and FFmpeg are needed
for the commands below. These API JSONs can also be imported through a frontend
that supports API-format import; they are not canvas-layout JSONs.
Model names in the supplied graphs:
- diffusion model: `minimax_h3/minimax_h3_fl2va_pruned_int8_convrot.safetensors`
- text encoder: `minimax_h3/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors`
- VAE: `minimax_h3/minimax_h3_video_vae_fp16.safetensors`
- LoRA: `minimax_h3/minimax_h3_motion_adapter_pilot_r16.safetensors`
Place this repository's adapter in `models/loras/minimax_h3/`. The historical
training filename `p4_pilot_k100.safetensors` refers to this published adapter.
Use the exact base/encoder/VAE variants for historical comparison; adjust only
folder paths if your installation uses different folders.
## Render the three processed tiles
Download this directory, then run these commands inside it with ComfyUI running.
Use `--server http://127.0.0.1:8189` on each command if that is your server port.
The runner uses the public `/prompt`, `/history`, and `/view` endpoints, assigns
unique output prefixes, waits for success, and downloads the resulting video.
```bash
python3 run.py full_clip.api.json --output full_clip.mp4
python3 run.py windowed.api.json --output windowed.mp4
python3 run.py adapter.api.json --output adapter.mp4
```
| Tile | Input / workflow | Settings |
|---|---|---|
| Top left | `seedhunt_20260963_00001_.mp4` | Original source, unchanged |
| Top right | `full_clip.api.json` | T2C arm A, full-clip pixel smear and regeneration, inject 0.60, no adapter |
| Bottom left | `windowed.api.json` | Windowed v3.1, inject 0.45, no adapter |
| Bottom right | `adapter.api.json` | Same window, adapter 0.75, inject 0.30 |
All processing arms use seed 20260817, beta schedule, 25 total steps, and
`gradient_estimation`. The windowed arms crop world frames 68–123, expand the
56-frame window to 107 frames, retain the historical tail guide, recover to
56 frames, and splice after the untouched 68-frame head. These are the old
recipes, not the later two-boundary pinned graph.
`expand_to_end=false` is explicit so current node defaults cannot rewrite the
historical hold maps. The windowed map already expands through its final frame.
Descriptive node titles from the archive were removed because some still said
inject 0.45 even when the actual input was 0.30.
To queue the combined graph and download its assembled comparison directly:
```bash
python3 run.py fight_quad.api.json --output-node 9005 --output fight_quad.mp4
```
## Assemble the four tiles
This rebuilds the panel arrangement with the source audio. It deliberately has
no historical timing overlays: your new runs are not the old benchmark.
```bash
ffmpeg -n -i seedhunt_20260963_00001_.mp4 -i full_clip.mp4 -i windowed.mp4 -i adapter.mp4 \
-filter_complex '[0:v][1:v]hstack=inputs=2[top];[2:v][3:v]hstack=inputs=2[bottom];[top][bottom]vstack=inputs=2[v]' \
-map '[v]' -map '0:a?' -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac -movflags +faststart fight_quad.mp4
```
## Metadata, provenance, and verification
`winner_verified_20260921.mp4` is the September 21 functional rerun of the
adapter arm: success, 65.385 s server execution, all 124 frames decoded without
errors. Only the text-encoder and VAE loader nodes were reported cached; the
sampler executed. This does not replace the historical matched-cache 49.9 s
measurement. The full-clip and no-adapter workflows are recovered historical graphs.
See `verification.json` for the separate full-canvas verification run.
The verification used MAINodes checkout commit
`4e957e267795c23c4d5544cf7d62bb4fa736c4c3`. Hardware, software, kernels, and cache
state affect timing and pixels; identical results across environments are not
guaranteed. Historical benchmark labels remain in the original
[comparison video](../../assets/adapter_t2c_insert_quad.mp4).
Both supplied MP4s embed API JSON under the container's `prompt` tag. To inspect:
```bash
ffprobe -v error -show_entries format_tags=prompt -of json winner_verified_20260921.mp4
```
The embedded winner graph preserves the exact local run, including the original
LoRA training filename. The separate `adapter.api.json` uses the published
filename and a fresh output prefix. A `workflow` canvas-layout tag is not present.
The assembled FFmpeg quad does not automatically inherit all four graphs.
`source_generation.api.json` is extracted from the original source MP4 for
provenance and optional regeneration. It additionally requires the LightX2V
LoRA named in that graph. Reusing the supplied source is the way to reproduce
this comparison without changing the plate. Base-model licensing remains as
specified in the parent model card.
The complete browser-exported comparison graph also passed a fresh render: all
three samplers executed, all five saved videos decode, and the assembled output
is 2304 × 1280 at 24 fps with 124 frames. Total server execution was 279.09 s
for this combined run. [Watch the newly assembled comparison](comparison_verified_20260921.mp4).