Instructions to use MATLOWAI/MiniMax-H3-Motion-Adapter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use MATLOWAI/MiniMax-H3-Motion-Adapter with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("MATLOWAI/MiniMax-H3-Motion-Adapter") prompt = "A man with short gray hair plays a red electric guitar." input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png") output = pipe(image=input_image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
|
Download examples/fight_quad/README.md from MATLOWAI/MiniMax-H3-Motion-Adapter: direct link, hf CLI and curl.
- Browser
- Download file 6.51 kB
-
https://huggingface.co/MATLOWAI/MiniMax-H3-Motion-Adapter/resolve/main/examples/fight_quad/README.md
- Command line
-
hf download hf://MATLOWAI/MiniMax-H3-Motion-Adapter/examples/fight_quad/README.md
-
curl -L -o README.md https://huggingface.co/MATLOWAI/MiniMax-H3-Motion-Adapter/resolve/main/examples/fight_quad/README.md
6.51 kB
| # Reproduce the historical fight quad | |
| The top-left tile is the supplied source video, not another processing pass. | |
| Copy `seedhunt_20260963_00001_.mp4` into your ComfyUI `input` directory. Keep its | |
| filename unchanged. It is the original 124-frame H3 generation (1152 × 640, | |
| 24 fps, 5.167 s), recovered from the local experiment outputs. | |
| ## Open it on the ComfyUI canvas | |
| - **[Winner only — easy starting point](adapter.canvas.json)**: the verified | |
| adapter recipe with setup notes and editable controls. | |
| - **[All four tiles — complete comparison](fight_quad.canvas.json)**: shared | |
| source/model loaders, color-coded processing lanes, four individual video | |
| outputs, and an automatically assembled comparison with source audio. | |
| Download the JSON, then drag it onto ComfyUI or use **Workflow → Open**. Put the | |
| source MP4 in `ComfyUI/input`, select your model files, and click **Run**. The | |
| amber lane is the full-clip pass, rose is the windowed control, and green is the | |
| adapter winner. Advanced nodes are collapsed; double-click a title to expand. | |
| The full comparison takes longer and uses more memory than the winner alone. | |
| The normal Save Video nodes embed both the prompt and frontend workflow when | |
| run through the canvas with metadata enabled. | |
| ## Requirements | |
| Use a ComfyUI build with MiniMax-H3 support, [ComfyUI-MAINodes](https://github.com/matlowai/ComfyUI-MAINodes), | |
| and ComfyUI-KJNodes with SageAttention available. Python 3 and FFmpeg are needed | |
| for the commands below. These API JSONs can also be imported through a frontend | |
| that supports API-format import; they are not canvas-layout JSONs. | |
| Model names in the supplied graphs: | |
| - diffusion model: `minimax_h3/minimax_h3_fl2va_pruned_int8_convrot.safetensors` | |
| - text encoder: `minimax_h3/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | |
| - VAE: `minimax_h3/minimax_h3_video_vae_fp16.safetensors` | |
| - LoRA: `minimax_h3/minimax_h3_motion_adapter_pilot_r16.safetensors` | |
| Place this repository's adapter in `models/loras/minimax_h3/`. The historical | |
| training filename `p4_pilot_k100.safetensors` refers to this published adapter. | |
| Use the exact base/encoder/VAE variants for historical comparison; adjust only | |
| folder paths if your installation uses different folders. | |
| ## Render the three processed tiles | |
| Download this directory, then run these commands inside it with ComfyUI running. | |
| Use `--server http://127.0.0.1:8189` on each command if that is your server port. | |
| The runner uses the public `/prompt`, `/history`, and `/view` endpoints, assigns | |
| unique output prefixes, waits for success, and downloads the resulting video. | |
| ```bash | |
| python3 run.py full_clip.api.json --output full_clip.mp4 | |
| python3 run.py windowed.api.json --output windowed.mp4 | |
| python3 run.py adapter.api.json --output adapter.mp4 | |
| ``` | |
| | Tile | Input / workflow | Settings | | |
| |---|---|---| | |
| | Top left | `seedhunt_20260963_00001_.mp4` | Original source, unchanged | | |
| | Top right | `full_clip.api.json` | T2C arm A, full-clip pixel smear and regeneration, inject 0.60, no adapter | | |
| | Bottom left | `windowed.api.json` | Windowed v3.1, inject 0.45, no adapter | | |
| | Bottom right | `adapter.api.json` | Same window, adapter 0.75, inject 0.30 | | |
| All processing arms use seed 20260817, beta schedule, 25 total steps, and | |
| `gradient_estimation`. The windowed arms crop world frames 68–123, expand the | |
| 56-frame window to 107 frames, retain the historical tail guide, recover to | |
| 56 frames, and splice after the untouched 68-frame head. These are the old | |
| recipes, not the later two-boundary pinned graph. | |
| `expand_to_end=false` is explicit so current node defaults cannot rewrite the | |
| historical hold maps. The windowed map already expands through its final frame. | |
| Descriptive node titles from the archive were removed because some still said | |
| inject 0.45 even when the actual input was 0.30. | |
| To queue the combined graph and download its assembled comparison directly: | |
| ```bash | |
| python3 run.py fight_quad.api.json --output-node 9005 --output fight_quad.mp4 | |
| ``` | |
| ## Assemble the four tiles | |
| This rebuilds the panel arrangement with the source audio. It deliberately has | |
| no historical timing overlays: your new runs are not the old benchmark. | |
| ```bash | |
| ffmpeg -n -i seedhunt_20260963_00001_.mp4 -i full_clip.mp4 -i windowed.mp4 -i adapter.mp4 \ | |
| -filter_complex '[0:v][1:v]hstack=inputs=2[top];[2:v][3:v]hstack=inputs=2[bottom];[top][bottom]vstack=inputs=2[v]' \ | |
| -map '[v]' -map '0:a?' -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac -movflags +faststart fight_quad.mp4 | |
| ``` | |
| ## Metadata, provenance, and verification | |
| `winner_verified_20260921.mp4` is the September 21 functional rerun of the | |
| adapter arm: success, 65.385 s server execution, all 124 frames decoded without | |
| errors. Only the text-encoder and VAE loader nodes were reported cached; the | |
| sampler executed. This does not replace the historical matched-cache 49.9 s | |
| measurement. The full-clip and no-adapter workflows are recovered historical graphs. | |
| See `verification.json` for the separate full-canvas verification run. | |
| The verification used MAINodes checkout commit | |
| `4e957e267795c23c4d5544cf7d62bb4fa736c4c3`. Hardware, software, kernels, and cache | |
| state affect timing and pixels; identical results across environments are not | |
| guaranteed. Historical benchmark labels remain in the original | |
| [comparison video](../../assets/adapter_t2c_insert_quad.mp4). | |
| Both supplied MP4s embed API JSON under the container's `prompt` tag. To inspect: | |
| ```bash | |
| ffprobe -v error -show_entries format_tags=prompt -of json winner_verified_20260921.mp4 | |
| ``` | |
| The embedded winner graph preserves the exact local run, including the original | |
| LoRA training filename. The separate `adapter.api.json` uses the published | |
| filename and a fresh output prefix. A `workflow` canvas-layout tag is not present. | |
| The assembled FFmpeg quad does not automatically inherit all four graphs. | |
| `source_generation.api.json` is extracted from the original source MP4 for | |
| provenance and optional regeneration. It additionally requires the LightX2V | |
| LoRA named in that graph. Reusing the supplied source is the way to reproduce | |
| this comparison without changing the plate. Base-model licensing remains as | |
| specified in the parent model card. | |
| The complete browser-exported comparison graph also passed a fresh render: all | |
| three samplers executed, all five saved videos decode, and the assembled output | |
| is 2304 × 1280 at 24 fps with 124 frames. Total server execution was 279.09 s | |
| for this combined run. [Watch the newly assembled comparison](comparison_verified_20260921.mp4). | |