Instructions to use Comfy-Org/MiniMax-H3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusion Single File
How to use Comfy-Org/MiniMax-H3 with Diffusion Single File:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Why MiniMax H3 Ruins Faces on Wide Shots?
π Quick Tip (Low VRAM / RTX 3060 12GB): Want fast MiniMax H3 iterations? Start low-res, then upscale with LTX or any other latent upscaler.
Note: MiniMax heavily distorts faces on wide shots. Distortions happen regardless of input res (even at 720p, very bad). Close/medium shots look fine! Cheers!
As you said it's a wide shot, your character face takes much less space on screen than if It was a vertical aspect ratio for the same total pixel count, the more smaller the head will be on screen, the worse it will look.
I don't recommend the tip you gave, if you want better looking video, you can start low res to see if your prompt is on par with what you wan't, if it is, just restart the same gen at higher resolution.
Using upscalers that will re-encode to latents and decode again will destroy initial details, add artifacts, give worse temporal motion
Using 'EasyCache' node also degrades the quality of the final result.
UntMods ... Maybe you didn't read the post closely enough. I'm talking about fast iterations there. For instance, when you need to quickly put together a demo for a presentation. The issue that some other users might have noticed too is that render times grow exponentially as the resolution, video length, and the number of steps increase. Compared to LTX, iterating is quite a pain if you're stuck with limited, low-VRAM hardware.
So, this tip isn't for you, but it might save time for those who don't own an RTX 5090. And from what I've gathered, even people with an RTX 5090 aren't exactly thrilled with the render times for Minimax H3. Time will tell. It's open-source, and everyone can choose the workflow that suits them bestβwhether they use Wan, LTX, Minimax H3, or anything else.
Feel free to share your own samples and let us know whatβs working for you and what isnβt, so others can learn from it too!
Cheers!
That's generally the case with all generative AI models, including image generation.
In VFX, there are a couple of common ways to deal with it:
- Increase the base resolution before generation, if your hardware can handle it. (Unfortunately, this isn't an option for everyone due to VRAM and compute limitations.)
- Repair only the affected region. In your NLE, crop the area where the quality has degraded, render that crop, then run it through your preferred workflow again using a ControlNet (such as Depth) with fixed start/end frames to preserve temporal consistency. Afterward, bring the result back into your NLE and blend it using feathered masks or rotoscoping.
If the degradation is relatively minor (not as severe as in your example), you can often get away with a light 0.15 denoise V2V pass instead of the more involved ControlNet + fixed start/end frame workflow.
Hope that helps!
Mike, I think the confusion is you lead with a rhetorical question to the community. You already knew the answer and were offering a solution.
Minimax H3 is "Perfect"... If You Don't Care About Audio, Faces, or Reality.
For all the hype-train riders and influencer fanboys who think Minimax H3 is perfect: Iβm glad youβre enjoying your new toy. But before you start typing in the comments that Iβm doing it wrong, save your energy. Send me a video sample that proves your point. Until then, enjoy the face glitches and audio hallucinations while the rest of us actually get work done on LTX.
I am currently finalizing the materials for an extensive review of Minimax H3. I rendered some samples at a higher resolution and 24 steps, yet I am seeing the exact same issue. Recent tests show that resolution doesn't have as much of an impact on producing face glitches and errors as one might think. I have extensive, long-term experience in the VFX industry. So, before anyone tries to lecture me on what I'm doing wrong, save your time typing and send us a video sample instead. Based on that, we can immediately talk about pretty much anything regarding VFX, ComfyUI, post-production, etc.
What I'm writing about here likely relates to how Minimax H3 was trained, what datasets were used, and what post-training restrictions (safety guardrails) were subsequently applied. Everyone here probably remembers how it was in the beginning with Seedance, the subsequent limitations, and so on. Furthermore, the licensing terms around Minimax seem to be quite blurry. I noticed reports on Reddit that some LoRA creators for Minimax H3 were forced to delete their models. These are unverified rumors, of course. However, some users have noted that Minimax H3 barely restricts obscure and explicit content. The "corn" (NSFW) community is rejoicing; they have a new toy.
Putting the visual aspect aside, there is another major component: multimodal utility. And during my testing, this went horribly wrong. The verdict isn't final yet, but the signs are already clear. Minimax is simply not what some people are trying to hype it up to be while trashing LTX. The LTX 2.3 model remains highly usable, especially for those with limited hardware and low VRAM. For my workflow, LTX delivers everything I need, and my RTX 3060 12GB is more than enough for now. So, what else is there regarding Minimax H3? What's even worse, it completely ignores the voice-over prompt, which works flawlessly with LTX. Minimax H3 just generates random, barely recognizable words, as if it's suffering from complete hallucination. And it gets even worse with custom audio it glitches out mid-generation and produces a completely messed-up audio remix. As far as multimodality goes, Minimax H3 is absolute garbage and completely unusable for me right now. I'll keep testing and posting the results on my YT channel, but so far, compared to LTX, Minimax H3 is a total flop.
You came here to ask something, or at least it seemed that way. Then started yapping for no reason. You need to rewrite your soul.md. It's just tokens waste
Hey Matt, here is a quick test update.
RTX 3060 12GB | T2V | 1376x768 | 24 steps | 3 sec (20 min was crazy!)
Still the same issue: distorted faces and heavy artifacts.
Cheers!
Does anybody else have the same experience, or know how to fix this?
I have the same issue, I think we have to wait somebody making the maigc things happen.
yes MMH3has some great "features" , but the output quality is not great, takes way too long for just 5-10 secs of video, but even if some fixes the speed to be more inline with LTX, the quality just isnt going to get a magic fix. The people in the thread are right, I just rendered a LTX video where the woman is roughly the same distance away from the camera as the video above and the difference in facial clarity is night and day.
Any video I do with MMH3 takes wayyy too long and the faces turn to garbage if it isnt a close up, which you 'dont' get with LTX .
The text encoder is incredible... The video model is a janky mess. The way the it can take many refs and make a video from them is really fun... but...
Its not able to replicate what LTX can do. It's a budget seedance, not a replacement LTX. Its also a side step from wan 2.1 vace that can do longer segments and not drift. But it cant inpaint or upscale or extend properly and still cant do fast motions like vace could. Skipping still fails my smell tests where vace can easily do it.
Its a good model, but its not replacing LTX yet for me and licence is terribad... the audio is a mess, and faces mush but its fun i guess.
Use LTX if you
want 30+ second videos in minutes.
Want proper sounds with lipsyncs
want to copy simple dances and memes.
Use minimax if you
Want to ref many images and combine them to a story
Want your prompt to do the heavy lifting.
Use wan 2.1 vace when you
want to do complex fast motions with a video copy
Minimax H3 video vae int8 convrot and 8-step Turbo LoRA
More testing with Minimax H3. This time, I'm using the fresh, new experimental VAE from Kijai: minimax_h3_video_vae_int8_convrot.safetensors.
I also used the Turbo custom node by larryvrh ComfyUI-MiniMax-H3-Turbo. I did a quick test with the 8-step Turbo LoRA and the new VAE, and the results are better in terms of speed with slightly improved quality.
Everything was tested on an RTX 3060 12GB. Rendering a 5-second video at 864x480 resolution with Turbo at 8 steps took 4.5 minutes.
It's not as fast as LTX, but it's doable. This is just the first test, and I need more time to benchmark everything properly. This includes testing the audio and pushing Minimax H3 with the custom prompts I use as benchmarks for every AI video model.
Cheers!
I know this isn't a wide shot - this should be a proof of concept - a portrait with minimaly! 0.6Mpix CAN do a distant face, but at the greater distance there is a need for even higher res...
4060ti 16gb, 64gb ddr4 ram, OG VAE, no turbo, Spectrum Apply MiniMax H3, 62min @ ~ 150s/it
result : 0.6Mpix 3:4 portrait ratio, 22sec
inputs : 4xImage (around 2K without resizing), einstein audio (cca 2min)
metadata saved in video - workflow
