Image-Text-to-Video
Diffusers
Safetensors
MiniMaxH3ModularPipeline
text-to-video
image-to-video
video-to-video
text-to-audio-video
image-to-audio-video
image-text-to-audio-video
video-to-audio-video
audio-to-audio-video
audio-video-generation
multimodal
synchronized-audio-video
reference-to-audio-video
Instructions to use johnsonhuggingapi/MiniMax-H3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use johnsonhuggingapi/MiniMax-H3 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("johnsonhuggingapi/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Download Ref2VA/audio_vae/dac_utils.py from johnsonhuggingapi/MiniMax-H3: direct link, hf CLI and curl.
- Browser
- Download file 362 Bytes
-
https://huggingface.co/johnsonhuggingapi/MiniMax-H3/resolve/main/Ref2VA/audio_vae/dac_utils.py
- Command line
-
hf download hf://johnsonhuggingapi/MiniMax-H3/Ref2VA/audio_vae/dac_utils.py
-
curl -L -o dac_utils.py https://huggingface.co/johnsonhuggingapi/MiniMax-H3/resolve/main/Ref2VA/audio_vae/dac_utils.py
362 Bytes
| # SPDX-License-Identifier: MIT | |
| # Adapted from https://github.com/jik876/hifi-gan under the MIT license. | |
| def init_weights(m, mean=0.0, std=0.01): | |
| classname = m.__class__.__name__ | |
| if classname.find("Conv") != -1: | |
| m.weight.data.normal_(mean, std) | |
| def get_padding(kernel_size, dilation=1): | |
| return int((kernel_size * dilation - dilation) / 2) | |