Image-Text-to-Video
Diffusers
Safetensors
MiniMaxH3ModularPipeline
text-to-video
image-to-video
video-to-video
text-to-audio-video
image-to-audio-video
image-text-to-audio-video
video-to-audio-video
audio-to-audio-video
audio-video-generation
multimodal
synchronized-audio-video
reference-to-audio-video
Instructions to use johnsonhuggingapi/MiniMax-H3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use johnsonhuggingapi/MiniMax-H3 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("johnsonhuggingapi/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Download Ref2VA/audio_vae/metadata.json from johnsonhuggingapi/MiniMax-H3: direct link, hf CLI and curl.
- Browser
- Download file 440 Bytes
-
https://huggingface.co/johnsonhuggingapi/MiniMax-H3/resolve/main/Ref2VA/audio_vae/metadata.json
- Command line
-
hf download hf://johnsonhuggingapi/MiniMax-H3/Ref2VA/audio_vae/metadata.json
-
curl -L -o metadata.json https://huggingface.co/johnsonhuggingapi/MiniMax-H3/resolve/main/Ref2VA/audio_vae/metadata.json
440 Bytes
| { | |
| "metadata": { | |
| "kwargs": { | |
| "attn_proj": true, | |
| "decoder_dim": 1024, | |
| "decoder_rates": [ | |
| 5, | |
| 5, | |
| 2, | |
| 2, | |
| 2, | |
| 2, | |
| 2 | |
| ], | |
| "decoder_type": "bigvgan", | |
| "encoder_dim": 64, | |
| "encoder_rates": [ | |
| 2, | |
| 4, | |
| 4, | |
| 5, | |
| 5 | |
| ], | |
| "latent_dim": 2048, | |
| "sample_rate": 32000, | |
| "vae_latent_channels": 32 | |
| } | |
| } | |
| } | |