Upload 10Eros_Max_h3_TURBO-hybrid_beta5_poc_simple_edges_w4g16.safetensors
Summary
This PR adds a provisional quantized build of the MiniMax H3 10Eros_Max_h3_TURBO-hybrid_beta5 transformer for testing on GPUs with approximately 16 GB of VRAM.
This is a moderate-quality proof-of-concept rather than the definitive beta5 quantization. The source checkpoint is itself provisional and is expected to be replaced due to a known upstream issue.
Quantization layout
The checkpoint uses a conservative mixed-precision arrangement focused on the 200 large transformer matrices, which account for approximately 95.8% of the original file:
- Edge blocks 0, 1, 47, 48, and 49 use row-wise INT8 ConvRot.
- Interior blocks 2–46 use native W4A8 ConvRot with group size 16.
- All remaining tensors are preserved byte-for-byte in source BF16.
The targeted matrix families are:
attn.qkv_proj.weight
attn.out_proj.weight
mlp.fc1.weight
mlp.fc2.weight
Small and potentially sensitive infrastructure tensors—including normalization weights, biases, AdaLN components, patch and condition projections, time/rope tables, final projections, and token-refiner tensors—remain in BF16.
The stored formats use ComfyUI/comfy-kitchen-compatible quantization metadata:
- int8_tensorwise with row-wise scales and ConvRot group 256 for edge blocks 0, 1, 47, 48, and 49 heavy weight tensors.
- asym_w4a8_int8 with group size 16, FP8 E4M3 relative scales, and ConvRot group 256 for interior blocks 2–46 heavy weight tensors.
Intended use
The resulting transformer is approximately 12.46 GiB on disk and is intended to provide a practical starting point for 16 GB target GPUs. Actual VRAM usage depends on the workflow, resolution, frame count, text encoder, VAE, and other resident models.
Quality and validation status
This build has not undergone the full benchmark, learned-rounding, precision-allocation, rescue, or controlled visual-comparison process used for the previous beta4 release.
It should therefore be treated as:
- provisional and experimental;
- moderate-quality rather than quality-maximized;
- suitable for compatibility and workflow testing;
- not representative of the eventual definitive beta5 quantization.
The source topology was verified against the expected 50-block H3 layout before conversion. No broad perceptual-quality claim is made for this artifact.
Closing without merging as this provisional quantization was based on the previous retired Beta5