Mage-Flow-NVFP4-Quality-AJH

Mage-Flow-NVFP4-Quality-AJH is a portable, runnable quantized version of microsoft/Mage-Flow. Standalone Hugging Face Mage-Flow repository with all 24 image MLP projections kept in BF16 and the 24 text MLP projections stored as native NVFP4.

This repository is intended to remain discoverable as a quantized microsoft/Mage-Flow derivative while avoiding any runtime dependency on a separate local models/ checkout. The complete transformer component, quantized text encoder, VAE, scheduler, vendored inference code, and native runtime are all packaged inside the repository layout.

Related releases:

Showcase

Cyborg woman gazing toward a star-filled sky

Generated directly with this quality package, without upscaling or post-processing.

A 4K resolution high detail photo realistic image of the top half of a cyborg woman with dark black hair, striking blue eyes that have a very subtle glow in the iris, standing side profile, head tilted up towards the sky with a questioning expression, she has subtle gaps in her skin that hint at a robotic nature, outdoor forest night setting, sky filled with bright brilliant stars that glow against the dark setting, nebula visible

Settings: 1280×1280, 20 steps, CFG 5, static shift 6, seed 3334072683.

Matched BF16 comparison

This comparison uses the same frozen prompt, seed 1795681438, 1024×1024 resolution, 20 steps, CFG 5, and static shift 6. It compares the complete BF16 pipeline against the complete Quality package, including its quantized text encoder.

Full BF16 reference Quality: 24 NVFP4 + 24 BF16 image projections
Full BF16 reference Quality NVFP4 package
21.79 s denoise 18.48 s denoise

Transformer policy

  • Native NVFP4 transformer projections: 24
  • BF16 passthrough transformer projections: 24
  • Default native runtime activation search: amax
  • Default up activation multiplier: 1
  • Default down activation multiplier: 1

BF16 passthrough modules:

  • transformer_blocks.0.img_mlp.net.0.proj
  • transformer_blocks.0.img_mlp.net.2
  • transformer_blocks.1.img_mlp.net.0.proj
  • transformer_blocks.1.img_mlp.net.2
  • transformer_blocks.2.img_mlp.net.0.proj
  • transformer_blocks.2.img_mlp.net.2
  • transformer_blocks.3.img_mlp.net.0.proj
  • transformer_blocks.3.img_mlp.net.2
  • transformer_blocks.4.img_mlp.net.0.proj
  • transformer_blocks.4.img_mlp.net.2
  • transformer_blocks.5.img_mlp.net.0.proj
  • transformer_blocks.5.img_mlp.net.2
  • transformer_blocks.6.img_mlp.net.0.proj
  • transformer_blocks.6.img_mlp.net.2
  • transformer_blocks.7.img_mlp.net.0.proj
  • transformer_blocks.7.img_mlp.net.2
  • transformer_blocks.8.img_mlp.net.0.proj
  • transformer_blocks.8.img_mlp.net.2
  • transformer_blocks.9.img_mlp.net.0.proj
  • transformer_blocks.9.img_mlp.net.2
  • transformer_blocks.10.img_mlp.net.0.proj
  • transformer_blocks.10.img_mlp.net.2
  • transformer_blocks.11.img_mlp.net.0.proj
  • transformer_blocks.11.img_mlp.net.2

Notes

  • Keeps every image-side MLP projection in BF16.
  • Quantizes only the text-side MLP projections.
  • Intended as the most conservative standalone mixed-transformer package candidate.

The Qwen3-VL text encoder uses the same mixed policy as the Fast release: 224 NVFP4 projections in blocks 2–33, 14 FP8 projections in blocks 1 and 34, with blocks 0 and 35 plus embeddings, norms, biases, and the vision tower retained in BF16.

Requirements and generation

The tested stack is Linux x86-64, NVIDIA Blackwell SM120, CUDA 13.1, Python 3.11, PyTorch 2.13.0+cu130, comfy-kitchen==0.2.22, and flash-attn==2.8.3. Install into a virtual environment using the included requirements.txt; do not install these packages system-wide.

python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
CUDA_HOME=/usr/local/cuda-13.1 python -m pip install --no-build-isolation flash-attn==2.8.3

CUDA_VISIBLE_DEVICES=0 .venv/bin/python generate.py \
  --model ajh-code/Mage-Flow-NVFP4-Quality-AJH \
  --prompt 'A detailed watercolor fox reading under an old oak tree' \
  --output fox.png --height 1024 --width 1024 --steps 20 --seed 1

The included binaries target the tested stack. Run build_native.sh after changing PyTorch, CUDA, or the C++ ABI.

Measured tradeoff

The controlled matched benchmark measured the Balanced transformer at about 1.43x BF16 throughput (30% less generation time), and the Quality transformer at about 1.22x BF16 throughput (18% less generation time). Projected transformer plus text-checkpoint storage is about 9.89 GB for Balanced and 10.97 GB for Quality, versus 17.12 GB for BF16. These are policy-level measurements from the research suite, not universal hardware guarantees.

Runtime caveats

  • Requires an NVIDIA Blackwell SM120 GPU.
  • Uses the packaged loader and native runtime; stock Diffusers does not natively understand mage_flow_nvfp4_* transformer modules.
  • The package-level runtime defaults are applied only when the caller has not already set the corresponding MAGE_NVFP4_* environment variables.
  • Generation and text-to-image are tested; Base, Turbo, editing, CUDA graph compatibility, and non-SM120 GPUs are not claimed.

Validate the package

python validate_release.py

License and attribution

Mage-Flow and the vendored Mage inference source are Copyright (c) 2026 Microsoft and MIT licensed. Qwen3-VL and the mixed NVFP4/FP8 text checkpoint are Apache-2.0 licensed. See the included license files and third-party notices for details.

Downloads last month
14
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ajh-code/Mage-Flow-NVFP4-Quality-AJH

Quantized
(4)
this model