Text-to-Image
Diffusers
Safetensors
English
Chinese
QwenImage21Pipeline
bitsandbytes
int8
image-generation
image-editing
rgba
8-bit precision
Instructions to use ixim/Image21-INT8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ixim/Image21-INT8 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("ixim/Image21-INT8", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Download tests/gpu_offload.py from ixim/Image21-INT8: direct link, hf CLI and curl.
- Browser
- Download file 1.64 kB
-
https://huggingface.co/ixim/Image21-INT8/resolve/main/tests/gpu_offload.py
- Command line
-
hf download hf://ixim/Image21-INT8/tests/gpu_offload.py
-
curl -L -o gpu_offload.py https://huggingface.co/ixim/Image21-INT8/resolve/main/tests/gpu_offload.py
1.64 kB
| """Regression: nested INT8 offload must move auxiliary tensors and preserve output.""" | |
| import argparse | |
| import scripts | |
| import torch | |
| import bitsandbytes as bnb | |
| def main(): | |
| ap = argparse.ArgumentParser() | |
| ap.add_argument('--fixed', action='store_true') | |
| args = ap.parse_args() | |
| torch.manual_seed(17) | |
| layer = bnb.nn.Linear8bitLt(512, 256, has_fp16_weights=False, threshold=6.0).to('cuda') | |
| model = torch.nn.Sequential(layer).eval() | |
| if args.fixed: | |
| from scripts.runtime import patch_int8_device_moves | |
| patch_int8_device_moves(model) | |
| x = torch.randn(2, 512, device='cuda', dtype=torch.float16) | |
| with torch.inference_mode(): | |
| expected = model(x).clone() | |
| for _ in range(3): | |
| model.to('cpu') | |
| tensors = [layer.weight, layer.weight.CB, layer.weight.SCB, | |
| layer.state.CB, layer.state.SCB, layer.state.idx] | |
| assert all(t is None or t.device.type == 'cpu' for t in tensors), 'CUDA tensors retained after parent.to(cpu)' | |
| model.to('cuda') | |
| with torch.inference_mode(): | |
| actual = model(x) | |
| torch.testing.assert_close(actual, expected, rtol=0, atol=0) | |
| # Also cover the pre-forward quantized parameter CB/SCB path. | |
| fresh = torch.nn.Sequential(bnb.nn.Linear8bitLt(512, 256, has_fp16_weights=False).to('cuda')) | |
| if args.fixed: | |
| patch_int8_device_moves(fresh) | |
| fresh.to('cpu') | |
| assert fresh[0].weight.CB.device.type == 'cpu' | |
| assert fresh[0].weight.SCB.device.type == 'cpu' | |
| print('PASS: auxiliary tensors offloaded; three exact-output CPU/GPU round trips') | |
| if __name__ == '__main__': | |
| main() | |