Instructions to use syaffers/tiny-random-QwenImagePipeline with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use syaffers/tiny-random-QwenImagePipeline with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("syaffers/tiny-random-QwenImagePipeline", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
tiny-random-QwenImagePipeline
A tiny, randomly initialised QwenImagePipeline (diffusers) intended only for smoke testing. All weights are random: it does not produce usable images. It exists so that code paths (loading, serving, CI) can be exercised quickly without downloading the real multi-billion-parameter checkpoint.
The script that generated this repo is available in the Reproducing this model section below.
Model Details
The pipeline mirrors the structure of Qwen/Qwen-Image-2512, with every component shrunk dramatically (~10M parameters in total, ~33 MB on disk, weights in bfloat16):
| Component | Class | Notes |
|---|---|---|
text_encoder |
Qwen2_5_VLForConditionalGeneration (transformers) |
2 text layers, hidden size 64; 2-layer vision tower, hidden size 32 |
transformer |
QwenImageTransformer2DModel (diffusers) |
2 layers, 4 heads, head dim 16 |
vae |
AutoencoderKLQwenImage (diffusers) |
base dim 8, 16 latent channels, 8x spatial downsampling (as in the real VAE) |
scheduler |
FlowMatchEulerDiscreteScheduler |
default config |
tokenizer |
Qwen2Tokenizer |
copied from the real Qwen-Image tokenizer, so the vocab size (152,064) matches the real model |
Because the vocabulary is fixed by the real tokenizer, the text encoder ties its input and output embeddings to avoid a second 152k x hidden-size matrix.
Uses
Intended for testing pipelines, loaders, and serving stacks (e.g. vLLM-Omni) where only shapes, dtypes, and API behaviour matter. Do not use it to generate meaningful images.
How to Get Started with the Model
import torch
from diffusers import QwenImagePipeline
pipe = QwenImagePipeline.from_pretrained(
"syaffers/tiny-random-Qwen2_5_VLForConditionalGeneration",
torch_dtype=torch.bfloat16,
)
image = pipe(
"a cup of coffee",
height=256,
width=256,
num_inference_steps=2,
true_cfg_scale=1.0,
).images[0]
The output is noise, since the weights are random. Example outputs are in samples/.
Serving with vLLM-Omni
Tested with the vllm/vllm-omni:v0.28.0 Docker image on an NVIDIA GPU. From the root of this repo (or after downloading it):
docker run -d --name tiny-omni --gpus all --ipc=host -p 8091:8091 \
-v "$PWD":/model:ro \
--entrypoint vllm vllm/vllm-omni:v0.28.0 \
serve /model --omni --host 0.0.0.0 --port 8091
Once the log shows Application startup complete., request an image through the OpenAI-compatible endpoint:
curl -s localhost:8091/v1/images/generations \
-H 'Content-Type: application/json' \
-d '{"model": "/model", "prompt": "a cup of coffee", "size": "256x256",
"n": 1, "num_inference_steps": 2, "response_format": "b64_json"}'
This returns HTTP 200 with a valid 256x256 PNG (random noise, as expected).
Sample outputs
This is what to expect: noise, regardless of the prompt. The model is only meant to verify that the pipeline runs end to end.
a cup of coffee on a wooden table |
a red bicycle leaning against a brick wall |
|---|---|
![]() |
![]() |
Reproducing this model
The full creation script is below. Run it as python make_tiny_random_qwen_image.py <tokenizer_source> <output_dir>, where <tokenizer_source> is a repo or directory with the real Qwen tokenizer (e.g. Qwen/Qwen-Image-2512). It builds the pipeline with torch.manual_seed(0), saves it, then reloads it and generates a 256x256 image on CPU as a sanity check.
make_tiny_random_qwen_image.py (click to expand)
"""Build a tiny, randomly-initialised QwenImagePipeline that loads and serves out of the box."""
import sys
import torch
from diffusers import AutoencoderKLQwenImage, FlowMatchEulerDiscreteScheduler, QwenImagePipeline, QwenImageTransformer2DModel
from transformers import AutoTokenizer, Qwen2_5_VLConfig, Qwen2_5_VLForConditionalGeneration
SRC_TOKENIZER = sys.argv[1] # repo or dir providing the real Qwen tokenizer, e.g. Qwen/Qwen-Image-2512
OUT = sys.argv[2]
tok = AutoTokenizer.from_pretrained(SRC_TOKENIZER, subfolder="tokenizer")
TEXT_DIM = 64
HEADS = 4
HEAD_DIM = TEXT_DIM // HEADS # 16 -> rope dims per head are 8, split by mrope_section below
text_cfg = dict(
vocab_size=152064, hidden_size=TEXT_DIM, intermediate_size=128, num_hidden_layers=2,
num_attention_heads=HEADS, num_key_value_heads=2, max_position_embeddings=128000, rms_norm_eps=1e-6,
rope_theta=1000000.0, rope_scaling={"rope_type": "default", "type": "default", "mrope_section": [2, 3, 3]},
tie_word_embeddings=True, # vocab is fixed by the real tokenizer, so avoid a second 152k x dim matrix
)
vision_cfg = dict(
depth=2, hidden_size=32, intermediate_size=64, num_heads=2, out_hidden_size=TEXT_DIM, patch_size=14,
spatial_merge_size=2, temporal_patch_size=2, window_size=112, fullatt_block_indexes=[1], in_channels=3,
)
cfg = Qwen2_5_VLConfig(
text_config=text_cfg, vision_config=vision_cfg, image_token_id=151655, video_token_id=151656,
vision_start_token_id=151652, vision_end_token_id=151653, tie_word_embeddings=True,
)
torch.manual_seed(0)
text_encoder = Qwen2_5_VLForConditionalGeneration(cfg)
LATENT_CH = 16
vae = AutoencoderKLQwenImage(
base_dim=8, z_dim=LATENT_CH, dim_mult=[1, 1, 1, 1], num_res_blocks=1, attn_scales=[],
temperal_downsample=[False, True, True], # 8x spatial downsampling, same as the real VAE
latents_mean=[0.0] * LATENT_CH, latents_std=[1.0] * LATENT_CH,
)
transformer = QwenImageTransformer2DModel(
patch_size=2, in_channels=LATENT_CH * 4, out_channels=LATENT_CH, num_layers=2,
attention_head_dim=HEAD_DIM, num_attention_heads=HEADS, joint_attention_dim=TEXT_DIM,
guidance_embeds=False, axes_dims_rope=(4, 6, 6), # must sum to attention_head_dim
)
pipe = QwenImagePipeline(
transformer=transformer, scheduler=FlowMatchEulerDiscreteScheduler(), vae=vae,
text_encoder=text_encoder, tokenizer=tok,
)
pipe.to(torch.bfloat16) # the real repo ships bf16; vLLM-Omni derives its latent dtype from the checkpoint
pipe.save_pretrained(OUT, safe_serialization=True)
# sanity check: reload and generate on CPU
p = QwenImagePipeline.from_pretrained(OUT, torch_dtype=torch.bfloat16)
img = p("a cup of coffee", height=256, width=256, num_inference_steps=2, true_cfg_scale=1.0).images[0]
print("ok", img.size, sum(x.numel() for m in (p.transformer, p.vae, p.text_encoder) for x in m.parameters()) / 1e6, "M params")
- Downloads last month
- 7

