How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("syaffers/tiny-random-QwenImagePipeline", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

tiny-random-QwenImagePipeline

A tiny, randomly initialised QwenImagePipeline (diffusers) intended only for smoke testing. All weights are random: it does not produce usable images. It exists so that code paths (loading, serving, CI) can be exercised quickly without downloading the real multi-billion-parameter checkpoint.

The script that generated this repo is available in the Reproducing this model section below.

Model Details

The pipeline mirrors the structure of Qwen/Qwen-Image-2512, with every component shrunk dramatically (~10M parameters in total, ~33 MB on disk, weights in bfloat16):

Component Class Notes
text_encoder Qwen2_5_VLForConditionalGeneration (transformers) 2 text layers, hidden size 64; 2-layer vision tower, hidden size 32
transformer QwenImageTransformer2DModel (diffusers) 2 layers, 4 heads, head dim 16
vae AutoencoderKLQwenImage (diffusers) base dim 8, 16 latent channels, 8x spatial downsampling (as in the real VAE)
scheduler FlowMatchEulerDiscreteScheduler default config
tokenizer Qwen2Tokenizer copied from the real Qwen-Image tokenizer, so the vocab size (152,064) matches the real model

Because the vocabulary is fixed by the real tokenizer, the text encoder ties its input and output embeddings to avoid a second 152k x hidden-size matrix.

Uses

Intended for testing pipelines, loaders, and serving stacks (e.g. vLLM-Omni) where only shapes, dtypes, and API behaviour matter. Do not use it to generate meaningful images.

How to Get Started with the Model

import torch
from diffusers import QwenImagePipeline

pipe = QwenImagePipeline.from_pretrained(
    "syaffers/tiny-random-Qwen2_5_VLForConditionalGeneration",
    torch_dtype=torch.bfloat16,
)
image = pipe(
    "a cup of coffee",
    height=256,
    width=256,
    num_inference_steps=2,
    true_cfg_scale=1.0,
).images[0]

The output is noise, since the weights are random. Example outputs are in samples/.

Serving with vLLM-Omni

Tested with the vllm/vllm-omni:v0.28.0 Docker image on an NVIDIA GPU. From the root of this repo (or after downloading it):

docker run -d --name tiny-omni --gpus all --ipc=host -p 8091:8091 \
  -v "$PWD":/model:ro \
  --entrypoint vllm vllm/vllm-omni:v0.28.0 \
  serve /model --omni --host 0.0.0.0 --port 8091

Once the log shows Application startup complete., request an image through the OpenAI-compatible endpoint:

curl -s localhost:8091/v1/images/generations \
  -H 'Content-Type: application/json' \
  -d '{"model": "/model", "prompt": "a cup of coffee", "size": "256x256",
       "n": 1, "num_inference_steps": 2, "response_format": "b64_json"}'

This returns HTTP 200 with a valid 256x256 PNG (random noise, as expected).

Sample outputs

This is what to expect: noise, regardless of the prompt. The model is only meant to verify that the pipeline runs end to end.

a cup of coffee on a wooden table a red bicycle leaning against a brick wall

Reproducing this model

The full creation script is below. Run it as python make_tiny_random_qwen_image.py <tokenizer_source> <output_dir>, where <tokenizer_source> is a repo or directory with the real Qwen tokenizer (e.g. Qwen/Qwen-Image-2512). It builds the pipeline with torch.manual_seed(0), saves it, then reloads it and generates a 256x256 image on CPU as a sanity check.

make_tiny_random_qwen_image.py (click to expand)
"""Build a tiny, randomly-initialised QwenImagePipeline that loads and serves out of the box."""
import sys
import torch
from diffusers import AutoencoderKLQwenImage, FlowMatchEulerDiscreteScheduler, QwenImagePipeline, QwenImageTransformer2DModel
from transformers import AutoTokenizer, Qwen2_5_VLConfig, Qwen2_5_VLForConditionalGeneration

SRC_TOKENIZER = sys.argv[1]  # repo or dir providing the real Qwen tokenizer, e.g. Qwen/Qwen-Image-2512
OUT = sys.argv[2]

tok = AutoTokenizer.from_pretrained(SRC_TOKENIZER, subfolder="tokenizer")

TEXT_DIM = 64
HEADS = 4
HEAD_DIM = TEXT_DIM // HEADS  # 16 -> rope dims per head are 8, split by mrope_section below
text_cfg = dict(
    vocab_size=152064, hidden_size=TEXT_DIM, intermediate_size=128, num_hidden_layers=2,
    num_attention_heads=HEADS, num_key_value_heads=2, max_position_embeddings=128000, rms_norm_eps=1e-6,
    rope_theta=1000000.0, rope_scaling={"rope_type": "default", "type": "default", "mrope_section": [2, 3, 3]},
    tie_word_embeddings=True,  # vocab is fixed by the real tokenizer, so avoid a second 152k x dim matrix
)
vision_cfg = dict(
    depth=2, hidden_size=32, intermediate_size=64, num_heads=2, out_hidden_size=TEXT_DIM, patch_size=14,
    spatial_merge_size=2, temporal_patch_size=2, window_size=112, fullatt_block_indexes=[1], in_channels=3,
)
cfg = Qwen2_5_VLConfig(
    text_config=text_cfg, vision_config=vision_cfg, image_token_id=151655, video_token_id=151656,
    vision_start_token_id=151652, vision_end_token_id=151653, tie_word_embeddings=True,
)

torch.manual_seed(0)
text_encoder = Qwen2_5_VLForConditionalGeneration(cfg)

LATENT_CH = 16
vae = AutoencoderKLQwenImage(
    base_dim=8, z_dim=LATENT_CH, dim_mult=[1, 1, 1, 1], num_res_blocks=1, attn_scales=[],
    temperal_downsample=[False, True, True],  # 8x spatial downsampling, same as the real VAE
    latents_mean=[0.0] * LATENT_CH, latents_std=[1.0] * LATENT_CH,
)
transformer = QwenImageTransformer2DModel(
    patch_size=2, in_channels=LATENT_CH * 4, out_channels=LATENT_CH, num_layers=2,
    attention_head_dim=HEAD_DIM, num_attention_heads=HEADS, joint_attention_dim=TEXT_DIM,
    guidance_embeds=False, axes_dims_rope=(4, 6, 6),  # must sum to attention_head_dim
)
pipe = QwenImagePipeline(
    transformer=transformer, scheduler=FlowMatchEulerDiscreteScheduler(), vae=vae,
    text_encoder=text_encoder, tokenizer=tok,
)
pipe.to(torch.bfloat16)  # the real repo ships bf16; vLLM-Omni derives its latent dtype from the checkpoint
pipe.save_pretrained(OUT, safe_serialization=True)

# sanity check: reload and generate on CPU
p = QwenImagePipeline.from_pretrained(OUT, torch_dtype=torch.bfloat16)
img = p("a cup of coffee", height=256, width=256, num_inference_steps=2, true_cfg_scale=1.0).images[0]
print("ok", img.size, sum(x.numel() for m in (p.transformer, p.vae, p.text_encoder) for x in m.parameters()) / 1e6, "M params")
Downloads last month
7
Safetensors
Model size
340k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support