Instructions to use devon7y/WikiQwen-Illustrator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use devon7y/WikiQwen-Illustrator with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("black-forest-labs/FLUX.2-klein-base-4B", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("devon7y/WikiQwen-Illustrator") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("black-forest-labs/FLUX.2-klein-base-4B", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("devon7y/WikiQwen-Illustrator")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]WikiQwen-Illustrator
Preview weights. This repo currently holds round 2 at step 8,000 of 18,000 (full ~588k-image set, rank 128). Training continues and the final weights will replace these.
Give it a caption, get a step illustration. WikiQwen-Illustrator is a style LoRA for FLUX.2 [klein] base 4B. It turns a one-to-three-sentence scene description into a clean, friendly how-to illustration at 512x384. It is the picture half of WikiQwen: the chat models (4B, 9B, 35B-A3B) write an [IMAGE: caption] line above every step, and this LoRA draws it.
Try it first: the demo Space writes an article with WikiQwen-4B, then draws every step with this LoRA.
What It Does
Inside a WikiQwen article, a step looks like this. The first line is the Illustrator's prompt, and the second is the step it illustrates:
[IMAGE: Hands pour water from a small green watering can into the soil of a potted basil plant on a sunny white windowsill.]
**2. Water when the top of the soil feels dry.** Soak the pot until water runs out of the bottom, then empty the saucer.
- Input: a plain description of the scene: who is in the frame (or just hands), what they are doing, the objects, the colors, the background, and any arrows, check marks or red X marks. One to three sentences, present tense.
- Output: a 512x384 picture in the flat, bright look of instructional step illustrations.
- No trigger word. The training captions describe content only, with no style words, so the style comes along automatically.
How to Draw a Step Illustration
It takes a few minutes, and most of that is downloading the base model.
1. Install the libraries. You need a diffusers release with Flux2KleinPipeline; the demo Space uses 0.38.0. PEFT is needed to load LoRAs.
pip install torch "diffusers>=0.38.0" transformers accelerate peft
2. Load the base pipeline. Use the undistilled base model. The LoRA was trained on it, and the demo runs on it.
3. Load the LoRA. Always pass weight_name. The file in this repo is pytorch_lora_weights.safetensors.
4. Write a caption like the ones it learned. Start with the main subject, describe only what is visible, and keep it to one to three sentences (the training captions ran 20 to 45 words).
5. Draw at 512x384 with 16 steps and guidance 4.0. These are the demo Space's settings. In our sweep, 16 steps looked close to the base model's usual 50, in about a third of the time.
import torch
from diffusers import Flux2KleinPipeline
pipe = Flux2KleinPipeline.from_pretrained("black-forest-labs/FLUX.2-klein-base-4B", torch_dtype=torch.bfloat16)
pipe.load_lora_weights("devon7y/WikiQwen-Illustrator", weight_name="pytorch_lora_weights.safetensors",
adapter_name="wikiqwen")
pipe.to("cuda")
caption = ("Hands pour water from a small green watering can into the soil of a potted basil plant "
"on a sunny white windowsill.")
image = pipe(prompt=caption, height=384, width=512, num_inference_steps=16, guidance_scale=4.0,
generator=torch.Generator("cuda").manual_seed(0)).images[0]
image.save("step.png")
6. Fuse it for speed (optional). Serving one style all day? Bake it into the weights, as the demo does:
pipe.fuse_lora(lora_scale=1.0)
pipe.unload_lora_weights()
Tip: Draw several steps at once by passing a list of captions and one generator per caption. The demo draws four at a time.
Warnings
- There is no safety checker. The FLUX.2 [klein] pipeline doesn't include one. The demo screens captions against a block list before drawing, so do something similar if strangers can type prompts.
- 512x384 is the trained size. Other sizes (multiples of 16) will run, but the style was learned at 512x384.
- Leave
max_sequence_lengthand the text-encoder layers at the pipeline defaults. The LoRA was trained with exactly those (512 tokens; layers 9, 18 and 27).
Things You'll Need
- A CUDA GPU with room for about 16 GB of bf16 weights (transformer, Qwen3 text encoder and VAE), plus activations
- Python with
torch,diffusers0.38.0 or newer,transformers,accelerateandpeft - A caption, ideally one written by a WikiQwen chat model
Training
- Base: black-forest-labs/FLUX.2-klein-base-4B, the undistilled base model.
- Images: 587,815 Creative Commons step illustrations by wikiHow contributors, CC BY-NC-SA 3.0, from the Kiwix March 2023 archive (see Attribution and License).
- Preparation: only images a vision-language model typed as illustrations (not photos, screenshots or diagrams), with near-duplicates removed. Each image was cropped and resized to 512x384.
- Captions: written by Qwen3.6-35B-A3B from each image and its step text. They describe what is visible and never the art style. The chat models were trained on the same caption style, so their
[IMAGE: …]lines are prompts this LoRA already understands. - Recipe (round 2): LoRA rank 128 (alpha 128) on every linear layer of the transformer blocks, trained for about one epoch at an effective batch of 32. AdamW, learning rate 1e-4 with a cosine schedule and 500 warm-up steps, weight decay 1e-4, gradient clipping at 1.0, bf16 with gradient checkpointing, and a single 512x384 bucket. Text-encoder layers 9, 18 and 27, captions padded to 512 tokens and uniform timestep sampling, on one GPU. These preview weights are the step-8,000 checkpoint of the planned 18,000 (about 44% of the epoch); the finished round-2 LoRA will replace them.
- Round 1 was a bake-off on a 20k-image subset with rank-32 LoRAs to pick the base model and settings. Round 2 is this release, trained on the full set.
Evaluation
Each held-out caption was drawn once and compared with the real illustration for that caption. The held-out set comes from a 1% bucket of articles set aside for evaluation. Four measures:
- CMMD (CLIP embedding distance to the real illustrations) and KID (on DINOv2 features): lower is closer to the real style. The real-vs-real row, two halves of the real set compared with each other, is the noise floor.
- Real-vs-generated probe: accuracy of a linear classifier that tries to tell real from generated for the same captions. 0.5 means it can't tell.
- Caption alignment: SigLIP2 image-text similarity between each picture and its caption. The real illustrations' score comes after the slash; a big gap means prompts are being followed less closely.
| Setup | Images | CMMD ↓ | KID ×10³ ↓ | Real-vs-generated probe ↓ | Caption alignment ↑ (generated / real) |
|---|---|---|---|---|---|
| Preview: round 2 step 8000 | 300 | 0.145 | 96.44 ± 37.84 | 0.968 ± 0.012 | 0.2040 / 0.1965 |
| Round 1 bake-off (step 3000) | 300 | 0.184 | 108.05 ± 35.23 | 0.977 ± 0.006 | 0.2063 / 0.1965 |
| Real vs. real (two halves of the held-out set) | 150 + 150 | 0.041 | -23.53 | 0.5 = chance | 0.1965 (real) |
Scored against 300 held-out captions and their real illustrations.
Limitations
- Text inside pictures is garbled. Labels, signs and screens come out as squiggles, so leave words out of captions or expect nonsense.
- Details can be wrong: extra fingers, odd tool shapes, wrong counts, or a step shown in a way that would not work. The pictures decorate an article. They are not instructions to trust on their own.
- It is a style, not a general art model. It pulls everything toward one instructional look and may ignore requests for other styles.
- It inherits its training images' habits, including who tends to appear in them and how scenes are framed.
- English captions only, and only the 512x384 size was trained.
Attribution and License
- Source credit: trained on Creative Commons illustrations by wikiHow contributors, licensed CC BY-NC-SA 3.0, from the Kiwix March 2023 archive.
- This LoRA and the pictures it makes: CC BY-NC-SA 4.0. Section 4(b) of CC BY-NC-SA 3.0 allows an adaptation to be shared under a later version of the license with the same elements (Attribution, NonCommercial, ShareAlike). In short: credit the source, no commercial use, and share what you make under the same license.
- Base model: using this LoRA requires FLUX.2 [klein] base 4B, which has its own license: see the base model's license on black-forest-labs/FLUX.2-klein-base-4B.
- Suggested credit line for pictures you share: "Drawn with WikiQwen-Illustrator (CC BY-NC-SA 4.0), a LoRA trained on illustrations by wikiHow contributors, CC BY-NC-SA 3.0."
- Independence: WikiQwen is an independent, non-commercial research project. It is not affiliated with or endorsed by the source site or its contributors, or by Black Forest Labs.
Citation
@misc{wikiqwen_illustrator_2026,
title = {WikiQwen-Illustrator: a how-to illustration style LoRA for FLUX.2 [klein] base 4B},
author = {devon7y},
year = {2026},
url = {https://huggingface.co/devon7y/WikiQwen-Illustrator}
}
- Downloads last month
- 28
Model tree for devon7y/WikiQwen-Illustrator
Base model
black-forest-labs/FLUX.2-klein-base-4B