Instructions to use devon7y/WikiQwen-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use devon7y/WikiQwen-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="devon7y/WikiQwen-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("devon7y/WikiQwen-4B") model = AutoModelForCausalLM.from_pretrained("devon7y/WikiQwen-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use devon7y/WikiQwen-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "devon7y/WikiQwen-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devon7y/WikiQwen-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/devon7y/WikiQwen-4B
- SGLang
How to use devon7y/WikiQwen-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "devon7y/WikiQwen-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devon7y/WikiQwen-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "devon7y/WikiQwen-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devon7y/WikiQwen-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use devon7y/WikiQwen-4B with Docker Model Runner:
docker model run hf.co/devon7y/WikiQwen-4B
WikiQwen-4B
Ask anything, get a how-to article. WikiQwen-4B is Qwen3.5-4B fine-tuned to answer every message the same way: as a tidy, step-by-step how-to article in markdown, with a picture caption above each step. Hand those captions to WikiQwen-Illustrator and they become illustrations.
Try it first: the WikiQwen and WikiQwen Plus Spaces run this model. WikiQwen draws each step with WikiQwen-Illustrator, and WikiQwen Plus with WikiQwen-Illustrator-9B.
What It Does
Send it a question ("how do I keep basil alive?"), a problem ("my bike chain keeps falling off") or just "hi", and it sends back a guide. Every reply has the same shape:
- a
# How to …title and a short intro - one or more
## Method N:or## Part N:sections, or a single## Steps - an
[IMAGE: caption]line above each step, written for the Illustrator to draw - steps written as
**N. Bold summary.** Details…, with bullets and> **Tip:**boxes where they help - optional
## Tips,## Warningsand## Things You'll Needat the end
Here is that shape, trimmed. It was written by hand to show the format, so it is not a verbatim model output:
# How to Keep Basil Alive on a Windowsill
Basil wants three things: sun, warmth and a steady drink. Get those right and it will keep you in pesto all summer.
## Part 1: Giving It the Right Spot
[IMAGE: A potted basil plant with bright green leaves sits on a sunny white windowsill, with light streaming through the glass behind it.]
**1. Pick your sunniest window.** Basil needs 6 to 8 hours of direct light a day, so a south-facing window is ideal.
- Give the pot a quarter turn every few days so the plant grows evenly.
[IMAGE: Hands pour water from a small green watering can into the soil of a potted basil plant until it drips into the saucer below.]
**2. Water when the top of the soil feels dry.** Soak the pot until water runs out of the bottom, then empty the saucer.
> **Tip:** Water in the morning so the leaves dry off before the evening chill.
## Warnings
- Don't let the pot sit in standing water. Soggy roots rot quickly.
## Things You'll Need
- A pot with drainage holes
- Potting mix
- A watering can
How to Run WikiQwen-4B
Getting your first article takes about as long as boiling the kettle. The model is about 8 GB in bf16, so a single consumer GPU is enough.
1. Install the libraries. You need PyTorch, Transformers 5 and Accelerate. The demo Space uses transformers==5.14.1.
pip install torch "transformers>=5.14" accelerate
- Optional:
pip install flash-linear-attentionfor faster kernels on Qwen3.5's linear-attention layers (the demo Space uses it).
2. Load the model and tokenizer. The LoRA is already merged into the weights, so it loads like any Qwen3.5 checkpoint.
3. Ask your question with no system prompt. The article habit is trained in. There was no system prompt in training, and you don't need one now.
4. Turn thinking off. Qwen3.5's chat template can open a <think> block. WikiQwen was trained with it closed, so pass enable_thinking=False. Leave it on and the model muses to itself before it starts writing.
5. Sample like the demo does. Temperature 0.7, top_p 0.9, repetition penalty 1.05 and up to 1,800 new tokens. That is room for about two methods with captions.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "devon7y/WikiQwen-4B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "How do I keep basil alive on a windowsill?"}] # no system prompt
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=1800, do_sample=True, temperature=0.7, top_p=0.9,
repetition_penalty=1.05)
article = tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(article)
6. Pull out the captions if you want pictures. Each [IMAGE: …] line is a ready-made prompt for WikiQwen-Illustrator.
import re
captions = re.findall(r"^\[IMAGE: (.+?)\]\s*$", article, flags=re.M)
Tip: Follow-up messages work (append the reply as an assistant turn and ask again). The model was trained on single questions, though, so each answer is a fresh article, not a chatty reply.
Warnings
- Don't skip
enable_thinking=False. It is the one setting that changes everything. - Long answers can now and then repeat a step. The demo Space stops generating when the same step or caption shows up a third time, and trims an unfinished last line. Do the same if you serve it.
Things You'll Need
- A CUDA GPU with room for about 8 GB of bf16 weights plus the context
- Python with
torch,transformers5.x andaccelerate - A question (any question)
The Family
| Model | Base | Size | Notes |
|---|---|---|---|
| WikiQwen-4B (this one) | Qwen3.5-4B | 4B, dense | Fastest; in the WikiQwen and WikiQwen Plus Spaces |
| WikiQwen-9B | Qwen3.5-9B | 9B, dense | Same recipe, more capacity; in the WikiQwen Pro Space |
| WikiQwen-9B-FP8 | WikiQwen-9B | 9B, FP8 weights | 11 GB of weights instead of 18 GB; slower than bf16 on current kernels |
| WikiQwen-35B-A3B | Qwen3.6-35B-A3B | 35B MoE, about 3B active | The largest |
| WikiQwen-35B-A3B-FP8 | WikiQwen-35B-A3B | 35B MoE, FP8 experts | 37 GB instead of 69 GB: fits on one GPU |
| WikiQwen-Illustrator | FLUX.2 [klein] base 4B | LoRA | Draws the [IMAGE: …] captions |
| WikiQwen-Illustrator-9B | FLUX.2 [klein] base 9B | LoRA | Sharper hands and object placement; FLUX non-commercial license |
Training
- Source: about 118k how-to articles from the Kiwix March 2023 archive, with text and Creative Commons illustrations by wikiHow contributors, CC BY-NC-SA 3.0 (see Attribution and License).
- From article to chat: each article became one conversation. The user turn is a realistic message written for that article by Qwen3.6-35B-A3B. The styles range from direct questions to oblique remarks and messages that aren't how-to questions at all, which is how the model learned to answer anything with an article. The assistant turn is the article in the markdown format above, keeping at most its first two methods or parts.
- Captions: every step carries an
[IMAGE: …]line. (pending) of these captions describe the step's real Creative Commons illustration, written by a vision-language model (Qwen3.6-35B-A3B). The rest were written from the step text alone. The Illustrator was trained on captions in the same style, so the chat model writes prompts the image model already understands. - Size: 107,054 training conversations and 300 for validation. A 1% bucket of articles never entered training, and the held-out evaluation prompts below come from it.
- Recipe: supervised fine-tuning with LoRA (rank 32, alpha 64, dropout 0.05) on the attention and MLP projections (q, k, v, o, gate, up, down). Loss was on the assistant turn only. One epoch, learning rate 1e-4 with a cosine schedule and 3% warm-up, effective batch 16, sequences up to 4,096 tokens, bf16, on one GPU. The adapter was then merged into the base weights.
Evaluation
The model wrote one reply per prompt for two prompt sets:
- Held-out prompts: user messages for articles the model never saw in training.
- Everyday prompts: 110 messages that mostly aren't how-to questions (open-assistant questions, everyday small talk and a few handwritten ones like "tell me a joke").
Replies were sampled at temperature 0.7, top_p 0.9 and repetition penalty 1.05, then checked automatically against the article format. These checks measure the shape of the answer. They don't check whether the advice is right.
| Format check | Held-out prompts (n=100) | Everyday prompts (n=110) |
|---|---|---|
Starts with a # Title line |
100% | 100% |
| Title reads "How to …" | 100% | 92% |
| Steps numbered 1, 2, 3… in every section | 100% | 100% |
Steps with an [IMAGE: …] line above (mean) |
100% | 100% |
## sections per article (median) |
2 | 2 |
| Numbered steps per article (median) | 9 | 9 |
| Words per image caption (median) | 35.1 | 35.7 |
| Words per article (median) | 1,039 | 1,036.5 |
| Cut off by the token limit | 1% | 5% |
Limitations
- The advice can be wrong. It sounds confident and looks tidy either way. Be most careful with health, safety, legal and money questions: check anything that matters with a qualified source.
- Everything is an article. "hi", "tell me a joke" and "what's 2 + 2" all get a how-to article back. That is the point, but it means this is not a general assistant.
- Knowledge is dated and general. It knows what the base model knows plus a 2023 snapshot of how-to articles. It can't browse, and it may invent product names, measurements or steps.
- Generated articles are not real articles from the source site, and no contributor wrote or reviewed them.
- Pictures can mislead. If you draw the captions with WikiQwen-Illustrator, text inside the pictures comes out garbled, and a picture can show a step wrongly.
- No extra safety training. The fine-tune only taught a format, and the article habit is strong. It may write an article for a request it should turn down, so don't rely on it to refuse.
- English only, and long answers can be cut off at the token limit.
Attribution and License
- Source credit: fine-tuned on article text and Creative Commons illustrations by wikiHow contributors, licensed CC BY-NC-SA 3.0, from the Kiwix March 2023 archive.
- This model and what it writes: CC BY-NC-SA 4.0. Section 4(b) of CC BY-NC-SA 3.0 allows an adaptation to be shared under a later version of the license with the same elements (Attribution, NonCommercial, ShareAlike). In short: credit the source, no commercial use, and share what you make under the same license.
- Base model: Qwen3.5-4B is released by the Qwen team under Apache-2.0 (license).
- Suggested credit line for articles you share: "Generated by WikiQwen-4B (CC BY-NC-SA 4.0), a model trained on text and illustrations by wikiHow contributors, CC BY-NC-SA 3.0."
- Independence: WikiQwen is an independent, non-commercial research project. It is not affiliated with or endorsed by the source site or its contributors, or by the Qwen team.
Citation
@misc{wikiqwen_4b_2026,
title = {WikiQwen-4B: a Qwen3.5-4B fine-tune that answers as illustrated how-to articles},
author = {devon7y},
year = {2026},
url = {https://huggingface.co/devon7y/WikiQwen-4B}
}
- Downloads last month
- 975