Text Generation
Transformers
Safetensors
designcoder
ui-generation
front-end
html
css
javascript
code-generation
full-sft
Instructions to use xingxm/DesignCoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xingxm/DesignCoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="xingxm/DesignCoder")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("xingxm/DesignCoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use xingxm/DesignCoder with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "xingxm/DesignCoder" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xingxm/DesignCoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/xingxm/DesignCoder
- SGLang
How to use xingxm/DesignCoder with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xingxm/DesignCoder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xingxm/DesignCoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "xingxm/DesignCoder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xingxm/DesignCoder", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use xingxm/DesignCoder with Docker Model Runner:
docker model run hf.co/xingxm/DesignCoder
Add designcoder_evaluator_qwen3.8_27b_adamw_bs256_data37847_step296/README.md
Browse files
designcoder_evaluator_qwen3.8_27b_adamw_bs256_data37847_step296/README.md
ADDED
|
@@ -0,0 +1,171 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
base_model: Qwen/Qwen3.8-27B
|
| 4 |
+
tags:
|
| 5 |
+
- designcoder
|
| 6 |
+
- evaluator
|
| 7 |
+
- reward-model
|
| 8 |
+
- ui-evaluation
|
| 9 |
+
- multimodal
|
| 10 |
+
- image-text-to-text
|
| 11 |
+
language:
|
| 12 |
+
- en
|
| 13 |
+
pipeline_tag: image-text-to-text
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# DesignCoder Evaluator · Qwen3.8-27B · AdamW · bs256 · data37847 · step296
|
| 17 |
+
|
| 18 |
+
> **This is an evaluator, not a generation model.**
|
| 19 |
+
> Every other checkpoint in this repository *writes* HTML/CSS/JS. This one *scores*
|
| 20 |
+
> the rendered result. It takes a UI screenshot plus a rubric and returns structured
|
| 21 |
+
> per-criterion scores — it is intended as a reward / judge model for rollout scoring,
|
| 22 |
+
> not for page generation.
|
| 23 |
+
|
| 24 |
+
## Task
|
| 25 |
+
|
| 26 |
+
Input: one full-page UI screenshot + `<surface>` (landing | dashboard) +
|
| 27 |
+
`<generation_brief>` + `<frozen_rubric>`.
|
| 28 |
+
Output: `<think>` reasoning followed by a strict JSON payload with
|
| 29 |
+
|
| 30 |
+
| Field | Content |
|
| 31 |
+
|---|---|
|
| 32 |
+
| `dynamic_scores` | 5 prompt-fit criteria, scored 0/1/2 |
|
| 33 |
+
| `static_scores` | 27~35 integrity + detail rubrics, 0/2 (integrity) or 0/1/2 (detail), `"N/A"` where genuinely inapplicable |
|
| 34 |
+
| `interaction_rubrics` | free-text interaction test cases |
|
| 35 |
+
| `frozen_dynamic_scores.verdicts` | binary 0/1 verdict per supplied rubric item |
|
| 36 |
+
|
| 37 |
+
## Training
|
| 38 |
+
|
| 39 |
+
| | |
|
| 40 |
+
|---|---|
|
| 41 |
+
| Base model | Qwen3.8-27B (`Qwen3_5ForConditionalGeneration`) |
|
| 42 |
+
| Dataset | **DesignCoder-evaluate**, 37,847 filtered samples |
|
| 43 |
+
| Optimizer | AdamW, lr 5e-6, cosine, warmup 0.1, wd 0.0 |
|
| 44 |
+
| Precision / parallel | bf16 + DeepSpeed ZeRO-3, 8 nodes × 8 H20 = 64 GPUs |
|
| 45 |
+
| Global batch | 256 (per_device 1 × grad_accum 4 × 64) |
|
| 46 |
+
| Epochs / steps | 2.0 / 296 |
|
| 47 |
+
| `cutoff_len` / `image_max_pixels` | 32768 / 1048576 |
|
| 48 |
+
| Frozen | vision tower (LM + projector fully tuned) |
|
| 49 |
+
| **train_loss** | **0.4939** |
|
| 50 |
+
|
| 51 |
+
> ⚠️ `data37847` refers to the **DesignCoder-evaluate** dataset (UI screenshot scoring).
|
| 52 |
+
> It is **not** the generation dataset used by the `data37865` / `data41287` checkpoints
|
| 53 |
+
> in this repo — the similar sample count is a coincidence.
|
| 54 |
+
|
| 55 |
+
### Checkpoint selection
|
| 56 |
+
|
| 57 |
+
Both saved checkpoints were evaluated on a purpose-built **out-of-distribution suite**
|
| 58 |
+
(576 prompts: rubric lengths 5/10/15/20/25 plus three dimension-grouping shapes).
|
| 59 |
+
Unlike the generation runs in this repo — and unlike the 9B evaluator, which peaks at
|
| 60 |
+
34% of training — the 27B shows **no late-training degradation**:
|
| 61 |
+
|
| 62 |
+
| Step | OOD κ | In-dist κ | Overfit gap | OOD count-match |
|
| 63 |
+
|---:|---:|---:|---:|---:|
|
| 64 |
+
| 200 | 0.6625 | 0.6643 | +0.002 | 0.9977 |
|
| 65 |
+
| **296 (this one)** | **0.6640** | 0.6353 | **−0.029** | **1.0000** |
|
| 66 |
+
|
| 67 |
+
The two are statistically indistinguishable on κ (Δ 0.0015, well inside variant-to-variant
|
| 68 |
+
noise), but step 296 has the better overfit gap and perfect verdict counting, so the final
|
| 69 |
+
checkpoint is published.
|
| 70 |
+
|
| 71 |
+
## Evaluation
|
| 72 |
+
|
| 73 |
+
The dataset has no test split, so evaluation uses **held-out inputs**: 239 generated-page
|
| 74 |
+
screenshots from the DesignCoder benchmark, each already scored check-by-check by a
|
| 75 |
+
stronger external vision judge. Byte-level md5 check confirmed **0 overlap** between these
|
| 76 |
+
710 benchmark screenshots and the 37,851 training images.
|
| 77 |
+
|
| 78 |
+
Because the training rubric has a **fixed shape** (10 items:
|
| 79 |
+
`Alignment:2, Layout:2, Typography:2, Components:2, Assets:1, Aesthetics:1`, 100% of
|
| 80 |
+
3,000 sampled records) while benchmark rubrics carry 24~25 items, two variants were run:
|
| 81 |
+
**Set B** = original benchmark rubric (out-of-distribution length),
|
| 82 |
+
**Set C** = same screenshots, rubric down-sampled to the training shape.
|
| 83 |
+
|
| 84 |
+
| Metric | Set B | Set C |
|
| 85 |
+
|---|---|---|
|
| 86 |
+
| JSON parse rate | **1.0000** | **1.0000** |
|
| 87 |
+
| Verdict-count exact match | **1.0000** | **1.0000** |
|
| 88 |
+
| Per-check agreement | 0.9313 | 0.9347 |
|
| 89 |
+
| Cohen's κ | **0.6802** | 0.6645 |
|
| 90 |
+
| Defect recall (GT=0) | 0.6996 | 0.6958 |
|
| 91 |
+
| Defect precision | 0.7401 | 0.7066 |
|
| 92 |
+
| Predicted vs true defect rate | 0.119 / 0.126 | 0.108 / 0.110 |
|
| 93 |
+
| Case-level Pearson | **0.8572** | 0.7890 |
|
| 94 |
+
|
| 95 |
+
Agreement alone is misleading (87% of checks are `1`, so always-pass scores ~88%);
|
| 96 |
+
κ and defect recall are the meaningful signals.
|
| 97 |
+
|
| 98 |
+
### Reproducing the known quality ordering
|
| 99 |
+
|
| 100 |
+
Three generation checkpoints, ranked by the external judge — this model reproduces the
|
| 101 |
+
ordering with **<1pp absolute error**, which is what matters for reward use:
|
| 102 |
+
|
| 103 |
+
| Scorer | 4B-Muon | 9B-AdamW | 27B-v1 | Ordering |
|
| 104 |
+
|---|---|---|---|---|
|
| 105 |
+
| External judge | 0.833 | 0.884 | 0.905 | reference |
|
| 106 |
+
| **This model** | 0.840 | 0.893 | 0.909 | ✅ correct |
|
| 107 |
+
| 9B evaluator (Muon) | 0.896 | 0.940 | 0.933 | ❌ inverted |
|
| 108 |
+
|
| 109 |
+
## Usage
|
| 110 |
+
|
| 111 |
+
```python
|
| 112 |
+
from transformers import AutoModelForImageTextToText, AutoProcessor, AutoTokenizer
|
| 113 |
+
from PIL import Image
|
| 114 |
+
import torch, math
|
| 115 |
+
|
| 116 |
+
MODEL = "xingxm/DesignCoder"
|
| 117 |
+
SUB = "designcoder_evaluator_qwen3.8_27b_adamw_bs256_data37847_step296"
|
| 118 |
+
|
| 119 |
+
tok = AutoTokenizer.from_pretrained(MODEL, subfolder=SUB, trust_remote_code=True)
|
| 120 |
+
proc = AutoProcessor.from_pretrained(MODEL, subfolder=SUB, trust_remote_code=True)
|
| 121 |
+
model = AutoModelForImageTextToText.from_pretrained(
|
| 122 |
+
MODEL, subfolder=SUB, dtype=torch.bfloat16, device_map="cuda", trust_remote_code=True
|
| 123 |
+
)
|
| 124 |
+
|
| 125 |
+
img = Image.open("screenshot.png").convert("RGB")
|
| 126 |
+
w, h = img.size # match training preprocessing
|
| 127 |
+
if w * h > 1048576:
|
| 128 |
+
s = math.sqrt(1048576 / (w * h))
|
| 129 |
+
img = img.resize((int(w * s), int(h * s)))
|
| 130 |
+
|
| 131 |
+
user = (
|
| 132 |
+
"<image>\n"
|
| 133 |
+
"Evaluate the attached UI screenshot using the selected visual criteria.\n\n"
|
| 134 |
+
"<surface>\nlanding\n</surface>\n\n"
|
| 135 |
+
"<generation_brief>\n...brief...\n</generation_brief>\n\n"
|
| 136 |
+
'<frozen_rubric>\n{"rubric":{"Alignment":["..."],"Layout":["..."]}}\n</frozen_rubric>'
|
| 137 |
+
)
|
| 138 |
+
text = tok.apply_chat_template(
|
| 139 |
+
[{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user}],
|
| 140 |
+
tokenize=False, add_generation_prompt=True,
|
| 141 |
+
).replace("<image>", "<|vision_start|><|image_pad|><|vision_end|>")
|
| 142 |
+
|
| 143 |
+
inputs = proc(text=[text], images=[img], return_tensors="pt").to("cuda")
|
| 144 |
+
out = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
|
| 145 |
+
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
|
| 146 |
+
```
|
| 147 |
+
|
| 148 |
+
`SYSTEM_PROMPT` must be the one shipped with the DesignCoder-evaluate dataset — the model
|
| 149 |
+
relies on it to enumerate the rubric namespace. The chat template already ends with
|
| 150 |
+
`<think>\n`, so generated text continues *inside* the reasoning block and contains no
|
| 151 |
+
opening `<think>` tag.
|
| 152 |
+
|
| 153 |
+
## Limitations
|
| 154 |
+
|
| 155 |
+
- **`summary` field miscounts.** `frozen_dynamic_scores.summary` is self-consistent in only
|
| 156 |
+
2.1% of cases (e.g. writes `"24/24"` when 25 verdicts were emitted). Parse the
|
| 157 |
+
`verdicts` array; ignore `summary`.
|
| 158 |
+
- **`static_scores` / `dynamic_scores` are unvalidated on unseen data.** No ground truth
|
| 159 |
+
exists for these 40 items outside the training distribution; they were only checked on a
|
| 160 |
+
contaminated training sample (static exact-match 0.9196, dynamic 0.8376).
|
| 161 |
+
- **Trained on a single rubric shape.** All training rubrics have exactly 10 items. The
|
| 162 |
+
model handles 25-item rubrics correctly, but longer or structurally different rubrics are
|
| 163 |
+
untested.
|
| 164 |
+
- **Judgement quality is not where this model beats the 9B.** On the OOD suite the
|
| 165 |
+
[9B evaluator](../designcoder_evaluator_qwen3.5_9b_adamw_bs256_data37847_step100)
|
| 166 |
+
actually scores slightly higher on κ (0.694 vs 0.664). The 27B's real advantage is
|
| 167 |
+
**format robustness**: verdict-count match stays at 1.00 across rubric lengths 5→25,
|
| 168 |
+
where the 9B drops to 0.00 at length 5 and ~0.85 at length 25. Use this model when
|
| 169 |
+
rubric length varies; the 9B is the cheaper choice for fixed 10-item rubrics.
|
| 170 |
+
- Evidence is screenshot-only; the model cannot verify interaction behaviour, yet
|
| 171 |
+
`interaction_rubrics` asks it to describe interactions — treat that field as speculative.
|