xingxm commited on
Commit
fc0a557
·
verified ·
1 Parent(s): a9e2060

Add designcoder_evaluator_qwen3.8_27b_adamw_bs256_data37847_step296/README.md

Browse files
designcoder_evaluator_qwen3.8_27b_adamw_bs256_data37847_step296/README.md ADDED
@@ -0,0 +1,171 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ base_model: Qwen/Qwen3.8-27B
4
+ tags:
5
+ - designcoder
6
+ - evaluator
7
+ - reward-model
8
+ - ui-evaluation
9
+ - multimodal
10
+ - image-text-to-text
11
+ language:
12
+ - en
13
+ pipeline_tag: image-text-to-text
14
+ ---
15
+
16
+ # DesignCoder Evaluator · Qwen3.8-27B · AdamW · bs256 · data37847 · step296
17
+
18
+ > **This is an evaluator, not a generation model.**
19
+ > Every other checkpoint in this repository *writes* HTML/CSS/JS. This one *scores*
20
+ > the rendered result. It takes a UI screenshot plus a rubric and returns structured
21
+ > per-criterion scores — it is intended as a reward / judge model for rollout scoring,
22
+ > not for page generation.
23
+
24
+ ## Task
25
+
26
+ Input: one full-page UI screenshot + `<surface>` (landing | dashboard) +
27
+ `<generation_brief>` + `<frozen_rubric>`.
28
+ Output: `<think>` reasoning followed by a strict JSON payload with
29
+
30
+ | Field | Content |
31
+ |---|---|
32
+ | `dynamic_scores` | 5 prompt-fit criteria, scored 0/1/2 |
33
+ | `static_scores` | 27~35 integrity + detail rubrics, 0/2 (integrity) or 0/1/2 (detail), `"N/A"` where genuinely inapplicable |
34
+ | `interaction_rubrics` | free-text interaction test cases |
35
+ | `frozen_dynamic_scores.verdicts` | binary 0/1 verdict per supplied rubric item |
36
+
37
+ ## Training
38
+
39
+ | | |
40
+ |---|---|
41
+ | Base model | Qwen3.8-27B (`Qwen3_5ForConditionalGeneration`) |
42
+ | Dataset | **DesignCoder-evaluate**, 37,847 filtered samples |
43
+ | Optimizer | AdamW, lr 5e-6, cosine, warmup 0.1, wd 0.0 |
44
+ | Precision / parallel | bf16 + DeepSpeed ZeRO-3, 8 nodes × 8 H20 = 64 GPUs |
45
+ | Global batch | 256 (per_device 1 × grad_accum 4 × 64) |
46
+ | Epochs / steps | 2.0 / 296 |
47
+ | `cutoff_len` / `image_max_pixels` | 32768 / 1048576 |
48
+ | Frozen | vision tower (LM + projector fully tuned) |
49
+ | **train_loss** | **0.4939** |
50
+
51
+ > ⚠️ `data37847` refers to the **DesignCoder-evaluate** dataset (UI screenshot scoring).
52
+ > It is **not** the generation dataset used by the `data37865` / `data41287` checkpoints
53
+ > in this repo — the similar sample count is a coincidence.
54
+
55
+ ### Checkpoint selection
56
+
57
+ Both saved checkpoints were evaluated on a purpose-built **out-of-distribution suite**
58
+ (576 prompts: rubric lengths 5/10/15/20/25 plus three dimension-grouping shapes).
59
+ Unlike the generation runs in this repo — and unlike the 9B evaluator, which peaks at
60
+ 34% of training — the 27B shows **no late-training degradation**:
61
+
62
+ | Step | OOD κ | In-dist κ | Overfit gap | OOD count-match |
63
+ |---:|---:|---:|---:|---:|
64
+ | 200 | 0.6625 | 0.6643 | +0.002 | 0.9977 |
65
+ | **296 (this one)** | **0.6640** | 0.6353 | **−0.029** | **1.0000** |
66
+
67
+ The two are statistically indistinguishable on κ (Δ 0.0015, well inside variant-to-variant
68
+ noise), but step 296 has the better overfit gap and perfect verdict counting, so the final
69
+ checkpoint is published.
70
+
71
+ ## Evaluation
72
+
73
+ The dataset has no test split, so evaluation uses **held-out inputs**: 239 generated-page
74
+ screenshots from the DesignCoder benchmark, each already scored check-by-check by a
75
+ stronger external vision judge. Byte-level md5 check confirmed **0 overlap** between these
76
+ 710 benchmark screenshots and the 37,851 training images.
77
+
78
+ Because the training rubric has a **fixed shape** (10 items:
79
+ `Alignment:2, Layout:2, Typography:2, Components:2, Assets:1, Aesthetics:1`, 100% of
80
+ 3,000 sampled records) while benchmark rubrics carry 24~25 items, two variants were run:
81
+ **Set B** = original benchmark rubric (out-of-distribution length),
82
+ **Set C** = same screenshots, rubric down-sampled to the training shape.
83
+
84
+ | Metric | Set B | Set C |
85
+ |---|---|---|
86
+ | JSON parse rate | **1.0000** | **1.0000** |
87
+ | Verdict-count exact match | **1.0000** | **1.0000** |
88
+ | Per-check agreement | 0.9313 | 0.9347 |
89
+ | Cohen's κ | **0.6802** | 0.6645 |
90
+ | Defect recall (GT=0) | 0.6996 | 0.6958 |
91
+ | Defect precision | 0.7401 | 0.7066 |
92
+ | Predicted vs true defect rate | 0.119 / 0.126 | 0.108 / 0.110 |
93
+ | Case-level Pearson | **0.8572** | 0.7890 |
94
+
95
+ Agreement alone is misleading (87% of checks are `1`, so always-pass scores ~88%);
96
+ κ and defect recall are the meaningful signals.
97
+
98
+ ### Reproducing the known quality ordering
99
+
100
+ Three generation checkpoints, ranked by the external judge — this model reproduces the
101
+ ordering with **<1pp absolute error**, which is what matters for reward use:
102
+
103
+ | Scorer | 4B-Muon | 9B-AdamW | 27B-v1 | Ordering |
104
+ |---|---|---|---|---|
105
+ | External judge | 0.833 | 0.884 | 0.905 | reference |
106
+ | **This model** | 0.840 | 0.893 | 0.909 | ✅ correct |
107
+ | 9B evaluator (Muon) | 0.896 | 0.940 | 0.933 | ❌ inverted |
108
+
109
+ ## Usage
110
+
111
+ ```python
112
+ from transformers import AutoModelForImageTextToText, AutoProcessor, AutoTokenizer
113
+ from PIL import Image
114
+ import torch, math
115
+
116
+ MODEL = "xingxm/DesignCoder"
117
+ SUB = "designcoder_evaluator_qwen3.8_27b_adamw_bs256_data37847_step296"
118
+
119
+ tok = AutoTokenizer.from_pretrained(MODEL, subfolder=SUB, trust_remote_code=True)
120
+ proc = AutoProcessor.from_pretrained(MODEL, subfolder=SUB, trust_remote_code=True)
121
+ model = AutoModelForImageTextToText.from_pretrained(
122
+ MODEL, subfolder=SUB, dtype=torch.bfloat16, device_map="cuda", trust_remote_code=True
123
+ )
124
+
125
+ img = Image.open("screenshot.png").convert("RGB")
126
+ w, h = img.size # match training preprocessing
127
+ if w * h > 1048576:
128
+ s = math.sqrt(1048576 / (w * h))
129
+ img = img.resize((int(w * s), int(h * s)))
130
+
131
+ user = (
132
+ "<image>\n"
133
+ "Evaluate the attached UI screenshot using the selected visual criteria.\n\n"
134
+ "<surface>\nlanding\n</surface>\n\n"
135
+ "<generation_brief>\n...brief...\n</generation_brief>\n\n"
136
+ '<frozen_rubric>\n{"rubric":{"Alignment":["..."],"Layout":["..."]}}\n</frozen_rubric>'
137
+ )
138
+ text = tok.apply_chat_template(
139
+ [{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user}],
140
+ tokenize=False, add_generation_prompt=True,
141
+ ).replace("<image>", "<|vision_start|><|image_pad|><|vision_end|>")
142
+
143
+ inputs = proc(text=[text], images=[img], return_tensors="pt").to("cuda")
144
+ out = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
145
+ print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
146
+ ```
147
+
148
+ `SYSTEM_PROMPT` must be the one shipped with the DesignCoder-evaluate dataset — the model
149
+ relies on it to enumerate the rubric namespace. The chat template already ends with
150
+ `<think>\n`, so generated text continues *inside* the reasoning block and contains no
151
+ opening `<think>` tag.
152
+
153
+ ## Limitations
154
+
155
+ - **`summary` field miscounts.** `frozen_dynamic_scores.summary` is self-consistent in only
156
+ 2.1% of cases (e.g. writes `"24/24"` when 25 verdicts were emitted). Parse the
157
+ `verdicts` array; ignore `summary`.
158
+ - **`static_scores` / `dynamic_scores` are unvalidated on unseen data.** No ground truth
159
+ exists for these 40 items outside the training distribution; they were only checked on a
160
+ contaminated training sample (static exact-match 0.9196, dynamic 0.8376).
161
+ - **Trained on a single rubric shape.** All training rubrics have exactly 10 items. The
162
+ model handles 25-item rubrics correctly, but longer or structurally different rubrics are
163
+ untested.
164
+ - **Judgement quality is not where this model beats the 9B.** On the OOD suite the
165
+ [9B evaluator](../designcoder_evaluator_qwen3.5_9b_adamw_bs256_data37847_step100)
166
+ actually scores slightly higher on κ (0.694 vs 0.664). The 27B's real advantage is
167
+ **format robustness**: verdict-count match stays at 1.00 across rubric lengths 5→25,
168
+ where the 9B drops to 0.00 at length 5 and ~0.85 at length 25. Use this model when
169
+ rubric length varies; the 9B is the cheaper choice for fixed 10-item rubrics.
170
+ - Evidence is screenshot-only; the model cannot verify interaction behaviour, yet
171
+ `interaction_rubrics` asks it to describe interactions — treat that field as speculative.