xingxm commited on
Commit
5721771
·
verified ·
1 Parent(s): 3c6d824

Update model card: add 4 data41287 checkpoints, fix dataset revision, add bench-200 scores

Browse files
Files changed (1) hide show
  1. README.md +65 -9
README.md CHANGED
@@ -30,22 +30,69 @@ designcoder_{basemodel}_{size}_{optimizer}_bs{global_batch}[_{extra_axes}]_step{
30
  - `optimizer`: `muon` or `adamw`
31
  - `bs`: global batch size (`per_device × grad_accum × world_size`)
32
  - `extra_axes`: any hyper-parameter that deviates from the default recipe, e.g. `wd0.05`
33
- (weight decay, default 0.0) or `ep20` (epochs, default 2)
34
  - `step`: trainer `global_step` of the exported weights
35
 
 
 
 
 
 
 
 
 
 
 
36
  ## Checkpoints
37
 
38
- | Subfolder | Base model | Optimizer | LR | Global batch | Epochs | Weight decay | Step | Notes |
39
  |---|---|---|---|---|---|---|---|---|
40
- | `designcoder_qwen3.5_4b_muon_bs32_step1900` | Qwen3.5-4B | Muon | 1e-5 | 32 | 2 | 0.0 | 1900 | smallest release |
41
- | `designcoder_qwen3.5_9b_muon_bs16_step3800` | Qwen3.5-9B | Muon | 1e-5 | 16 | 2 | 0.0 | 3800 | optimizer ablation (Muon arm) |
42
- | `designcoder_qwen3.5_9b_adamw_bs16_step3800` | Qwen3.5-9B | AdamW | 2e-5 | 16 | 2 | 0.0 | 3800 | optimizer ablation (AdamW arm) |
43
- | `designcoder_qwen3.6_27b_adamw_bs32_step1900` | Qwen3.6-27B | AdamW | 1e-5 | 32 | 2 | 0.0 | 1900 | largest release |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
  ## Shared training setup
46
 
47
  - Objective: full-parameter supervised fine-tuning (no LoRA / adapters)
48
- - Dataset: `designcoder_sft_v2_train`, 41,287 ShareGPT-format records
49
  - Chat template: `qwen3_5` with thinking enabled
50
  - Context length: 32,768
51
  - Sequence packing: enabled, with neat packing (no cross-sample attention)
@@ -57,7 +104,7 @@ designcoder_{basemodel}_{size}_{optimizer}_bs{global_batch}[_{extra_axes}]_step{
57
  from transformers import AutoModelForCausalLM, AutoProcessor
58
 
59
  repo = "xingxm/DesignCoder"
60
- subfolder = "designcoder_qwen3.5_4b_muon_bs32_step1900"
61
 
62
  model = AutoModelForCausalLM.from_pretrained(repo, subfolder=subfolder, dtype="auto", device_map="auto")
63
  processor = AutoProcessor.from_pretrained(repo, subfolder=subfolder)
@@ -66,9 +113,18 @@ processor = AutoProcessor.from_pretrained(repo, subfolder=subfolder)
66
  To download a single checkpoint only:
67
 
68
  ```bash
69
- hf download xingxm/DesignCoder --include "designcoder_qwen3.5_4b_muon_bs32_step1900/*" --local-dir ./DesignCoder
70
  ```
71
 
 
 
 
 
 
 
 
 
 
72
  ## Provenance
73
 
74
  Each subfolder additionally ships `trainer_state.json` / `trainer_log.jsonl` (and
 
30
  - `optimizer`: `muon` or `adamw`
31
  - `bs`: global batch size (`per_device × grad_accum × world_size`)
32
  - `extra_axes`: any hyper-parameter that deviates from the default recipe, e.g. `wd0.05`
33
+ (weight decay, default 0.0), `ep20` (epochs, default 2), or `data41287` (dataset revision)
34
  - `step`: trainer `global_step` of the exported weights
35
 
36
+ ## Dataset revisions
37
+
38
+ Checkpoints in this repository come from two different dataset revisions. **Scores and loss
39
+ values are only comparable within the same revision.**
40
+
41
+ | Tag | Samples | Used by |
42
+ |---|---:|---|
43
+ | *(untagged)* `data37865` | 37,865 | `*_step1900`, `*_step3800` |
44
+ | `data41287` | 41,287 | `*_data41287_step200`, `*_data41287_step400` |
45
+
46
  ## Checkpoints
47
 
48
+ | Subfolder | Base model | Optimizer | LR | Global batch | Dataset | Step | bench-200 | Notes |
49
  |---|---|---|---|---|---|---|---|---|
50
+ | `designcoder_qwen3.5_4b_muon_bs32_step1900` | Qwen3.5-4B | Muon | 1e-5 | 32 | 37,865 | 1900 | – | smallest of the first release |
51
+ | `designcoder_qwen3.5_9b_muon_bs16_step3800` | Qwen3.5-9B | Muon | 1e-5 | 16 | 37,865 | 3800 | – | optimizer ablation (Muon arm) |
52
+ | `designcoder_qwen3.5_9b_adamw_bs16_step3800` | Qwen3.5-9B | AdamW | 2e-5 | 16 | 37,865 | 3800 | – | optimizer ablation (AdamW arm) |
53
+ | `designcoder_qwen3.6_27b_adamw_bs32_step1900` | Qwen3.6-27B | AdamW | 1e-5 | 32 | 37,865 | 1900 | – | largest of the first release |
54
+ | `designcoder_qwen3.5_4b_adamw_bs256_data41287_step200` | Qwen3.5-4B | AdamW | 2e-5 | 256 | 41,287 | 200 | 84.22 | best 4B / AdamW |
55
+ | `designcoder_qwen3.5_4b_muon_bs256_data41287_step200` | Qwen3.5-4B | Muon | 2e-5 | 256 | 41,287 | 200 | 83.36 | best 4B / Muon; degrades less late in training |
56
+ | `designcoder_qwen3.5_9b_adamw_bs256_data41287_step200` | Qwen3.5-9B | AdamW | 2e-5 | 256 | 41,287 | 200 | **84.40** | only checkpoint scored on all 200 cases |
57
+ | `designcoder_qwen3.8_27b_adamw_bs128_data41287_step400` | Qwen3.8-27B | AdamW | 1e-5 | 128 | 41,287 | 400 | **91.19** | strongest checkpoint in the collection |
58
+
59
+ ## Benchmark
60
+
61
+ `bench-200` is the frozen 200-case DesignCoder benchmark (100 Track A landing, 40 Track A
62
+ dashboard, 30 Track B landing, 30 Track B dashboard). Every rubric item is a binary
63
+ screenshot check scored by a vision judge over full-page renders; the reported number is the
64
+ unweighted mean of Prompt Fit and the six rubric dimensions (Alignment, Layout, Typography,
65
+ Components, Assets, Aesthetics).
66
+
67
+ The 9B score is a full 200-case run. The 4B and 27B scores come from an 8-case subset
68
+ reweighted to the benchmark's real landing/dashboard split, so they are indicative rather
69
+ than final — and the subset was sampled around the 9B mid-range, which understates 9B
70
+ relative to 4B. Use the 9B full-run number (84.40) when comparing against 27B (91.19).
71
+
72
+ ### Checkpoint selection
73
+
74
+ The `data41287` checkpoints were selected by **running the benchmark**, not by taking the
75
+ lowest training loss. In all four runs the best checkpoint sits at roughly 75% of training,
76
+ and loss kept improving while benchmark scores fell:
77
+
78
+ | Run | Step | Train loss | bench-200 |
79
+ |---|---:|---:|---:|
80
+ | 4B AdamW | 200 | 0.2696 | **84.22** |
81
+ | 4B AdamW | 266 | 0.2682 | 68.35 |
82
+ | 4B Muon | 200 | 0.3339 | **83.36** |
83
+ | 4B Muon | 266 | 0.3349 | 81.27 |
84
+ | 9B AdamW | 200 | 0.2518 | **84.40** |
85
+ | 9B AdamW | 266 | 0.2504 | lowest of the three |
86
+ | 27B AdamW | 400 | 0.2067 | **91.19** |
87
+ | 27B AdamW | 530 | 0.2059 | 86.37 |
88
+
89
+ The 4B AdamW pair is the clearest example: loss improved from 0.2696 to 0.2682 while the
90
+ score collapsed from 84.22 to 68.35. **Do not pick checkpoints from this family by loss.**
91
 
92
  ## Shared training setup
93
 
94
  - Objective: full-parameter supervised fine-tuning (no LoRA / adapters)
95
+ - Dataset: `designcoder_sft_v2_train` in ShareGPT format (see revision table above)
96
  - Chat template: `qwen3_5` with thinking enabled
97
  - Context length: 32,768
98
  - Sequence packing: enabled, with neat packing (no cross-sample attention)
 
104
  from transformers import AutoModelForCausalLM, AutoProcessor
105
 
106
  repo = "xingxm/DesignCoder"
107
+ subfolder = "designcoder_qwen3.8_27b_adamw_bs128_data41287_step400"
108
 
109
  model = AutoModelForCausalLM.from_pretrained(repo, subfolder=subfolder, dtype="auto", device_map="auto")
110
  processor = AutoProcessor.from_pretrained(repo, subfolder=subfolder)
 
113
  To download a single checkpoint only:
114
 
115
  ```bash
116
+ hf download xingxm/DesignCoder --include "designcoder_qwen3.8_27b_adamw_bs128_data41287_step400/*" --local-dir ./DesignCoder
117
  ```
118
 
119
+ ### Inference contract
120
+
121
+ These models are trained as tool-using agents, not single-turn generators. A case runs
122
+ `design_search` → (`websearch`, landing only) → a final answer containing exactly three code
123
+ blocks in the order `html`, `css`, `js`. Reproduce the system prompts and tool observation
124
+ format from `examples/designcoder/runtime/infer_designcoder.py`; prompting with a bare
125
+ instruction and no tool turns does not match the training distribution and will score far
126
+ below the numbers above.
127
+
128
  ## Provenance
129
 
130
  Each subfolder additionally ships `trainer_state.json` / `trainer_log.jsonl` (and