xingxm commited on
Commit
b8f37a9
·
verified ·
1 Parent(s): f9e65a9

Add model card: naming convention and checkpoint index

Browse files
Files changed (1) hide show
  1. README.md +75 -0
README.md CHANGED
@@ -1,3 +1,78 @@
1
  ---
2
  license: mit
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ library_name: transformers
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - designcoder
7
+ - ui-generation
8
+ - front-end
9
+ - html
10
+ - css
11
+ - javascript
12
+ - code-generation
13
+ - full-sft
14
  ---
15
+
16
+ # DesignCoder
17
+
18
+ Checkpoint collection for **DesignCoder**, a family of full-parameter SFT models for UI design
19
+ research and end-to-end HTML/CSS/JavaScript implementation.
20
+
21
+ Each subfolder in this repository is a self-contained, directly loadable checkpoint.
22
+
23
+ ## Naming convention
24
+
25
+ ```
26
+ designcoder_{basemodel}_{size}_{optimizer}_bs{global_batch}[_ep{epochs}]_step{global_step}[_r{rerun}]
27
+ ```
28
+
29
+ - `basemodel` / `size`: base model family and parameter scale
30
+ - `optimizer`: `muon` or `adamw`
31
+ - `bs`: global batch size (`per_device × grad_accum × world_size`)
32
+ - `ep`: only present when epochs differ from the default 2
33
+ - `step`: trainer `global_step` of the exported weights
34
+ - `r`: rerun index, only present for repeated runs of an identical configuration
35
+
36
+ ## Checkpoints
37
+
38
+ | Subfolder | Base model | Optimizer | LR | Global batch | Epochs | Step | Notes |
39
+ |---|---|---|---|---|---|---|---|
40
+ | `designcoder_qwen3.5_4b_muon_bs32_step1900` | Qwen3.5-4B | Muon | 1e-5 | 32 | 2 | 1900 | smallest release |
41
+ | `designcoder_qwen3.5_9b_muon_bs16_step3800` | Qwen3.5-9B | Muon | 1e-5 | 16 | 2 | 3800 | optimizer ablation (Muon arm) |
42
+ | `designcoder_qwen3.5_9b_adamw_bs16_step3800` | Qwen3.5-9B | AdamW | 2e-5 | 16 | 2 | 3800 | optimizer ablation (AdamW arm) |
43
+ | `designcoder_qwen3.5_9b_adamw_bs16_ep20_step38000` | Qwen3.5-9B | AdamW | 2e-5 | 16 | 20 | 38000 | epoch-scaling ablation |
44
+ | `designcoder_qwen3.6_27b_adamw_bs32_step1900` | Qwen3.6-27B | AdamW | 1e-5 | 32 | 2 | 1900 | largest release |
45
+ | `designcoder_qwen3.6_27b_adamw_bs32_step1900_r2` | Qwen3.6-27B | AdamW | 1e-5 | 32 | 2 | 1900 | rerun of the 27B configuration |
46
+
47
+ ## Shared training setup
48
+
49
+ - Objective: full-parameter supervised fine-tuning (no LoRA / adapters)
50
+ - Dataset: `designcoder_sft_v2_train`, 41,287 ShareGPT-format records
51
+ - Chat template: `qwen3_5` with thinking enabled
52
+ - Context length: 32,768
53
+ - Sequence packing: enabled, with neat packing (no cross-sample attention)
54
+ - LR schedule: cosine, warmup ratio 0.1
55
+
56
+ ## Usage
57
+
58
+ ```python
59
+ from transformers import AutoModelForCausalLM, AutoProcessor
60
+
61
+ repo = "xingxm/DesignCoder"
62
+ subfolder = "designcoder_qwen3.5_4b_muon_bs32_step1900"
63
+
64
+ model = AutoModelForCausalLM.from_pretrained(repo, subfolder=subfolder, dtype="auto", device_map="auto")
65
+ processor = AutoProcessor.from_pretrained(repo, subfolder=subfolder)
66
+ ```
67
+
68
+ To download a single checkpoint only:
69
+
70
+ ```bash
71
+ hf download xingxm/DesignCoder --include "designcoder_qwen3.5_4b_muon_bs32_step1900/*" --local-dir ./DesignCoder
72
+ ```
73
+
74
+ ## Provenance
75
+
76
+ Each subfolder additionally ships `trainer_state.json` / `trainer_log.jsonl` (and
77
+ `training_loss.png` where available) so that the loss curve and exact step schedule of the run
78
+ can be recovered from the checkpoint itself.