Shuibai12138 commited on
Commit
535b369
·
verified ·
1 Parent(s): 0e52f8e

Model card: correct usage, base-model attribution, gradient implementation note, reference results

Browse files
Files changed (1) hide show
  1. README.md +24 -6
README.md CHANGED
@@ -33,23 +33,41 @@ This model is part of the code release of *Corrective Diffusion Language Models*
33
  | Frozen | `lm_head`, `embed_tokens` |
34
  | Seed | 42 |
35
  | Hardware | 4 x A100-PCIE-40GB, about 19 minutes |
36
- | Code | [zhangshuibai/CDLM](https://github.com/zhangshuibai/CDLM), trained at commit `5e52812`; tag `v1.0-corrective-training` contains the same training code. Command: `ARM=mdlm bash training/scripts/train_0.5b_opencodeinstruct.sh` |
37
 
38
  `training_config.yaml` in this repository is the fully resolved configuration the trainer saved for this run.
39
 
 
 
 
 
 
 
 
 
 
 
40
  ## Usage
41
 
42
- This is a masked diffusion language model with bidirectional attention. Loading it with `AutoModelForCausalLM` gives a causal Qwen2 model and wrong outputs. Use the evaluation pipeline of the [code repository](https://github.com/zhangshuibai/CDLM), which loads models whose name contains `open-dcoder` with the diffusion Qwen2 implementation (bidirectional attention, shifted logits):
43
 
44
  ```bash
45
- git clone https://github.com/zhangshuibai/CDLM && cd CDLM
46
- # see README.md for the evaluation environment and commands, then pass
47
- # --model_name Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct
 
 
48
  ```
49
 
 
 
 
 
 
 
50
  ## Licence and attribution
51
 
52
- Model weights: MIT. Trained on [nvidia/OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) (CC BY 4.0, NVIDIA). Base model: [fredzzp/open-dcoder-0.5B](https://huggingface.co/fredzzp/open-dcoder-0.5B).
53
 
54
  ## Citation
55
 
 
33
  | Frozen | `lm_head`, `embed_tokens` |
34
  | Seed | 42 |
35
  | Hardware | 4 x A100-PCIE-40GB, about 19 minutes |
36
+ | Code | [zhangshuibai/CDLM](https://github.com/zhangshuibai/CDLM), trained at commit `5e52812`; tags `v1.0-corrective-training` and `v1.0.1-corrective-training` contain the same training code. Command: `ARM=mdlm bash training/scripts/train_0.5b_opencodeinstruct.sh` |
37
 
38
  `training_config.yaml` in this repository is the fully resolved configuration the trainer saved for this run.
39
 
40
+ ## Reference results
41
+
42
+ Scored with the repository's CRB launcher (4 datasets x 3 error types at n_replace 1, macro over the 12 cells, confidence threshold 0.9). The 8 s execution timeout can flip one slow HumanEval+ program, which moves a Pass@1 by at most about 0.001.
43
+
44
+ | | Pass@1 T=1 | T=2 | T=3 | T=4 | confidence gap | Top-1 | Top-3 | Top-5 |
45
+ |---|---|---|---|---|---|---|---|---|
46
+ | CRB | 0.1401 | 0.2226 | 0.2352 | 0.2387 | 0.0987 | 0.1445 | 0.3575 | 0.5199 |
47
+
48
+ The matched CDLM model on the same data scores 0.2196 / 0.3018 / 0.3029 / 0.3070, gap 0.1919, Top-1 0.2517.
49
+
50
  ## Usage
51
 
52
+ This is a masked diffusion language model with bidirectional attention. Loading it with `AutoModelForCausalLM` gives a causal Qwen2 model and wrong outputs. Evaluate it with the launchers of the [code repository](https://github.com/zhangshuibai/CDLM) (environment: `evaluation/ENVIRONMENT.md`; pin the model revision listed in `evaluation/README.md`):
53
 
54
  ```bash
55
+ git clone --branch v1.0.1-corrective-training https://github.com/zhangshuibai/CDLM && cd CDLM
56
+ # CRB localisation and correction (n_replace 1: the 48 headline cells)
57
+ bash evaluation/crb/run_crb.sh Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct mdlm_oci --gpus 0 --nr 1
58
+ # from-scratch code generation (HumanEval, HumanEval+, MBPP, MBPP+)
59
+ bash evaluation/codegen/run_codegen_eval.sh Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct outputs/mdlm_oci 0
60
  ```
61
 
62
+ The launchers load the diffusion Qwen2 implementation because the repository id contains `open-dcoder`.
63
+
64
+ ## Implementation note on the gradient
65
+
66
+ The 0.5B training code computes the per-token cross-entropy terms with `LigerFusedLinearCrossEntropyLoss(reduction="none")` from liger-kernel 0.5.8. Its backward pass scales every token's gradient by the upstream gradient of the first token, so the per-token weights of the objective (the 1/|S| and 1/t factors and the noise-term weight) are not applied in the update: each micro-batch receives the first token's weight times the unweighted sum of the per-token gradients, a micro-batch whose first target token is unsupervised receives no gradient, and when the first target token is a replaced one the noise term adds gradient on every valid position, clean tokens included. Logged losses are correct. This model was trained with that code; see `training/README.md` ("Effective gradient") in the code repository.
67
+
68
  ## Licence and attribution
69
 
70
+ Model weights: MIT. Base model: [fredzzp/open-dcoder-0.5B](https://huggingface.co/fredzzp/open-dcoder-0.5B), licensed under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0); this model is a derivative (continued training) of it, and the base model's licence and notices apply to the parts derived from it. Trained on [nvidia/OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) (CC BY 4.0, NVIDIA).
71
 
72
  ## Citation
73