Model card: correct usage, base-model attribution, gradient implementation note, reference results
Browse files
README.md
CHANGED
|
@@ -33,23 +33,41 @@ This model is part of the code release of *Corrective Diffusion Language Models*
|
|
| 33 |
| Frozen | `lm_head`, `embed_tokens` |
|
| 34 |
| Seed | 42 |
|
| 35 |
| Hardware | 4 x A100-PCIE-40GB, about 19 minutes |
|
| 36 |
-
| Code | [zhangshuibai/CDLM](https://github.com/zhangshuibai/CDLM), trained at commit `5e52812`;
|
| 37 |
|
| 38 |
`training_config.yaml` in this repository is the fully resolved configuration the trainer saved for this run.
|
| 39 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
## Usage
|
| 41 |
|
| 42 |
-
This is a masked diffusion language model with bidirectional attention. Loading it with `AutoModelForCausalLM` gives a causal Qwen2 model and wrong outputs.
|
| 43 |
|
| 44 |
```bash
|
| 45 |
-
git clone https://github.com/zhangshuibai/CDLM && cd CDLM
|
| 46 |
-
#
|
| 47 |
-
|
|
|
|
|
|
|
| 48 |
```
|
| 49 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 50 |
## Licence and attribution
|
| 51 |
|
| 52 |
-
Model weights: MIT.
|
| 53 |
|
| 54 |
## Citation
|
| 55 |
|
|
|
|
| 33 |
| Frozen | `lm_head`, `embed_tokens` |
|
| 34 |
| Seed | 42 |
|
| 35 |
| Hardware | 4 x A100-PCIE-40GB, about 19 minutes |
|
| 36 |
+
| Code | [zhangshuibai/CDLM](https://github.com/zhangshuibai/CDLM), trained at commit `5e52812`; tags `v1.0-corrective-training` and `v1.0.1-corrective-training` contain the same training code. Command: `ARM=mdlm bash training/scripts/train_0.5b_opencodeinstruct.sh` |
|
| 37 |
|
| 38 |
`training_config.yaml` in this repository is the fully resolved configuration the trainer saved for this run.
|
| 39 |
|
| 40 |
+
## Reference results
|
| 41 |
+
|
| 42 |
+
Scored with the repository's CRB launcher (4 datasets x 3 error types at n_replace 1, macro over the 12 cells, confidence threshold 0.9). The 8 s execution timeout can flip one slow HumanEval+ program, which moves a Pass@1 by at most about 0.001.
|
| 43 |
+
|
| 44 |
+
| | Pass@1 T=1 | T=2 | T=3 | T=4 | confidence gap | Top-1 | Top-3 | Top-5 |
|
| 45 |
+
|---|---|---|---|---|---|---|---|---|
|
| 46 |
+
| CRB | 0.1401 | 0.2226 | 0.2352 | 0.2387 | 0.0987 | 0.1445 | 0.3575 | 0.5199 |
|
| 47 |
+
|
| 48 |
+
The matched CDLM model on the same data scores 0.2196 / 0.3018 / 0.3029 / 0.3070, gap 0.1919, Top-1 0.2517.
|
| 49 |
+
|
| 50 |
## Usage
|
| 51 |
|
| 52 |
+
This is a masked diffusion language model with bidirectional attention. Loading it with `AutoModelForCausalLM` gives a causal Qwen2 model and wrong outputs. Evaluate it with the launchers of the [code repository](https://github.com/zhangshuibai/CDLM) (environment: `evaluation/ENVIRONMENT.md`; pin the model revision listed in `evaluation/README.md`):
|
| 53 |
|
| 54 |
```bash
|
| 55 |
+
git clone --branch v1.0.1-corrective-training https://github.com/zhangshuibai/CDLM && cd CDLM
|
| 56 |
+
# CRB localisation and correction (n_replace 1: the 48 headline cells)
|
| 57 |
+
bash evaluation/crb/run_crb.sh Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct mdlm_oci --gpus 0 --nr 1
|
| 58 |
+
# from-scratch code generation (HumanEval, HumanEval+, MBPP, MBPP+)
|
| 59 |
+
bash evaluation/codegen/run_codegen_eval.sh Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct outputs/mdlm_oci 0
|
| 60 |
```
|
| 61 |
|
| 62 |
+
The launchers load the diffusion Qwen2 implementation because the repository id contains `open-dcoder`.
|
| 63 |
+
|
| 64 |
+
## Implementation note on the gradient
|
| 65 |
+
|
| 66 |
+
The 0.5B training code computes the per-token cross-entropy terms with `LigerFusedLinearCrossEntropyLoss(reduction="none")` from liger-kernel 0.5.8. Its backward pass scales every token's gradient by the upstream gradient of the first token, so the per-token weights of the objective (the 1/|S| and 1/t factors and the noise-term weight) are not applied in the update: each micro-batch receives the first token's weight times the unweighted sum of the per-token gradients, a micro-batch whose first target token is unsupervised receives no gradient, and when the first target token is a replaced one the noise term adds gradient on every valid position, clean tokens included. Logged losses are correct. This model was trained with that code; see `training/README.md` ("Effective gradient") in the code repository.
|
| 67 |
+
|
| 68 |
## Licence and attribution
|
| 69 |
|
| 70 |
+
Model weights: MIT. Base model: [fredzzp/open-dcoder-0.5B](https://huggingface.co/fredzzp/open-dcoder-0.5B), licensed under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0); this model is a derivative (continued training) of it, and the base model's licence and notices apply to the parts derived from it. Trained on [nvidia/OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) (CC BY 4.0, NVIDIA).
|
| 71 |
|
| 72 |
## Citation
|
| 73 |
|