Model card: correct usage, base-model attribution, gradient implementation note, reference results
535b369 verified |
Download README.md from Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct: direct link, hf CLI and curl.
- Browser
- Download file 5.68 kB
-
https://huggingface.co/Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct/resolve/main/README.md
- Command line
-
hf download hf://Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct/README.md
-
curl -L -o README.md https://huggingface.co/Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct/resolve/main/README.md
5.68 kB
| license: mit | |
| base_model: fredzzp/open-dcoder-0.5B | |
| datasets: | |
| - nvidia/OpenCodeInstruct | |
| language: | |
| - code | |
| pipeline_tag: text-generation | |
| tags: | |
| - masked-diffusion | |
| - diffusion-language-model | |
| - code | |
| - code-correction | |
| - cdlm | |
| # Open-Dcoder-0.5B-MDLM-OpenCodeInstruct | |
| Open-dCoder-0.5B continued for 2,000 steps on [nvidia/OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) with the standard absorbing-only masked-diffusion objective (mixture_prob 0, noise_token_wt 0), i.e. the matched MDLM control for the CDLM objective. It is one of a matched pair; Shuibai12138/Open-Dcoder-0.5B-CDLM-OpenCodeInstruct is the CDLM model trained identically except for the objective. | |
| This model is part of the code release of *Corrective Diffusion Language Models* (NeurIPS 2026). **It is not a model from the paper.** The paper's 0.5B models ([Shuibai12138/Open-Dcoder-0.5B-baseline-mdm-step2000](https://huggingface.co/Shuibai12138/Open-Dcoder-0.5B-baseline-mdm-step2000)) were trained on Nemotron-SFT-Code, which is gated and licensed for internal training only. This pair uses the same code and hyperparameters on a public, ungated corpus so that the recipe can be reproduced and compared by anyone. Results on it are not the paper's results. | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Initialisation | [fredzzp/open-dcoder-0.5B](https://huggingface.co/fredzzp/open-dcoder-0.5B) (revision `d0d86d5b9996`) | | |
| | Objective | `mixture_prob=0.0`, `noise_token_wt=0.0`, `clean_token_wt=0.0` | | |
| | Data | nvidia/OpenCodeInstruct, revision `8f3ba5bafe4d`, all 50 shards in sorted order, each row rendered as `"input: " + input + " output: " + output` (the text format of the paper's corpus), no filtering | | |
| | Steps | 2,000 (step 2,000 of a 20,345,053-step schedule, the same truncated long-horizon schedule as the paper's CDLM-0.5B) | | |
| | Optimiser | AdamW, peak lr 3e-4, cosine, 20,345 warmup steps (lr at step 2,000 = 2.95e-5), weight decay 0.01, grad clip 1.0, bf16 | | |
| | Batch | global 12 sequences x 4,096 packed tokens (micro 3 x 4 GPUs) | | |
| | Frozen | `lm_head`, `embed_tokens` | | |
| | Seed | 42 | | |
| | Hardware | 4 x A100-PCIE-40GB, about 19 minutes | | |
| | Code | [zhangshuibai/CDLM](https://github.com/zhangshuibai/CDLM), trained at commit `5e52812`; tags `v1.0-corrective-training` and `v1.0.1-corrective-training` contain the same training code. Command: `ARM=mdlm bash training/scripts/train_0.5b_opencodeinstruct.sh` | | |
| `training_config.yaml` in this repository is the fully resolved configuration the trainer saved for this run. | |
| ## Reference results | |
| Scored with the repository's CRB launcher (4 datasets x 3 error types at n_replace 1, macro over the 12 cells, confidence threshold 0.9). The 8 s execution timeout can flip one slow HumanEval+ program, which moves a Pass@1 by at most about 0.001. | |
| | | Pass@1 T=1 | T=2 | T=3 | T=4 | confidence gap | Top-1 | Top-3 | Top-5 | | |
| |---|---|---|---|---|---|---|---|---| | |
| | CRB | 0.1401 | 0.2226 | 0.2352 | 0.2387 | 0.0987 | 0.1445 | 0.3575 | 0.5199 | | |
| The matched CDLM model on the same data scores 0.2196 / 0.3018 / 0.3029 / 0.3070, gap 0.1919, Top-1 0.2517. | |
| ## Usage | |
| This is a masked diffusion language model with bidirectional attention. Loading it with `AutoModelForCausalLM` gives a causal Qwen2 model and wrong outputs. Evaluate it with the launchers of the [code repository](https://github.com/zhangshuibai/CDLM) (environment: `evaluation/ENVIRONMENT.md`; pin the model revision listed in `evaluation/README.md`): | |
| ```bash | |
| git clone --branch v1.0.1-corrective-training https://github.com/zhangshuibai/CDLM && cd CDLM | |
| # CRB localisation and correction (n_replace 1: the 48 headline cells) | |
| bash evaluation/crb/run_crb.sh Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct mdlm_oci --gpus 0 --nr 1 | |
| # from-scratch code generation (HumanEval, HumanEval+, MBPP, MBPP+) | |
| bash evaluation/codegen/run_codegen_eval.sh Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct outputs/mdlm_oci 0 | |
| ``` | |
| The launchers load the diffusion Qwen2 implementation because the repository id contains `open-dcoder`. | |
| ## Implementation note on the gradient | |
| The 0.5B training code computes the per-token cross-entropy terms with `LigerFusedLinearCrossEntropyLoss(reduction="none")` from liger-kernel 0.5.8. Its backward pass scales every token's gradient by the upstream gradient of the first token, so the per-token weights of the objective (the 1/|S| and 1/t factors and the noise-term weight) are not applied in the update: each micro-batch receives the first token's weight times the unweighted sum of the per-token gradients, a micro-batch whose first target token is unsupervised receives no gradient, and when the first target token is a replaced one the noise term adds gradient on every valid position, clean tokens included. Logged losses are correct. This model was trained with that code; see `training/README.md` ("Effective gradient") in the code repository. | |
| ## Licence and attribution | |
| Model weights: MIT. Base model: [fredzzp/open-dcoder-0.5B](https://huggingface.co/fredzzp/open-dcoder-0.5B), licensed under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0); this model is a derivative (continued training) of it, and the base model's licence and notices apply to the parts derived from it. Trained on [nvidia/OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) (CC BY 4.0, NVIDIA). | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{zhang2026corrective, | |
| title = {Corrective Diffusion Language Models}, | |
| author = {Zhang, Shuibai and Peng, Fred Zhangzhi and Zhang, Yiheng and Pan, Jin and Chrysos, Grigorios G.}, | |
| booktitle = {Advances in Neural Information Processing Systems}, | |
| year = {2026} | |
| } | |
| ``` | |