Open-Dcoder-0.5B-CDLM-OpenCodeInstruct

Open-dCoder-0.5B continued for 2,000 steps on nvidia/OpenCodeInstruct with the CDLM corrective objective: absorbing (mask) corruption plus uniform replacement of 10% of the still-visible tokens, with a cross-entropy term on the replaced positions (weight 0.1). It is one of a matched pair; Shuibai12138/Open-Dcoder-0.5B-MDLM-OpenCodeInstruct is its matched MDLM control (absorbing-only objective, otherwise identical).

This model is part of the code release of Corrective Diffusion Language Models (NeurIPS 2026). It is not a model from the paper. The paper's 0.5B models (Shuibai12138/CDLM-0.5B) were trained on Nemotron-SFT-Code, which is gated and licensed for internal training only. This pair uses the same code and hyperparameters on a public, ungated corpus so that the recipe can be reproduced and compared by anyone. Results on it are not the paper's results.

Training

Initialisation fredzzp/open-dcoder-0.5B (revision d0d86d5b9996)
Objective mixture_prob=0.1, noise_token_wt=0.1, clean_token_wt=0.0
Data nvidia/OpenCodeInstruct, revision 8f3ba5bafe4d, all 50 shards in sorted order, each row rendered as "input: " + input + " output: " + output (the text format of the paper's corpus), no filtering
Steps 2,000 (step 2,000 of a 20,345,053-step schedule, the same truncated long-horizon schedule as the paper's CDLM-0.5B)
Optimiser AdamW, peak lr 3e-4, cosine, 20,345 warmup steps (lr at step 2,000 = 2.95e-5), weight decay 0.01, grad clip 1.0, bf16
Batch global 12 sequences x 4,096 packed tokens (micro 3 x 4 GPUs)
Frozen lm_head, embed_tokens
Seed 42
Hardware 4 x A100-PCIE-40GB, about 19 minutes
Code zhangshuibai/CDLM, trained at commit 5e52812; tags v1.0-corrective-training and v1.0.1-corrective-training contain the same training code. Command: ARM=cdlm bash training/scripts/train_0.5b_opencodeinstruct.sh

training_config.yaml in this repository is the fully resolved configuration the trainer saved for this run.

Reference results

Scored with the repository's launchers (CRB: 4 datasets x 3 error types at n_replace 1, macro over the 12 cells, confidence threshold 0.9; code generation: n=10 samples, vanilla decoding). The 8 s execution timeout can flip one slow HumanEval+ program, which moves a CRB Pass@1 by at most about 0.001.

Pass@1 T=1 T=2 T=3 T=4 confidence gap Top-1 Top-3 Top-5
CRB 0.2196 0.3018 0.3029 0.3070 0.1919 0.2517 0.4883 0.6495

HumanEval (vanilla decoding): pass@1 0.2165, pass@10 0.4146.

The matched MDLM control on the same data scores 0.1401 / 0.2226 / 0.2352 / 0.2387 (Pass@1 T=1..4), gap 0.0987, Top-1 0.1445 on the same CRB cells.

Usage

This is a masked diffusion language model with bidirectional attention. Loading it with AutoModelForCausalLM gives a causal Qwen2 model and wrong outputs. Evaluate it with the launchers of the code repository (environment: evaluation/ENVIRONMENT.md; pin the model revision listed in evaluation/README.md):

git clone --branch v1.0.1-corrective-training https://github.com/zhangshuibai/CDLM && cd CDLM
# CRB localisation and correction (n_replace 1: the 48 headline cells)
bash evaluation/crb/run_crb.sh Shuibai12138/Open-Dcoder-0.5B-CDLM-OpenCodeInstruct cdlm_oci --gpus 0 --nr 1
# from-scratch code generation (HumanEval, HumanEval+, MBPP, MBPP+)
bash evaluation/codegen/run_codegen_eval.sh Shuibai12138/Open-Dcoder-0.5B-CDLM-OpenCodeInstruct outputs/cdlm_oci 0

The launchers load the diffusion Qwen2 implementation because the repository id contains open-dcoder.

Implementation note on the gradient

The 0.5B training code computes the per-token cross-entropy terms with LigerFusedLinearCrossEntropyLoss(reduction="none") from liger-kernel 0.5.8. Its backward pass scales every token's gradient by the upstream gradient of the first token, so the per-token weights of the objective (the 1/|S| and 1/t factors and the noise-term weight) are not applied in the update: each micro-batch receives the first token's weight times the unweighted sum of the per-token gradients, a micro-batch whose first target token is unsupervised receives no gradient, and when the first target token is a replaced one the noise term adds gradient on every valid position, clean tokens included. Logged losses are correct. This model was trained with that code; see training/README.md ("Effective gradient") in the code repository.

Licence and attribution

Model weights: MIT. Base model: fredzzp/open-dcoder-0.5B, licensed under the Apache License 2.0; this model is a derivative (continued training) of it, and the base model's licence and notices apply to the parts derived from it. Trained on nvidia/OpenCodeInstruct (CC BY 4.0, NVIDIA).

Citation

@inproceedings{zhang2026corrective,
  title     = {Corrective Diffusion Language Models},
  author    = {Zhang, Shuibai and Peng, Fred Zhangzhi and Zhang, Yiheng and Pan, Jin and Chrysos, Grigorios G.},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}
Downloads last month
577
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shuibai12138/Open-Dcoder-0.5B-CDLM-OpenCodeInstruct

Finetuned
(5)
this model

Dataset used to train Shuibai12138/Open-Dcoder-0.5B-CDLM-OpenCodeInstruct

Collection including Shuibai12138/Open-Dcoder-0.5B-CDLM-OpenCodeInstruct