Instructions to use Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
LLaDA-8B-CDLM-LoRA-OpenCodeInstruct
LoRA adapter for GSAI-ML/LLaDA-8B-Base trained for 2,000 steps on nvidia/OpenCodeInstruct with the CDLM corrective objective. It is one of a matched pair; the other arm is Shuibai12138/LLaDA-8B-MDLM-LoRA-OpenCodeInstruct. The two arms share the data, steps, schedule, seed, batch and LoRA configuration and differ only in mixture_prob and noise_token_wt.
This adapter is part of the code release of Corrective Diffusion Language Models (NeurIPS 2026). It is not a model from the paper. The paper's LLaDA-8B adapters (Shuibai12138/LLaDA-8B-CDLM-LoRA) were trained on Nemotron-SFT-Code, which is gated and licensed for internal training only. This pair uses the same trainer and hyperparameters on a public, ungated corpus, so that the 8B experiment can be reproduced and compared by anyone. Results on it are not the paper's results.
Code: zhangshuibai/CDLM.
Training
| Base model | GSAI-ML/LLaDA-8B-Base, revision 0f2787f2d87e (8.02B parameters) |
| Objective | mixture_prob=0.1, noise_token_wt=0.1, clean_token_wt=0 (CDLM: absorbing corruption plus uniform replacement of 10% of the still-visible tokens, supervised with weight 0.1). The absorbing loss is LLaDA's: t ~ U(0,1) per sequence, mask probability (1-1e-3)t + 1e-3, cross-entropy on masked tokens divided by the mask probability |
| LoRA | r = 64, alpha = 128, dropout 0.05 on q_proj, k_proj, v_proj, attn_out, ff_proj, up_proj, ff_out inside transformer.blocks.*; 167.77M trainable parameters (2.05%); embeddings and output layer frozen |
| Data | nvidia/OpenCodeInstruct, revision 8f3ba5bafe4d, all 50 shards, each row rendered as "input: " + input + " output: " + output (the text format of the paper's corpus) by training/data_prep/prepare_opencodeinstruct.py, no filtering; the shards are shuffled with seed 42 and assigned round-robin to the 4 ranks, and each document is tokenised, given an EOS token and split into chunks of at most 4,096 tokens. As in the paper's 8B runs, a 2,000-step run reads the first 3,000 documents of one shard per rank: 12,000 documents, 6.58M tokens, 548 tokens per document on average (the paper's runs: 12,000 documents, 3.55M tokens). OpenCodeInstruct's shards are not mixed by category, so these rows are not a sample of the whole corpus: they are 9,000 generic/evol-instruct and 3,000 algorithmic/self-instruct rows, while the corpus is 38.9% generic/evol-instruct, 33.4% generic/self-instruct, 19.6% algorithmic/self-instruct and 8.1% algorithmic/evol-instruct. data_shards.tsv lists the rank and position at which each shard is read |
| Steps | 2,000 |
| Optimiser | AdamW, betas (0.9, 0.95), weight decay 0.01, grad clip 1.0; peak lr 3e-4, 200 warmup steps, cosine over a 20,000-step horizon (lr at step 2,000 = 2.94e-4) |
| Batch | global 12 sequences of at most 4,096 tokens (micro 1 x grad-accum 3 x 4 GPUs); documents are not packed |
| Precision | bf16 base weights, fp32 LoRA parameters and optimiser states, activation checkpointing |
| Seed | 42 |
| Hardware | 4 x A100-PCIE-40GB, 56.1 min |
| Code | bash training/llada8b_lora/run_train_opencodeinstruct.sh cdlm at commit 96906ed; tag v1.0.2-corrective-training has the same trainer and run_train.sh, and its launcher passes them the same data and arguments (later commits added comments, a shards-per-rank check, rank and position columns in data_shards.tsv and stricter link handling) |
The trainer's DataLoader runs two workers over an unsharded IterableDataset, so each chunk is used twice in a
row with independent noise; 2,000 steps see 12,000 distinct chunks. This is how every LLaDA-8B adapter of the
release was trained. The loss is computed with PyTorch cross-entropy, so the liger-kernel gradient issue of the
0.5B trainer does not apply. Training is not bitwise deterministic on GPUs: a re-run matches the step-1 loss
exactly, but by step 20 its loss differs by 5e-4 to 7e-3 relative in the re-runs we compared, and the difference
grows after that (the Nemotron CDLM 200-step run differs from the 2,000-step run by 0.7% at step 20 and 1.9% at
step 200). Use this adapter, not a retrained one, to reproduce the numbers exactly.
train_config.json and train_log.json are the configuration and loss log the trainer wrote for this run.
Results
Localisation is measured on CRB HumanEval with n_replace=1 and all three error types (n = 541). Confidence is
the probability the model assigns to the token currently at each position, which is the quantity the
self_conf-remask:vanilla sampler uses; the gap is mean confidence on clean positions minus mean confidence on
the injected error positions, and Top-K is the fraction of samples with an error among the K least confident
positions. Pass@1 (%) is CRB repair at T=1 and T=4 refinement steps with confidence threshold 0.9, remove_all,
temperature 0 and batch size 1, macro-averaged over {HumanEval, HumanEval+, MBPP, MBPP+} x {operator, var,
literal} at n_replace=1.
Raw vs. de-fenced. Fine-tuning on either corpus teaches the model markdown code fences. The HumanEval+ buggy bodies end with a spare blank line. Under remove_all refinement CDLM-OCI fills it with a fence in 99.6% of HumanEval+ completions at T=1 (99.0% at T=4) and MDLM-OCI in 10.4% (27.8%), and a program with a fence fails to parse. defence.py in the code repository truncates each completion at the first fence and strips trailing whitespace, with the same rule for every model; both numbers are given.
| model | confidence gap | Top-1 hit (%) | Top-3 hit (%) | Pass@1 T=1, de-fenced | Pass@1 T=1, raw | Pass@1 T=4, de-fenced | Pass@1 T=4, raw |
|---|---|---|---|---|---|---|---|
| LLaDA-8B-Base (no fine-tuning) | 0.265 | 16.5 | 54.7 | 55.1 | 55.1 | 59.5 | 59.6 |
| + MDLM LoRA, OpenCodeInstruct | 0.338 | 10.2 | 73.0 | 57.6 | 54.6 | 59.0 | 52.1 |
| + CDLM LoRA, OpenCodeInstruct | 0.835 | 59.3 | 90.8 | 74.1 | 51.8 | 72.6 | 50.6 |
| + MDLM LoRA, Nemotron-SFT-Code (paper) | 0.402 | 15.2 | 73.4 | 60.0 | 47.5 | 60.9 | 46.0 |
| + CDLM LoRA, Nemotron-SFT-Code (paper) | 0.777 | 60.8 | 92.8 | 72.5 | 53.0 | 62.3 | 51.8 |
Pass@1 differences below about 0.5 points are within the evaluation noise (execution timeouts). The base-model row is from the same evaluation run as the two OpenCodeInstruct rows; the Nemotron rows are the paper's.
Paired comparison. Over the 3,184 test-failing records, CDLM-OCI repairs more programs than MDLM-OCI: de-fenced +17.0 points (p = 2e-59) at T=1 and +14.5 points (p = 4e-44) at T=4, raw +4.2 points (p = 2e-04) and +4.6 points (p = 2e-05) (pooled difference, exact McNemar test). The raw macro average nevertheless puts CDLM-OCI below MDLM-OCI, because 99.6% of its HumanEval+ completions contain a fence at T=1, so its three raw HumanEval+ cells are close to zero and weigh a quarter of the average. Without HumanEval+, the raw macro Pass@1 at T=1 is 69.1 (CDLM-OCI), 53.8 (MDLM-OCI) and 53.3 (base).
Usage
The adapter is a PEFT LoRA adapter for the bidirectional masked diffusion model LLaDA-8B-Base. Merge it into the base model before use:
import torch
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
base = AutoModel.from_pretrained("GSAI-ML/LLaDA-8B-Base", revision="0f2787f2d87eac5eed8a087d5ecd24277e6255b2",
trust_remote_code=True, dtype=torch.bfloat16)
tok = AutoTokenizer.from_pretrained("GSAI-ML/LLaDA-8B-Base", revision="0f2787f2d87eac5eed8a087d5ecd24277e6255b2", trust_remote_code=True)
model = PeftModel.from_pretrained(base, "Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct").merge_and_unload().eval()
To reproduce the numbers above, use the code repository (zhangshuibai/CDLM,
training/llada8b_lora/); its scripts accept a Hub adapter id, optionally with a subfolder and a revision
(<owner>/<name>[/<subfolder>][@<revision>]). The commands below pin the revision that holds these weights.
git clone --branch v1.0.2-corrective-training https://github.com/zhangshuibai/CDLM && cd CDLM
pip install -r training/llada8b_lora/requirements.txt
bash training/llada8b_lora/fetch_crb_inputs.sh # the 24 CRB input files of this experiment
# localisation (confidence gap and Top-K hit rate)
python training/llada8b_lora/eval_confidence.py --label cdlm_oci --adapter Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct@d7a86e388333a3b5cdece3ed32cba388c1127175 --out outputs/conf_cdlm_oci.json
# one CRB repair cell (HumanEval, operator errors, n_replace 1, T=1)
bash training/llada8b_lora/run_crb_cell.sh 0 cdlm_oci Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct@d7a86e388333a3b5cdece3ed32cba388c1127175 human-eval operator 1 2
# all cells of the base model and both arms, then the Pass@1 table
ARMS="base:NONE mdlm_oci:Shuibai12138/LLaDA-8B-MDLM-LoRA-OpenCodeInstruct@47ba0fb36df497c944df2e3b93d2d878a198d701 cdlm_oci:Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct@d7a86e388333a3b5cdece3ed32cba388c1127175" \
GPUS="0 1 2 3" bash training/llada8b_lora/run_crb_parallel.sh 1
python training/llada8b_lora/aggregate_crb.py --n_replace 1 --labels base mdlm_oci cdlm_oci
Licence and attribution
Adapter: MIT. Base model: GSAI-ML/LLaDA-8B-Base (MIT). Trained on nvidia/OpenCodeInstruct (CC BY 4.0, NVIDIA).
Citation
@inproceedings{zhang2026corrective,
title = {Corrective Diffusion Language Models},
author = {Zhang, Shuibai and Peng, Fred Zhangzhi and Zhang, Yiheng and Pan, Jin and Chrysos, Grigorios G.},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}
- Downloads last month
- -
Model tree for Shuibai12138/LLaDA-8B-CDLM-LoRA-OpenCodeInstruct
Base model
GSAI-ML/LLaDA-8B-Base