|
Download README.md from AndreiRabau/diffusion-llm: direct link, hf CLI and curl.
- Browser
- Download file 7.8 kB
-
https://huggingface.co/AndreiRabau/diffusion-llm/resolve/main/README.md
- Command line
-
hf download hf://AndreiRabau/diffusion-llm/README.md
-
curl -L -o README.md https://huggingface.co/AndreiRabau/diffusion-llm/resolve/main/README.md
7.8 kB
| language: | |
| - en | |
| library_name: pytorch | |
| base_model: openai-community/gpt2 | |
| datasets: | |
| - Salesforce/wikitext | |
| tags: | |
| - diffusion-language-model | |
| - discrete-diffusion | |
| - text-denoising | |
| - gpt2 | |
| - research | |
| pipeline_tag: text-generation | |
| # Diffusion-LLM | |
| > **A small, experimental discrete diffusion language model built on GPT-2.** | |
| This repository contains checkpoints from an experiment in iterative text | |
| denoising. Instead of predicting only the next token from a left-to-right | |
| context, the model receives a corrupted sequence and predicts the original | |
| tokens across the whole sequence with bidirectional attention. | |
| The checkpoints are intended for research and reproducibility. They are not | |
| drop-in `transformers` causal language models and should not be treated as | |
| production-ready text generators. | |
| ## What is inside? | |
| The model starts from [`openai-community/gpt2`](https://huggingface.co/openai-community/gpt2) | |
| and adapts its Transformer backbone for discrete diffusion-style denoising: | |
| - GPT-2 embeddings and Transformer blocks are retained. | |
| - Causal attention is replaced with bidirectional attention. | |
| - A learned diffusion-timestep embedding is added to every token embedding. | |
| - The language-model head predicts the clean token at each position. | |
| - Denoising is performed by repeatedly updating corrupted positions. | |
| The training code adds two special tokens to the GPT-2 tokenizer: | |
| `<|pad|>` and `<|mask|>`. The vocabulary is resized before training, so the | |
| same tokenizer construction is required when loading a checkpoint. | |
| ## Checkpoints | |
| The uploaded weights follow this layout: | |
| | Path | Training corruption | Refinement steps during training | | |
| | --- | --- | ---: | | |
| | `default.pt` | Initial/reference checkpoint | — | | |
| | `1_step/similar_100.pt` | Semantically similar replacements | 1 | | |
| | `1_step/mask_100.pt` | Mask-token corruption | 1 | | |
| | `1_step/random_100.pt` | Random-token replacement | 1 | | |
| | `1_step/mix_30_30_40.pt` | 30 training epochs similar → 30 mask → 40 random | 1 | | |
| | `4_step/similar_100_4_steps.pt` | Semantically similar replacements | 4 | | |
| | `4_step/mask_100_4_steps.pt` | Mask-token corruption | 4 | | |
| | `4_step/random_100_4_steps.pt` | Random-token replacement | 4 | | |
| | `4_step/mix_30_30_40_4_steps.pt` | 30 training epochs similar → 30 mask → 40 random | 4 | | |
| All listed experimental runs use the same main setup: GPT-2, 100 diffusion | |
| timesteps, maximum sequence length 64, learning rate `1e-5`, and WikiText-2 | |
| (`Salesforce/wikitext`, `wikitext-2-raw-v1`). The 4-step runs use a rollout | |
| loss decay of `0.5`. | |
| ## Corruption strategies | |
| The training pipeline supports three discrete corruption operators: | |
| 1. **Similar** — replaces tokens with nearby tokens in the GPT-2 embedding | |
| space. This creates difficult, semantically close mistakes. | |
| 2. **Mask** — replaces tokens with `<|mask|>`, the classic masked-denoising | |
| condition. | |
| 3. **Random** — replaces tokens with random vocabulary entries. | |
| The mixed checkpoints use these methods as a sequential training curriculum: | |
| **30 epochs of similar-token corruption, 30 epochs of masking, and 40 epochs | |
| of random replacement**. Corruption probability is sampled between `0.01` and | |
| `0.95`; the similar-token operator uses 20 nearest neighbours. | |
| ## Quick start | |
| Clone this repository so that the model wrapper and corruption utilities are | |
| available alongside the downloaded checkpoint: | |
| ```bash | |
| git clone https://huggingface.co/AndreiRabau/diffusion-llm | |
| cd diffusion-llm | |
| pip install torch transformers | |
| ``` | |
| The following example loads the 4-step mixed checkpoint and denoises a | |
| sequence containing mask tokens: | |
| ```python | |
| from pathlib import Path | |
| import torch | |
| from transformers import AutoTokenizer | |
| from model_wrappers import GPT2DiffusionTransformer | |
| MODEL_NAME = "openai-community/gpt2" | |
| CHECKPOINT = Path("4_step/mix_30_30_40_4_steps.pt") | |
| tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME) | |
| tokenizer.add_special_tokens({ | |
| "pad_token": "<|pad|>", | |
| "mask_token": "<|mask|>", | |
| }) | |
| model = GPT2DiffusionTransformer.from_file_path( | |
| file_path=CHECKPOINT, | |
| model_name=MODEL_NAME, | |
| num_diffusion_steps=100, | |
| vocabulary_size=len(tokenizer), | |
| device="cpu", | |
| ).eval() | |
| text = "Diffusion models can refine [MASK] sequences over several steps." | |
| inputs = tokenizer(text, return_tensors="pt") | |
| # For a real masked input, create the sequence with tokenizer.mask_token_id. | |
| # Here we replace one token to keep the example self-contained. | |
| corrupted_ids = inputs["input_ids"].clone() | |
| corrupted_positions = torch.zeros_like(corrupted_ids, dtype=torch.bool) | |
| corrupted_ids[0, 4] = tokenizer.mask_token_id | |
| corrupted_positions[0, 4] = True | |
| reconstructed_ids = model.denoise( | |
| corrupted_ids=corrupted_ids, | |
| attention_mask=inputs["attention_mask"], | |
| corrupted_positions=corrupted_positions, | |
| num_iterations=4, | |
| ) | |
| print(tokenizer.decode(reconstructed_ids[0], skip_special_tokens=True)) | |
| ``` | |
| The `.pt` files contain a PyTorch `state_dict`. Loading them therefore | |
| requires the repository's `GPT2DiffusionTransformer` implementation and the | |
| same model configuration used during training (`num_diffusion_steps=100` and | |
| the resized tokenizer vocabulary). | |
| ## Training and evaluation details | |
| Training uses cross-entropy only on corrupted, non-padding positions. For | |
| multi-step training, each step feeds the model's predictions back into the | |
| next step, exposing the model to its own errors. At evaluation time, the | |
| denoiser uses confidence-guided refinement: the most confident corrupted | |
| positions are updated first, while less certain positions remain available for | |
| later iterations. This difference is intentional. | |
| The included experiments evaluate on the `test` split of | |
| [`cimec/lambada`](https://huggingface.co/datasets/cimec/lambada) with | |
| corruption rates from 25% to 95% and 2, 4, 5, 8, 10, 20, or 50 denoising | |
| iterations, depending on the run. Evaluation artifacts are stored in | |
| `results/eval/` in the source project. | |
| ## Intended use | |
| This release is useful for: | |
| - studying corruption policies for discrete diffusion language models; | |
| - comparing one-step and iterative denoising rollouts; | |
| - experimenting with confidence-based token refinement; | |
| - reproducing the accompanying small-scale WikiText-2 experiments. | |
| It is **not** intended for factual question answering, safety-critical | |
| generation, deployment, or direct comparison with large-scale diffusion LMs. | |
| The checkpoints inherit the limitations of GPT-2 and the small experimental | |
| training setup, including potential memorization, bias, and unstable outputs. | |
| ## Limitations and open questions | |
| - These are research checkpoints rather than a fully packaged inference | |
| pipeline. | |
| - The `.pt` format stores weights only; configuration and tokenizer files are | |
| reconstructed by the loader. | |
| - The model was trained on short sequences (maximum length 64), so longer | |
| contexts are outside the validated setup. | |
| - The corruption curriculum and the semantic-neighbour operator have not been | |
| established as generally superior to standard masked or random corruption. | |
| - Reported results should be interpreted as exploratory measurements, not as | |
| a leaderboard claim. | |
| ## Citation | |
| If you use these checkpoints or the training setup, please cite this | |
| repository and the upstream GPT-2 work: | |
| ```bibtex | |
| @misc{diffusion_llm_experiment, | |
| title = {Diffusion-LLM: Experimental GPT-2 Discrete Denoising Checkpoints}, | |
| author = {Andrei Rabau}, | |
| year = {2026}, | |
| url = {https://huggingface.co/AndreiRabau/diffusion-llm} | |
| } | |
| ``` | |
| The implementation and experiment configurations are available in the source | |
| repository, including [`model_wrappers/gpt2_diffusion_transformer_wrapper.py`](model_wrappers/gpt2_diffusion_transformer_wrapper.py), | |
| [`train_utils/train.py`](train_utils/train.py), and the files under | |
| [`experiments/`](experiments/). | |