Add links to paper and project page
#1
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -2,16 +2,18 @@
|
|
| 2 |
license: cc-by-nc-4.0
|
| 3 |
pipeline_tag: text-generation
|
| 4 |
tags:
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
---
|
| 12 |
|
| 13 |
# ALoDLM: Adaptively Looped Diffusion Language Models
|
| 14 |
|
|
|
|
|
|
|
| 15 |
ALoDLM is a family of diffusion language models developed at Amazon, available as **ALoDLM-1.7B** and **ALoDLM-8B** and initialized from the corresponding Qwen3 backbones. The models generate text and code by repeatedly applying shared transformer layers, preserving unresolved tokens' latent states, and adaptively allocating computation across token positions.
|
| 16 |
|
| 17 |
Training data includes mathematical problems and worked solutions, programming tasks and code solutions, and instruction-formatted text.
|
|
@@ -169,4 +171,4 @@ Please report model quality, risk, security vulnerabilities or Amazon AI Concern
|
|
| 169 |
|
| 170 |
The released model weights are licensed under [CC BY-NC 4.0](LICENSE). Use and redistribution must comply with the license, including its attribution and noncommercial requirements.
|
| 171 |
|
| 172 |
-
ALoDLM-1.7B is derived from [Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B), and ALoDLM-8B is derived from [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B). Both base models were developed by the Qwen team and released under Apache 2.0. The [upstream Apache 2.0 license](licenses/Qwen-Apache-2.0.txt) and [attribution notice](NOTICE) accompany each model. Retain applicable upstream license and attribution notices when redistributing the models.
|
|
|
|
| 2 |
license: cc-by-nc-4.0
|
| 3 |
pipeline_tag: text-generation
|
| 4 |
tags:
|
| 5 |
+
- alodlm
|
| 6 |
+
- diffusion-language-model
|
| 7 |
+
- looped-transformer
|
| 8 |
+
- adaptive-computation
|
| 9 |
+
- text-generation
|
| 10 |
+
- code
|
| 11 |
---
|
| 12 |
|
| 13 |
# ALoDLM: Adaptively Looped Diffusion Language Models
|
| 14 |
|
| 15 |
+
**Paper:** [ALoDLM: Adaptively Looped Diffusion Language Models](https://huggingface.co/papers/2610.04198) · **Project page:** [https://alo-dlm.github.io/](https://alo-dlm.github.io/)
|
| 16 |
+
|
| 17 |
ALoDLM is a family of diffusion language models developed at Amazon, available as **ALoDLM-1.7B** and **ALoDLM-8B** and initialized from the corresponding Qwen3 backbones. The models generate text and code by repeatedly applying shared transformer layers, preserving unresolved tokens' latent states, and adaptively allocating computation across token positions.
|
| 18 |
|
| 19 |
Training data includes mathematical problems and worked solutions, programming tasks and code solutions, and instruction-formatted text.
|
|
|
|
| 171 |
|
| 172 |
The released model weights are licensed under [CC BY-NC 4.0](LICENSE). Use and redistribution must comply with the license, including its attribution and noncommercial requirements.
|
| 173 |
|
| 174 |
+
ALoDLM-1.7B is derived from [Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B), and ALoDLM-8B is derived from [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B). Both base models were developed by the Qwen team and released under Apache 2.0. The [upstream Apache 2.0 license](licenses/Qwen-Apache-2.0.txt) and [attribution notice](NOTICE) accompany each model. Retain applicable upstream license and attribution notices when redistributing the models.
|