Add model card
Browse filesThis PR adds a model card for DLM-AN. It links the paper [Controllable Accent Normalization via Discrete Diffusion](https://huggingface.co/papers/2603.14275), the project page, and the GitHub repository, and sets the `pipeline_tag` to `audio-to-audio` so the model is discoverable under audio-to-audio models.
README.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
pipeline_tag: audio-to-audio
|
| 3 |
+
---
|
| 4 |
+
|
| 5 |
+
# DLM-AN: Controllable Accent Normalization via Discrete Diffusion
|
| 6 |
+
|
| 7 |
+
This repository contains the pretrained checkpoints for **DLM-AN**, the inference release for the paper [*Controllable Accent Normalization via Discrete Diffusion*](https://huggingface.co/papers/2603.14275) (Interspeech 2026).
|
| 8 |
+
|
| 9 |
+
**DLM-AN** (Diffusion Language Model for Accent Normalization) is a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. It provides tunable accent strength through a Common Token Predictor that selectively reuses source tokens, and includes a flow-matching Duration Ratio Predictor to better match native rhythm.
|
| 10 |
+
|
| 11 |
+
- Project page: https://P1ping.github.io/dlman-demo
|
| 12 |
+
- Code: https://github.com/P1ping/DLM-AN
|