Papers
arxiv:2610.01016

Scaling and Distilling Text Embeddings for Better Diffusibility

Published on Oct 1
· Submitted by
Zekai Zhang
on Oct 2
Authors:
,
,
,
,
,
,

Abstract

Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.

Community

Paper submitter

Glad to share our recent paper on the latent space for continuous diffusion language models (DLMs)!

In this paper, we find that scaling text embeddings can greatly boost the performance of continuous DLMs; for instance, by replacing the T5-small embeddings used in the recent ELF models with the advanced T5Gemma-2-270M, we can reduce Gen. PPL by about 40%.

But scaling alone is not enough, as the scaled embeddings can be hard to generate. They are so distinctive and informative that embeddings of similar and interchangeable words are far apart. Thus, continuous diffusion struggles to generate such separated and discrete targets. To mitigate this, we distill them into a more connected and robust latent space, making them easier for diffusion to generate.

As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL. Our results shed light on language representation learning, and also on a reliable way to scale DLMs.

Code and project page are released too: code | project page. Welcome to check them out!

Sign up or log in to comment

Models citing this paper 5

Browse 5 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.01016 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.01016 in a Space README.md to link it from this page.

Collections including this paper 1