Consideration of fine-tuning on DiffusionGemma for faster inference?

#2
by bandageshi - opened

Hi

Thanks a lot for releasing this model!

I was wondering if you have considered experimenting with fine-tuning on DiffusionGemma to improve the rewriting speed. I’d love to hear your thoughts on this direction if it's something you've looked into.

Thank you for the kind words and great suggestion!

I haven't experimented with DiffusionGemma yet. It's a neat idea for speed improvements. But there's a few reasons why this isn't on my near term roadmap:

  1. I use DPO and GRPO during training which requires token level log probability of an autoregressive model. So I'd need different RL/preference methods for diffusion LMs (like diffu-GRPO, etc), which would require rewriting a large part of my pipeline (it's not just a fine-tune)

  2. My focus is fact fidelity (preserving all numbers, names, dates and original intent from the text). I'd like to confirm parallel denoising will hold for fact fidelity before downgrading quality for speed. Also DiffusionGemma trails Gemma 4 12B on benchmarks

  3. Memory consumption for DiffusionGemma is around 18GB. But my Q4_K_M is around 7.6GB. And I'm developing even smaller 12B quants (4-6 GB) for lower RAM systems.

For speed, my nearer-term roadmap is MTP (multi-token prediction) and speculative decoding with DFlash / DSpark. These keep the same model, so output quality stays the same, while generation gets faster.

But yeah, I'm interested in faster local rewrites. So I'll be following it as tooling gets developed. Let me know if you experiment with it!

Sign up or log in to comment