We're deeply intrigued by Diffusion and Large Language models
Qwen-Image-2.1 editing with a 0.8B text encoder