bench-labs/PixelModel-v4
Text-to-Image • 40.1M • Updated • 16 • 5
ai models that create images from text prompts
Note a tiny latent diffusion transformer, and the result is a roughly 10x jump in FID.
Note New architecture, not a scale-up. SIREN decoder + FiLM + learned embeddings. Beats v1 FID.
Note scaled up
Note x8.5 smaller, more efficient, better prompt understanding and larger training dataset
Note The first of its kind.