bench-labs/PixelModel-v5
Text-to-Image • 0.2B • Updated • 2
ai models that create images from text prompts
Note almost the same as v4 but has more training data
Note a tiny latent diffusion transformer, and the result is a roughly 10x jump in FID.
Note New architecture, not a scale-up. SIREN decoder + FiLM + learned embeddings. Beats v1 FID.
Note scaled up
Note x8.5 smaller, more efficient, better prompt understanding and larger training dataset
Note The first of its kind.