An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
Abstract
Researchers propose a latent-to-pixel training strategy that accelerates convergence and improves inference speed for large-scale pixel-space diffusion models.
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.
Community
Z-Image-Pixel & Empirical Insight of Training Pixel-Space Diffusion Models
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion (2026)
- Pixel-Space Diffusion Transformers (2026)
- V-RAE: Rethinking Video Latent Spaces for Generation (2026)
- Where Does Generative Difficulty Reside? An Empirical Study of Target Representations (2026)
- DiT-Reward: Generative Representations for Text-to-Image Reward Modeling (2026)
- UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction (2026)
- Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.16887 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper