Diffusion Models for High-Fidelity Image Synthesis: From Score Matching to Latent Generative Modeling
Keywords:
Diffusion models, GEN AI, text-to-image generation, DDIMAbstract
Deep generative modeling underwent a substantial transition by 2022 as diffusion-based approaches emerged as strong alternatives to generative adversarial networks, variational autoencoders, and autoregressive image models. Diffusion models formulate generation as the reversal of a gradual noise-corruption process, allowing complex data distributions to be learned through repeated denoising operations. This review examines the development of diffusion-based generative modeling through 2022, beginning with nonequilibrium thermodynamic formulations and denoising score matching and progressing toward denoising diffusion probabilistic models, score-based stochastic differential equations, deterministic implicit sampling, classifier guidance, and latent diffusion. Denoising diffusion probabilistic models demonstrated high-quality synthesis using parameterized reverse diffusion, while improved parameterizations and classifier guidance substantially enhanced sample fidelity. Score-based generative modeling unified several related approaches through stochastic differential equations. Denoising diffusion implicit models reduced sampling requirements by defining non-Markovian reverse processes. Subsequent classifier-free conditioning enabled controllable generation without a separate classifier. Latent diffusion further addressed the substantial computational cost of pixel-space models by performing the generative process within compressed perceptual representations. Concurrent text-conditioned systems illustrated the capacity of diffusion architectures to connect natural-language representations with high-resolution image synthesis. The review compares these methods in terms of likelihood formulation, sampling efficiency, conditioning mechanisms, image fidelity, latent representation, and computational requirements. Challenges remaining in 2022 included slow iterative inference, expensive large-scale training, evaluation of semantic alignment, dataset bias, memorization, and the computational accessibility of high-resolution generation.
References
Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S. Deep unsupervised learning using nonequilibrium thermodynamics. Proc Mach Learn Res. 2015;37:2256-2265.
Vincent P. A connection between score matching and denoising autoencoders. Neural Comput. 2011;23(7):1661-1674.
Song Y, Ermon S. Generative modeling by estimating gradients of the data distribution. Adv Neural Inf Process Syst. 2019;32:11918-11930.
Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Adv Neural Inf Process Syst. 2020;33:6840-6851.
Song J, Meng C, Ermon S. Denoising diffusion implicit models. In: International Conference on Learning Representations. 2021.
Song Y, Sohl-Dickstein J, Kingma DP, Kumar A, Ermon S, Poole B. Score-based generative modeling through stochastic differential equations. In: International Conference on Learning Representations. 2021.
Nichol AQ, Dhariwal P. Improved denoising diffusion probabilistic models. Proc Mach Learn Res. 2021;139:8162-8171.
Dhariwal P, Nichol A. Diffusion models beat GANs on image synthesis. Adv Neural Inf Process Syst. 2021;34:8780-8794.
Ho J, Salimans T. Classifier-free diffusion guidance. arXiv. 2022;2207.12598.
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. p. 10684-10695.
Ramesh A, Dhariwal P, Nichol A, Chu C, Chen M. Hierarchical text-conditional image generation with CLIP latents. arXiv. 2022;2204.06125.
Saharia C, Chan W, Saxena S, Li L, Whang J, Denton E, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv. 2022;2205.11487.
Nichol A, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, et al. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. arXiv. 2021;2112.10741.
Ho J, Saharia C, Chan W, Fleet DJ, Norouzi M, Salimans T. Cascaded diffusion models for high fidelity image generation. J Mach Learn Res. 2022;23(47):1-33.
Choi J, Kim S, Jeong Y, Gwon Y, Yoon S. ILVR: conditioning method for denoising diffusion probabilistic models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 14367-14376.
Song Y, Ermon S. Improved techniques for training score-based generative models. Adv Neural Inf Process Syst. 2020;33:12438-12448.
Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, et al. Generative adversarial nets. Adv Neural Inf Process Syst. 2014;27:2672-2680.
Kingma DP, Welling M. Auto-encoding variational Bayes. In: International Conference on Learning Representations. 2014.
Esser P, Rombach R, Ommer B. Taming Transformers for high-resolution image synthesis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021. p. 12873-12883.
Karras T, Laine S, Aittala M, Hellsten J, Lehtinen J, Aila T. Analyzing and improving the image quality of StyleGAN. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. p. 8110-8119.