Controllable and Personalized Text-to-Image Diffusion Models: Spatial Conditioning, Subject Adaptation, and Semantic Image Editing
Keywords:
latent diffusion, personalization, text-to-image synthesis, generative AI, image editingAbstract
Diffusion-based text-to-image generation achieved substantial advances in visual quality and semantic expressiveness by 2022, but early systems provided users with comparatively limited control over spatial organization, identity preservation, object placement, and local image editing. Research during 2022 and 2023 increasingly shifted toward controllable and personalized diffusion systems capable of adapting large pretrained generative models without retraining them from scratch. This review examines major methods for controllable image synthesis available through 2023. Latent Diffusion Models established computationally efficient text-conditioned generation within compressed latent spaces, while classifier-free guidance supported flexible text conditioning. SDEdit demonstrated image transformation through stochastic differential equations, and ILVR introduced conditioning using reference-image information. Textual Inversion learned new concepts through optimized text embeddings, whereas DreamBooth adapted pretrained diffusion models to specific subjects from only a few reference images. Prompt-to-Prompt enabled semantic editing by manipulating cross-attention maps, while Null-text Inversion improved reconstruction and editing of real images. InstructPix2Pix introduced instruction-based image editing using paired synthetic supervision. Custom Diffusion reduced the parameter requirements of multi-concept personalization, while ControlNet provided explicit structural conditioning based on edges, depth maps, human poses, segmentation maps, and related signals. This review compares fine-tuning, embedding optimization, attention manipulation, inversion, structural conditioning, and plug-and-play approaches. Challenges include identity drift, concept overfitting, compositional failure, unauthorized personalization, memorization, computational cost, and reliable preservation of non-target image content.
References
Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. Adv Neural Inf Process Syst. 2020;33:6840-6851.
Song Y, Sohl-Dickstein J, Kingma DP, Kumar A, Ermon S, Poole B. Score-based generative modeling through stochastic differential equations. In: International Conference on Learning Representations. 2021.
Dhariwal P, Nichol A. Diffusion models beat GANs on image synthesis. Adv Neural Inf Process Syst. 2021;34:8780-8794.
Song J, Meng C, Ermon S. Denoising diffusion implicit models. In: International Conference on Learning Representations. 2021.
Choi J, Kim S, Jeong Y, Gwon Y, Yoon S. ILVR: conditioning method for denoising diffusion probabilistic models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. p. 14367-14376.
Meng C, He Y, Song Y, Song J, Wu J, Zhu JY, Ermon S. SDEdit: guided image synthesis and editing with stochastic differential equations. In: International Conference on Learning Representations. 2022.
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. p. 10684-10695.
Ho J, Salimans T. Classifier-free diffusion guidance. arXiv. 2022;2207.12598.
Nichol A, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, et al. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. arXiv. 2021;2112.10741.
Saharia C, Chan W, Saxena S, Li L, Whang J, Denton E, et al. Photorealistic text-to-image diffusion models with deep language understanding. Adv Neural Inf Process Syst. 2022;35:36479-36494.
Gal R, Alaluf Y, Atzmon Y, Patashnik O, Bermano AH, Chechik G, Cohen-Or D. An image is worth one word: personalizing text-to-image generation using textual inversion. In: International Conference on Learning Representations. 2023.
Ruiz N, Li Y, Jampani V, Pritch Y, Rubinstein M, Aberman K. DreamBooth: fine tuning text-to-image diffusion models for subject-driven generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. p. 22500-22510.
Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D. Prompt-to-Prompt image editing with cross attention control. In: International Conference on Learning Representations. 2023.
Mokady R, Hertz A, Aberman K, Pritch Y, Cohen-Or D. Null-text inversion for editing real images using guided diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. p. 6038-6047.
Brooks T, Holynski A, Efros AA. InstructPix2Pix: learning to follow image editing instructions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. p. 18392-18402.
Kumari N, Zhang B, Zhang R, Shechtman E, Zhu JY. Multi-concept customization of text-to-image diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. p. 1931-1941.
Zhang L, Rao A, Agrawala M. Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. p. 3836-3847.
Tumanyan N, Geyer M, Bagon S, Dekel T. Plug-and-Play diffusion features for text-driven image-to-image translation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. p. 1921-1930.
Chefer H, Alaluf Y, Vinker Y, Wolf L, Cohen-Or D. Attend-and-Excite: attention-based semantic guidance for text-to-image diffusion models. ACM Trans Graph. 2023;42(4):148.
Kawar B, Elad M, Ermon S, Song J. Denoising diffusion restoration models. Adv Neural Inf Process Syst. 2022;35:23593-23606.