Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Guiding Text-to-Image Diffusion Model Towards Grounded Generation

Jan 12, 2023

Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, Weidi Xie

Figure 1 for Guiding Text-to-Image Diffusion Model Towards Grounded Generation

Figure 2 for Guiding Text-to-Image Diffusion Model Towards Grounded Generation

Figure 3 for Guiding Text-to-Image Diffusion Model Towards Grounded Generation

Figure 4 for Guiding Text-to-Image Diffusion Model Towards Grounded Generation

Share this with someone who'll enjoy it:

Abstract:The goal of this paper is to augment a pre-trained text-to-image diffusion model with the ability of open-vocabulary objects grounding, i.e., simultaneously generating images and segmentation masks for the corresponding visual entities described in the text prompt. We make the following contributions: (i) we insert a grounding module into the existing diffusion model, that can be trained to align the visual and textual embedding space of the diffusion model with only a small number of object categories; (ii) we propose an automatic pipeline for constructing a dataset, that consists of {image, segmentation mask, text prompt} triplets, to train the proposed grounding module; (iii) we evaluate the performance of open-vocabulary grounding on images generated from the text-to-image diffusion model and show that the module can well segment the objects of categories beyond seen ones at training time; (iv) we adopt the guided diffusion model to build a synthetic semantic segmentation dataset, and show that training a standard segmentation model on such dataset demonstrates competitive performance on zero-shot segmentation(ZS3) benchmark, which opens up new opportunities for adopting the powerful diffusion model for discriminative tasks.

View paper on

Share this with someone who'll enjoy it:

Title:Guiding Text-to-Image Diffusion Model Towards Grounded Generation

Paper and Code