Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Distinctive Image Captioning via CLIP Guided Group Optimization

Aug 14, 2022

Youyuan Zhang, Jiuniu Wang, Hao Wu, Wenjia Xu

Figure 1 for Distinctive Image Captioning via CLIP Guided Group Optimization

Figure 2 for Distinctive Image Captioning via CLIP Guided Group Optimization

Figure 3 for Distinctive Image Captioning via CLIP Guided Group Optimization

Figure 4 for Distinctive Image Captioning via CLIP Guided Group Optimization

Share this with someone who'll enjoy it:

Abstract:Image captioning models are usually trained according to human annotated ground-truth captions, which could generate accurate but generic captions. In this paper, we focus on generating the distinctive captions that can distinguish the target image from other similar images. To evaluate the distinctiveness of captions, we introduce a series of metrics that use large-scale vision-language pre-training model CLIP to quantify the distinctiveness. To further improve the distinctiveness of captioning models, we propose a simple and effective training strategy which trains the model by comparing target image with similar image group and optimizing the group embedding gap. Extensive experiments are conducted on various baseline models to demonstrate the wide applicability of our strategy and the consistency of metric results with human evaluation. By comparing the performance of our best model with existing state-of-the-art models, we claim that our model achieves new state-of-the-art towards distinctiveness objective.

View paper on

Share this with someone who'll enjoy it:

Title:Distinctive Image Captioning via CLIP Guided Group Optimization

Paper and Code