Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:EPIC: Efficient Prompt Interaction for Text-Image Classification

Jul 10, 2025

Xinyao Yu, Hao Sun, Zeyu Ling, Ziwei Niu, Zhenjia Bai, Rui Qin, Yen-Wei Chen, Lanfen Lin

Figure 1 for EPIC: Efficient Prompt Interaction for Text-Image Classification

Figure 2 for EPIC: Efficient Prompt Interaction for Text-Image Classification

Figure 3 for EPIC: Efficient Prompt Interaction for Text-Image Classification

Figure 4 for EPIC: Efficient Prompt Interaction for Text-Image Classification

Share this with someone who'll enjoy it:

Abstract:In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tasks, such as text-image classification. The growing size of LMMs, however, results in a significant computational cost for fine-tuning these models for downstream tasks. Hence, prompt-based interaction strategy is studied to align modalities more efficiently. In this context, we propose a novel efficient prompt-based multimodal interaction strategy, namely Efficient Prompt Interaction for text-image Classification (EPIC). Specifically, we utilize temporal prompts on intermediate layers, and integrate different modalities with similarity-based prompt interaction, to leverage sufficient information exchange between modalities. Utilizing this approach, our method achieves reduced computational resource consumption and fewer trainable parameters (about 1\% of the foundation model) compared to other fine-tuning strategies. Furthermore, it demonstrates superior performance on the UPMC-Food101 and SNLI-VE datasets, while achieving comparable performance on the MM-IMDB dataset.

* arXiv admin note: substantial text overlap with arXiv:2401.14856

View paper on

Share this with someone who'll enjoy it:

Title:EPIC: Efficient Prompt Interaction for Text-Image Classification

Paper and Code