Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Tianwei Zhao

Can Vision Language Models Infer Human Gaze Direction? A Controlled Study

Jun 04, 2025

Zory Zhang, Pinyuan Feng, Bingyang Wang, Tianwei Zhao, Suyang Yu, Qingying Gao, Hokin Deng, Ziqiao Ma, Yijiang Li, Dezhi Luo

Figure 1 for Can Vision Language Models Infer Human Gaze Direction? A Controlled Study

Figure 2 for Can Vision Language Models Infer Human Gaze Direction? A Controlled Study

Figure 3 for Can Vision Language Models Infer Human Gaze Direction? A Controlled Study

Figure 4 for Can Vision Language Models Infer Human Gaze Direction? A Controlled Study

Abstract:Gaze-referential inference--the ability to infer what others are looking at--is a critical component of a theory of mind that underpins natural human-AI interaction. In a controlled study, we evaluated this skill across 111 Vision Language Models (VLMs) using photos taken with manipulated difficulty and variability, comparing performance with that of human participants (N = 65), and analyzed behaviors using mixed-effects models. We found that 94 of the 111 VLMs failed to do better than random guessing, while humans achieved near-ceiling accuracy. VLMs even respond with each choice almost equally frequently. Are they randomly guessing? Although most VLMs struggle, when we zoom in on five of the top-tier VLMs with above-chance performance, we find that their performance declined with increasing task difficulty but varied only slightly across different prompts and scene objects. These behavioral features cannot be explained by considering them as random guessers. Instead, they likely use a combination of heuristics and guessing such that their performance is subject to the task difficulty but robust to perceptual variations. This suggests that VLMs, lacking gaze inference capability, have yet to become technologies that can naturally interact with humans, but the potential remains.

* Preprint under review. Project page at https://grow-ai-like-a-child.github.io/gaze/

Via

Access Paper or Ask Questions

Machine Psychophysics: Cognitive Control in Vision-Language Models

May 25, 2025

Dezhi Luo, Maijunxian Wang, Bingyang Wang, Tianwei Zhao, Yijiang Li, Hokin Deng

Figure 1 for Machine Psychophysics: Cognitive Control in Vision-Language Models

Figure 2 for Machine Psychophysics: Cognitive Control in Vision-Language Models

Figure 3 for Machine Psychophysics: Cognitive Control in Vision-Language Models

Figure 4 for Machine Psychophysics: Cognitive Control in Vision-Language Models

Abstract:Cognitive control refers to the ability to flexibly coordinate thought and action in pursuit of internal goals. A standard method for assessing cognitive control involves conflict tasks that contrast congruent and incongruent trials, measuring the ability to prioritize relevant information while suppressing interference. We evaluate 108 vision-language models on three classic conflict tasks and their more demanding "squared" variants across 2,220 trials. Model performance corresponds closely to human behavior under resource constraints and reveals individual differences. These results indicate that some form of human-like executive function have emerged in current multi-modal foundational models.

Via

Access Paper or Ask Questions

Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Mar 25, 2025

Max W. Y. Lam, Yijin Xing, Weiya You, Jingcheng Wu, Zongyu Yin, Fuqiang Jiang, Hangyu Liu, Feng Liu, Xingda Li, Wei-Tsung Lu(+7 more)

Figure 1 for Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Figure 2 for Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Figure 3 for Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Figure 4 for Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Abstract:Autoregressive (AR) models have demonstrated impressive capabilities in generating high-fidelity music. However, the conventional next-token prediction paradigm in AR models does not align with the human creative process in music composition, potentially compromising the musicality of generated samples. To overcome this limitation, we introduce MusiCoT, a novel chain-of-thought (CoT) prompting technique tailored for music generation. MusiCoT empowers the AR model to first outline an overall music structure before generating audio tokens, thereby enhancing the coherence and creativity of the resulting compositions. By leveraging the contrastive language-audio pretraining (CLAP) model, we establish a chain of "musical thoughts", making MusiCoT scalable and independent of human-labeled data, in contrast to conventional CoT methods. Moreover, MusiCoT allows for in-depth analysis of music structure, such as instrumental arrangements, and supports music referencing -- accepting variable-length audio inputs as optional style references. This innovative approach effectively addresses copying issues, positioning MusiCoT as a vital practical method for music prompting. Our experimental results indicate that MusiCoT consistently achieves superior performance across both objective and subjective metrics, producing music quality that rivals state-of-the-art generation models. Our samples are available at https://MusiCoT.github.io/.

* Preprint

Via

Access Paper or Ask Questions